Querying Big Datasets
flyrank-bih/flyrank-ml-internship-starter
Works with datasets far too big to download or load in pandas — SQL over remote Parquet with DuckDB, aggregate-then-model, iterate on samples.
Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot…
$ npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mapcentia/geocloud2 centia-snapshot-catalog --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mapcentia/geocloud2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/centia-snapshot-catalog .claude/skills/centia-snapshot-catalog && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "centia-snapshot-catalog" agent skill from https://github.com/mapcentia/geocloud2/tree/master/skills/centia-snapshot-catalog into .claude/skills/centia-snapshot-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "centia-snapshot-catalog", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mapcentia/geocloud2/tree/master/skills/centia-snapshot-catalogType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mapcentia/geocloud2 centia-snapshot-catalog --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mapcentia/geocloud2.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/centia-snapshot-catalog .agents/skills/centia-snapshot-catalog && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "centia-snapshot-catalog" agent skill from https://github.com/mapcentia/geocloud2/tree/master/skills/centia-snapshot-catalog into .agents/skills/centia-snapshot-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "centia-snapshot-catalog", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mapcentia/geocloud2 centia-snapshot-catalog --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mapcentia/geocloud2.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/centia-snapshot-catalog .cursor/skills/centia-snapshot-catalog && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "centia-snapshot-catalog" agent skill from https://github.com/mapcentia/geocloud2/tree/master/skills/centia-snapshot-catalog into .cursor/skills/centia-snapshot-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "centia-snapshot-catalog", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mapcentia/geocloud2.git --path skills/centia-snapshot-catalog--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mapcentia/geocloud2 centia-snapshot-catalog --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mapcentia/geocloud2.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/centia-snapshot-catalog .gemini/skills/centia-snapshot-catalog && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "centia-snapshot-catalog" agent skill from https://github.com/mapcentia/geocloud2/tree/master/skills/centia-snapshot-catalog into .gemini/skills/centia-snapshot-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "centia-snapshot-catalog", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mapcentia/geocloud2 centia-snapshot-catalogInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mapcentia/geocloud2.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/centia-snapshot-catalog .github/skills/centia-snapshot-catalog && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "centia-snapshot-catalog" agent skill from https://github.com/mapcentia/geocloud2/tree/master/skills/centia-snapshot-catalog into .github/skills/centia-snapshot-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "centia-snapshot-catalog", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mapcentia/geocloud2 centia-snapshot-catalog --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mapcentia/geocloud2.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/centia-snapshot-catalog .opencode/skills/centia-snapshot-catalog && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "centia-snapshot-catalog" agent skill from https://github.com/mapcentia/geocloud2/tree/master/skills/centia-snapshot-catalog into .opencode/skills/centia-snapshot-catalog/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "centia-snapshot-catalog", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
centia-snapshot-catalogAnalyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot…
Centia Snapshot Catalog is an agent skill from mapcentia/geocloud2. Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot, the newest snapshot or the whole history over Hive partitions, and run spatial and non-spatial SQL on them.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering DataFrames. It works with DuckDB and SQL. The repository describes itself as: The GC2 framework helps you build a spatial data infrastructure quickly and easily. Powered using open source components for a scalable solution focused on freedom rather than… The licence is AGPL-3.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a18093f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
gc2-parquet.s3.eu-west-1.amazonaws.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
BEARER_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Centia Snapshot Catalog loads about 2.3k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 601 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mapcentia/geocloud2 at commit a18093f, republished under its AGPL-3.0 licence (© mapcentia). 601 words, ~2,308 tokens.
.claude/skills/centia-snapshot-catalog/SKILL.md (or your agent's skills folder).A GC2 snapshot store is a static STAC 1.1.0 catalog over (Geo)Parquet files. Everything an agent needs to plan an analysis is in three small JSON files; the data itself is read directly with DuckDB over HTTP or S3.
Example root: https://gc2-parquet.s3.eu-west-1.amazonaws.com/centia-io/dk/catalog.json
({base} below is the directory that holds catalog.json).
{base}/catalog.json Catalog: one "child" link per dataset
{base}/schema=S/relation=R/collection.json Collection: title, extent, CRS, one "item" link per snapshot date
{base}/schema=S/relation=R/latest.json Pointer to the newest snapshot (newer stores only)
{base}/schema=S/relation=R/_gc2_snapshot_date=D/item.json Item: one snapshot
{base}/schema=S/relation=R/_gc2_snapshot_date=D/data-<uuid>.parquet the data (GeoParquet when spatial)
{base}/schema=S/relation=R/_gc2_snapshot_date=D/data-<uuid>.fgb optional FlatGeobuf
{base}/schema=S/relation=R/_gc2_snapshot_date=D/metadata-<uuid>.json column schema, row count, crs, bbox, formatsschema.relation (a PostGIS table or view). Every href is
relative to the file it appears in._gc2_snapshot_date=D is a Hive partition key: reading several dates at
once with hive_partitioning = true adds a _gc2_snapshot_date column.*.parquet globs are safe.INSTALL httpfs; LOAD httpfs;
INSTALL spatial; LOAD spatial; -- GEOMETRY type, ST_* functions, GeoParquet awareness
-- Public bucket, anonymous access (needed for globs; HTTP URLs cannot glob):
CREATE SECRET pub (TYPE s3, REGION 'eu-west-1', KEY_ID '', SECRET '');Single files work with plain https://… URLs. Wildcards need the
s3://bucket/prefix/… form (the bucket is the first host label of the
HTTPS URL: s3://gc2-parquet/centia-io/dk/…).
Read the STAC JSON with DuckDB instead of guessing paths:
-- every dataset in the store
SELECT l.title AS dataset, l.href
FROM (SELECT unnest(links) AS l FROM read_json('{base}/catalog.json'))
WHERE l.rel = 'child';
-- one dataset: extent, CRS, snapshot dates
SELECT id, title, description, keywords,
extent.spatial.bbox[1] AS bbox_wgs84,
extent.temporal.interval[1] AS first_last_snapshot,
summaries -- {"gc2:schema_version": [...], "proj:code": ["EPSG:25832"]}
FROM read_json('{base}/schema=S/relation=R/collection.json');
SELECT l.href FROM (SELECT unnest(links) AS l FROM read_json('{base}/schema=S/relation=R/collection.json'))
WHERE l.rel = 'item' ORDER BY l.href DESC; -- newest first is how they are listedThe item tells you what one snapshot holds:
SELECT id, bbox, geometry.type AS footprint,
properties."gc2:row_count" AS rows,
properties."proj:code" AS crs,
properties.datetime AS snapshot_date,
assets.data.href, assets.data.title, -- "GeoParquet" or "Parquet"
assets.flatgeobuf.href -- NULL when not produced
FROM read_json('{base}/schema=S/relation=R/_gc2_snapshot_date=D/item.json');Title, description and keywords come from the layer metadata in GC2; when a
layer has no title the collection is titled schema.relation.
Check in this order; stop at the first definitive answer.
item.json): bbox and a Polygon geometry present → spatial;
geometry: null and no bbox → non-spatial. assets.data.title is
GeoParquet for spatial data and Parquet otherwise. properties."proj:code"
gives the CRS (EPSG:25832 etc.). Items published before September 2026 may
lack proj:code; the collection extent.spatial.bbox of [-180,-90,180,90]
then means "unknown or non-spatial", not "world-wide data".<uuid>.json (next to the data): crs ("EPSG:25832" or null),
bbox, and schema[] where a column with data_type starting with
geometry( or geography( is the geometry column.-- GeoParquet metadata: present only for spatial files
SELECT json_extract_string(decode(value), '$.primary_column') AS geom_column,
json_extract_string(decode(value), '$.columns.the_geom.crs.id.code') AS epsg,
json_extract_string(decode(value), '$.columns.the_geom.geometry_types') AS geometry_types
FROM parquet_kv_metadata('{file}.parquet') WHERE key = 'geo'; -- 0 rows => non-spatial
-- or simply look at the column type DuckDB infers (spatial extension loaded)
SELECT column_name, column_type FROM (DESCRIBE SELECT * FROM read_parquet('{file}.parquet'))
WHERE column_type LIKE 'GEOMETRY%'; -- e.g. GEOMETRY('EPSG:25832')GC2 names the geometry column the_geom in imported tables; other tables
may use another name, so read primary_column rather than assuming.
-- one snapshot (HTTP is fine for single files)
SELECT count(*) FROM read_parquet('{base}/schema=S/relation=R/_gc2_snapshot_date=D/data-<uuid>.parquet');
-- the newest snapshot through the GC2 API (fixed URL, needs a Bearer token unless the layer is public)
CREATE SECRET api (TYPE http, BEARER_TOKEN '<token>');
SELECT * FROM read_parquet('https://<gc2-host>/api/v4/schemas/S/relations/R/snapshots/latest/data') LIMIT 10;
-- the newest snapshot from the store: follow latest.json (newer stores) or the first item link
SELECT item, assets.data.href FROM read_json('{base}/schema=S/relation=R/latest.json');
-- the whole history, one row per snapshot date (s3:// + secret from §2)
SELECT _gc2_snapshot_date, count(*)
FROM read_parquet('s3://gc2-parquet/centia-io/dk/schema=S/relation=R/_gc2_snapshot_date=*/data-*.parquet',
hive_partitioning = true)
GROUP BY 1 ORDER BY 1;Rules of thumb:
properties."gc2:row_count" tells you the size
before you read.latest moves
after every publish.gc2:schema_version (a hash of
column names and types) between items before UNIONing them; use
union_by_name = true in read_parquet when columns differ.TIMESTAMP WITH TIME ZONE; Time and binary
columns were exported as strings.With the spatial extension the geometry column arrives as
GEOMETRY('EPSG:xxxx'), so ST_* functions work directly:
SELECT ST_GeometryType(the_geom) AS type, count(*) FROM read_parquet('{file}') GROUP BY 1;
-- area per category in the native CRS (planar units of the CRS, metres for EPSG:25832)
SELECT temanavn, sum(ST_Area(the_geom)) / 1e6 AS km2 FROM read_parquet('{file}') GROUP BY 1 ORDER BY 2 DESC;
-- reproject to WGS84 for a map / bbox filter; source and target CRS from proj:code / the geo metadata
SELECT ST_Transform(the_geom, 'EPSG:25832', 'EPSG:4326', always_xy := true) AS geom_wgs84 FROM read_parquet('{file}');
-- spatial join between two datasets in the same CRS
SELECT b.*, k.navn AS kommune
FROM read_parquet('{buildings}') b
JOIN read_parquet('{municipalities}') k ON ST_Intersects(b.the_geom, k.the_geom);Join datasets in different CRSs only after ST_Transform-ing one of them.
ST_Area/ST_Length are planar: use a projected CRS (EPSG:25832 for
Denmark), never EPSG:4326, for metric results.
Non-spatial datasets (no geo metadata) are plain tables: join them to a
spatial one on a key column (id, bfe, kommunekode, …) to map them.
assets.flatgeobuf (when present) is the same rows as a .fgb file for web
maps and GIS clients. In DuckDB use ST_Read('{file}.fgb'); for analysis
prefer the Parquet asset, which is faster to scan and carries the same CRS.
s3:// with an anonymous secret for
_gc2_snapshot_date=* reads, or list the item hrefs from the collection.parquet_kv_metadata values are BLOBs: wrap them in decode() before
json_extract_string.latest.json/collection.json JSON files are the only exceptions);
Hive readers treat every directory under relation=R/ as a partition./api/v4/schemas/ {schema}/relations/{relation}/snapshots/…) enforces privileges and
geofence rules. Use the API URLs when the data is not public.© mapcentia, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/centia-snapshot-catalog of mapcentia/geocloud2.
Open the folder on GitHubat commit a18093f
Centia Snapshot Catalog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Centia Snapshot Catalog this skillmapcentia/geocloud2 | 152 | — | ~2.3k | Automated safety check: Pass | AGPL-3.0 | |
| Querying Big Datasetsflyrank-bih/flyrank-ml-internship-starter | 140 | — | ~750 | Automated safety check: Pass | Custom licence | |
| Duckdb EnaAAaqwq/AGI-Super-Team | 105 | 1 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Data Processingjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Chdb Datastorevemetric/vemetric | 395 | 2 repos | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Dinobase Business Data Querieskappa90/dinobase | 263 | — | ~1.5k | Automated safety check: Pass | Custom licence |
flyrank-bih/flyrank-ml-internship-starter
Works with datasets far too big to download or load in pandas — SQL over remote Parquet with DuckDB, aggregate-then-model, iterate on samples.
aAAaqwq/AGI-Super-Team
DuckDB CLI specialist for SQL analysis, data processing and file conversion.
jeremylongshore/tons-of-skills-marketplace
A skill your agent uses when working with structured data files (CSV, JSON, YAML, TOML, Parquet) — querying, transforming, filtering, aggregating, or converting between formats
vemetric/vemetric
A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.
kappa90/dinobase
Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.
boundless-xyz/boundless
Internal — for Boundless team members only. An agent skill from boundless-xyz/boundless.
Categories
Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot…. Centia Snapshot Catalog is an agent skill from mapcentia/geocloud2.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot, the newest snapshot or the whole history over Hive partitions, and run spatial and non-spatial SQL on them.
Centia Snapshot Catalog fits situations like: tasks that involve DataFrames.
Run `npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a claude-code`. Or copy the skill folder (skills/centia-snapshot-catalog in mapcentia/geocloud2) into .claude/skills/centia-snapshot-catalog in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a codex`. Or copy the skill folder (skills/centia-snapshot-catalog in mapcentia/geocloud2) into .agents/skills/centia-snapshot-catalog in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/centia-snapshot-catalog, .gemini/skills/centia-snapshot-catalog, .github/skills/centia-snapshot-catalog and .opencode/skills/centia-snapshot-catalog in your project.
Going by SKILL.md and its folder, Centia Snapshot Catalog needs credentials named BEARER_TOKEN. Our summary lists: A credential in BEARER_TOKEN.
SKILL.md names 1 domain. In commands or code: gc2-parquet.s3.eu-west-1.amazonaws.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Centia Snapshot Catalog is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Centia Snapshot Catalog: Querying Big Datasets (flyrank-bih/flyrank-ml-internship-starter, 140 stars), Duckdb En (aAAaqwq/AGI-Super-Team, 105 stars), Data Processing (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Chdb Datastore (vemetric/vemetric, 395 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mapcentia (a GitHub user) maintains it in mapcentia/geocloud2, which has 152 GitHub stars. The repository was last updated on October 9, 2026.
Source: mapcentia/geocloud2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.