Agent skill

Centia Snapshot Catalog

by mapcentia in mapcentia/geocloud2

Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot…

AGPL-3.0Auto-check passedData & Analytics

Install Centia Snapshot Catalog

skills CLI
$ npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mapcentia/geocloud2 centia-snapshot-catalog --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mapcentia/geocloud2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/centia-snapshot-catalog .claude/skills/centia-snapshot-catalog && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
centia-snapshot-catalog
GitHub stars
152
Token cost
~2.3k tokens
SKILL.md length
601 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot…

  • Works in 8 steps: Layout → Setup → Find datasets → …
  • Tasks that involve DataFrames
  • SKILL.md covers 1. Layout, 2. Setup, 3. Find datasets and 4. Does the dataset have…, plus 4 more sections
  • Reaches gc2-parquet.s3.eu-west-1.amazonaws.com; needs BEARER_TOKEN

What it does

Centia Snapshot Catalog is an agent skill from mapcentia/geocloud2. Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot, the newest snapshot or the whole history over Hive partitions, and run spatial and non-spatial SQL on them.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering DataFrames. It works with DuckDB and SQL. The repository describes itself as: The GC2 framework helps you build a spatial data infrastructure quickly and easily. Powered using open source components for a scalable solution focused on freedom rather than… The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve DataFrames

Example prompts

  • “/centia-snapshot-catalog”

Requirements

  • A credential in BEARER_TOKEN

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Layout
  2. Setup
  3. Find datasets
  4. Does the dataset have geometry?
  5. Read the data
  6. Spatial analysis
  7. FlatGeobuf assets
  8. Pitfalls

What it can do on your machine

Read from SKILL.md and the folder at commit a18093f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gc2-parquet.s3.eu-west-1.amazonaws.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BEARER_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Centia Snapshot Catalog loads about 2.3k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 601 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mapcentia/geocloud2 at commit a18093f, republished under its AGPL-3.0 licence (© mapcentia). 601 words, ~2,308 tokens.

Download SKILL.mdSave it as .claude/skills/centia-snapshot-catalog/SKILL.md (or your agent's skills folder).
name
centia-snapshot-catalog
description
Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot, the newest snapshot or the whole history over Hive partitions, and run spatial and non-spatial SQL on them.

Centia snapshot catalog + DuckDB

A GC2 snapshot store is a static STAC 1.1.0 catalog over (Geo)Parquet files. Everything an agent needs to plan an analysis is in three small JSON files; the data itself is read directly with DuckDB over HTTP or S3.

Example root: https://gc2-parquet.s3.eu-west-1.amazonaws.com/centia-io/dk/catalog.json ({base} below is the directory that holds catalog.json).

1. Layout

{base}/catalog.json                                   Catalog: one "child" link per dataset
{base}/schema=S/relation=R/collection.json            Collection: title, extent, CRS, one "item" link per snapshot date
{base}/schema=S/relation=R/latest.json                Pointer to the newest snapshot (newer stores only)
{base}/schema=S/relation=R/_gc2_snapshot_date=D/item.json          Item: one snapshot
{base}/schema=S/relation=R/_gc2_snapshot_date=D/data-<uuid>.parquet   the data (GeoParquet when spatial)
{base}/schema=S/relation=R/_gc2_snapshot_date=D/data-<uuid>.fgb       optional FlatGeobuf
{base}/schema=S/relation=R/_gc2_snapshot_date=D/metadata-<uuid>.json  column schema, row count, crs, bbox, formats
  • A dataset is schema.relation (a PostGIS table or view). Every href is relative to the file it appears in.
  • _gc2_snapshot_date=D is a Hive partition key: reading several dates at once with hive_partitioning = true adds a _gc2_snapshot_date column.
  • Only JSON files live outside the partitions, so *.parquet globs are safe.

2. Setup

sql
INSTALL httpfs; LOAD httpfs;
INSTALL spatial; LOAD spatial;          -- GEOMETRY type, ST_* functions, GeoParquet awareness
-- Public bucket, anonymous access (needed for globs; HTTP URLs cannot glob):
CREATE SECRET pub (TYPE s3, REGION 'eu-west-1', KEY_ID '', SECRET '');

Single files work with plain https://… URLs. Wildcards need the s3://bucket/prefix/… form (the bucket is the first host label of the HTTPS URL: s3://gc2-parquet/centia-io/dk/…).

3. Find datasets

Read the STAC JSON with DuckDB instead of guessing paths:

sql
-- every dataset in the store
SELECT l.title AS dataset, l.href
FROM (SELECT unnest(links) AS l FROM read_json('{base}/catalog.json'))
WHERE l.rel = 'child';

-- one dataset: extent, CRS, snapshot dates
SELECT id, title, description, keywords,
       extent.spatial.bbox[1]              AS bbox_wgs84,
       extent.temporal.interval[1]         AS first_last_snapshot,
       summaries                           -- {"gc2:schema_version": [...], "proj:code": ["EPSG:25832"]}
FROM read_json('{base}/schema=S/relation=R/collection.json');

SELECT l.href FROM (SELECT unnest(links) AS l FROM read_json('{base}/schema=S/relation=R/collection.json'))
WHERE l.rel = 'item' ORDER BY l.href DESC;          -- newest first is how they are listed

The item tells you what one snapshot holds:

sql
SELECT id, bbox, geometry.type AS footprint,
       properties."gc2:row_count"    AS rows,
       properties."proj:code"        AS crs,
       properties.datetime           AS snapshot_date,
       assets.data.href, assets.data.title,          -- "GeoParquet" or "Parquet"
       assets.flatgeobuf.href                        -- NULL when not produced
FROM read_json('{base}/schema=S/relation=R/_gc2_snapshot_date=D/item.json');

Title, description and keywords come from the layer metadata in GC2; when a layer has no title the collection is titled schema.relation.

4. Does the dataset have geometry?

Check in this order; stop at the first definitive answer.

  1. Item (item.json): bbox and a Polygon geometry present → spatial; geometry: null and no bbox → non-spatial. assets.data.title is GeoParquet for spatial data and Parquet otherwise. properties."proj:code" gives the CRS (EPSG:25832 etc.). Items published before September 2026 may lack proj:code; the collection extent.spatial.bbox of [-180,-90,180,90] then means "unknown or non-spatial", not "world-wide data".
  2. metadata-<uuid>.json (next to the data): crs ("EPSG:25832" or null), bbox, and schema[] where a column with data_type starting with geometry( or geography( is the geometry column.
  3. The Parquet file itself (definitive, one HTTP range request):
sql
-- GeoParquet metadata: present only for spatial files
SELECT json_extract_string(decode(value), '$.primary_column')                       AS geom_column,
       json_extract_string(decode(value), '$.columns.the_geom.crs.id.code')          AS epsg,
       json_extract_string(decode(value), '$.columns.the_geom.geometry_types')       AS geometry_types
FROM parquet_kv_metadata('{file}.parquet') WHERE key = 'geo';      -- 0 rows => non-spatial

-- or simply look at the column type DuckDB infers (spatial extension loaded)
SELECT column_name, column_type FROM (DESCRIBE SELECT * FROM read_parquet('{file}.parquet'))
WHERE column_type LIKE 'GEOMETRY%';                                -- e.g. GEOMETRY('EPSG:25832')

GC2 names the geometry column the_geom in imported tables; other tables may use another name, so read primary_column rather than assuming.

5. Read the data

sql
-- one snapshot (HTTP is fine for single files)
SELECT count(*) FROM read_parquet('{base}/schema=S/relation=R/_gc2_snapshot_date=D/data-<uuid>.parquet');

-- the newest snapshot through the GC2 API (fixed URL, needs a Bearer token unless the layer is public)
CREATE SECRET api (TYPE http, BEARER_TOKEN '<token>');
SELECT * FROM read_parquet('https://<gc2-host>/api/v4/schemas/S/relations/R/snapshots/latest/data') LIMIT 10;

-- the newest snapshot from the store: follow latest.json (newer stores) or the first item link
SELECT item, assets.data.href FROM read_json('{base}/schema=S/relation=R/latest.json');

-- the whole history, one row per snapshot date (s3:// + secret from §2)
SELECT _gc2_snapshot_date, count(*)
FROM read_parquet('s3://gc2-parquet/centia-io/dk/schema=S/relation=R/_gc2_snapshot_date=*/data-*.parquet',
                  hive_partitioning = true)
GROUP BY 1 ORDER BY 1;

Rules of thumb:

  • Select only the columns you need; Parquet is columnar and the files are read over the network. properties."gc2:row_count" tells you the size before you read.
  • Prefer a dated URL when an analysis must be reproducible; latest moves after every publish.
  • Schema drift across dates: compare gc2:schema_version (a hash of column names and types) between items before UNIONing them; use union_by_name = true in read_parquet when columns differ.
  • Timestamps are stored as UTC TIMESTAMP WITH TIME ZONE; Time and binary columns were exported as strings.
Show full SKILL.md (211 more words)Show less

6. Spatial analysis

With the spatial extension the geometry column arrives as GEOMETRY('EPSG:xxxx'), so ST_* functions work directly:

sql
SELECT ST_GeometryType(the_geom) AS type, count(*) FROM read_parquet('{file}') GROUP BY 1;

-- area per category in the native CRS (planar units of the CRS, metres for EPSG:25832)
SELECT temanavn, sum(ST_Area(the_geom)) / 1e6 AS km2 FROM read_parquet('{file}') GROUP BY 1 ORDER BY 2 DESC;

-- reproject to WGS84 for a map / bbox filter; source and target CRS from proj:code / the geo metadata
SELECT ST_Transform(the_geom, 'EPSG:25832', 'EPSG:4326', always_xy := true) AS geom_wgs84 FROM read_parquet('{file}');

-- spatial join between two datasets in the same CRS
SELECT b.*, k.navn AS kommune
FROM read_parquet('{buildings}') b
JOIN read_parquet('{municipalities}') k ON ST_Intersects(b.the_geom, k.the_geom);

Join datasets in different CRSs only after ST_Transform-ing one of them. ST_Area/ST_Length are planar: use a projected CRS (EPSG:25832 for Denmark), never EPSG:4326, for metric results.

Non-spatial datasets (no geo metadata) are plain tables: join them to a spatial one on a key column (id, bfe, kommunekode, …) to map them.

7. FlatGeobuf assets

assets.flatgeobuf (when present) is the same rows as a .fgb file for web maps and GIS clients. In DuckDB use ST_Read('{file}.fgb'); for analysis prefer the Parquet asset, which is faster to scan and carries the same CRS.

8. Pitfalls

  • Plain HTTPS URLs cannot glob; use s3:// with an anonymous secret for _gc2_snapshot_date=* reads, or list the item hrefs from the collection.
  • parquet_kv_metadata values are BLOBs: wrap them in decode() before json_extract_string.
  • Do not create files or directories beside the partitions in the store (the latest.json/collection.json JSON files are the only exceptions); Hive readers treat every directory under relation=R/ as a partition.
  • The store lists every published relation of the database; access to the files is whatever the bucket allows, while the GC2 API (/api/v4/schemas/ {schema}/relations/{relation}/snapshots/…) enforces privileges and geofence rules. Use the API URLs when the data is not public.

© mapcentia, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/centia-snapshot-catalog of mapcentia/geocloud2.

Open the folder on GitHubat commit a18093f

Compare with similar skills

Centia Snapshot Catalog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Centia Snapshot Catalog compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Centia Snapshot Catalog this skillmapcentia/geocloud2152—~2.3kAutomated safety check: PassAGPL-3.0
Querying Big Datasetsflyrank-bih/flyrank-ml-internship-starter140—~750Automated safety check: PassCustom licence
Duckdb EnaAAaqwq/AGI-Super-Team1051 repos~1.6kAutomated safety check: PassMIT
Data Processingjeremylongshore/tons-of-skills-marketplace2.8k—~1.3kAutomated safety check: PassMIT
Chdb Datastorevemetric/vemetric3952 repos~1.4kAutomated safety check: PassApache-2.0
Dinobase Business Data Querieskappa90/dinobase263—~1.5kAutomated safety check: PassCustom licence

Similar skills

  • Querying Big Datasets

    flyrank-bih/flyrank-ml-internship-starter

    Works with datasets far too big to download or load in pandas — SQL over remote Parquet with DuckDB, aggregate-then-model, iterate on samples.

    140 GitHub stars~750 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Duckdb En

    aAAaqwq/AGI-Super-Team

    DuckDB CLI specialist for SQL analysis, data processing and file conversion.

    105 GitHub starsUsed in 1 repo~1.6k tokens
    DatabasesAuto-check passed
  • Data Processing

    jeremylongshore/tons-of-skills-marketplace

    A skill your agent uses when working with structured data files (CSV, JSON, YAML, TOML, Parquet) — querying, transforming, filtering, aggregating, or converting between formats

    2.8k GitHub stars~1.3k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    395 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.

    263 GitHub stars~1.5k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Ops Telemetry Query

    boundless-xyz/boundless

    Internal — for Boundless team members only. An agent skill from boundless-xyz/boundless.

    193 GitHub stars~3.8k tokensUpdated 1 mo ago
    DatabasesAuto-check passed

Works with

Questions about Centia Snapshot Catalog

What does Centia Snapshot Catalog do?

Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot…. Centia Snapshot Catalog is an agent skill from mapcentia/geocloud2.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot, the newest snapshot or the whole history over Hive partitions, and run spatial and non-spatial SQL on them.

When should I use Centia Snapshot Catalog?

Centia Snapshot Catalog fits situations like: tasks that involve DataFrames.

How do I install Centia Snapshot Catalog in Claude Code?

Run `npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a claude-code`. Or copy the skill folder (skills/centia-snapshot-catalog in mapcentia/geocloud2) into .claude/skills/centia-snapshot-catalog in your project. Claude Code loads it when a task matches its description.

How do I install Centia Snapshot Catalog in Codex?

Run `npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a codex`. Or copy the skill folder (skills/centia-snapshot-catalog in mapcentia/geocloud2) into .agents/skills/centia-snapshot-catalog in your project. Codex loads it when a task matches its description.

Can I use Centia Snapshot Catalog in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mapcentia/geocloud2 --skill centia-snapshot-catalog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/centia-snapshot-catalog, .gemini/skills/centia-snapshot-catalog, .github/skills/centia-snapshot-catalog and .opencode/skills/centia-snapshot-catalog in your project.

What does Centia Snapshot Catalog need to run?

Going by SKILL.md and its folder, Centia Snapshot Catalog needs credentials named BEARER_TOKEN. Our summary lists: A credential in BEARER_TOKEN.

Does Centia Snapshot Catalog access the network?

SKILL.md names 1 domain. In commands or code: gc2-parquet.s3.eu-west-1.amazonaws.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Centia Snapshot Catalog safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Centia Snapshot Catalog use?

Centia Snapshot Catalog is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Centia Snapshot Catalog use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Centia Snapshot Catalog?

Skills that share tags, products or a category with Centia Snapshot Catalog: Querying Big Datasets (flyrank-bih/flyrank-ml-internship-starter, 140 stars), Duckdb En (aAAaqwq/AGI-Super-Team, 105 stars), Data Processing (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Chdb Datastore (vemetric/vemetric, 395 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Centia Snapshot Catalog?

mapcentia (a GitHub user) maintains it in mapcentia/geocloud2, which has 152 GitHub stars. The repository was last updated on October 9, 2026.

Source: mapcentia/geocloud2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.