Agent skill

Dataset Metadata

by opensanctions in opensanctions/opensanctions

Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments).

MITAuto-check: notesData & Analytics

Install Dataset Metadata

skills CLI
$ npx skills add opensanctions/opensanctions --skill dataset-metadata -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install opensanctions/opensanctions dataset-metadata --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/opensanctions/opensanctions.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/dataset-metadata .claude/skills/dataset-metadata && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-metadata
GitHub stars
831
Token cost
~604 tokens
SKILL.md length
265 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments).

  • Works in 3 steps: Read → Check each field against the doc → Maintainer comments
  • The user asks to fix
  • SKILL.md covers 1. Read, 2. Check each field against…, 3. Maintainer comments and Verify
  • Calls python

What it does

Dataset Metadata is an agent skill from opensanctions/opensanctions. Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments). Use when the user asks to fix, improve or standardise a dataset's metadata.

Its SKILL.md is about 600 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics. The repository describes itself as: An open database of international sanctions data, persons of interest and politically exposed persons. The licence is MIT.

When your agent uses it

  • The user asks to fix
  • Standardise a datasets metadata

Example prompts

  • “/dataset-metadata”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Edit, Glob, Grep, Bash

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Read
  2. Check each field against the doc
  3. Maintainer comments

What it can do on your machine

Read from SKILL.md and the folder at commit 4499adf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Edit
    • Glob
    • Grep
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Metadata loads about 604 tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~604

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Edit, Glob, Grep, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from opensanctions/opensanctions at commit 4499adf, republished under its MIT licence (© opensanctions). 265 words, ~604 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-metadata/SKILL.md (or your agent's skills folder).
name
dataset-metadata
description
Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments). Use when the user asks to fix, improve or standardise a dataset's metadata.
allowed-tools
Read, Edit, Glob, Grep, Bash
argument-hint
[dataset .yml path]

Dataset metadata

Standardise the metadata of the dataset at $ARGUMENTS.

The rules live in zavod/docs/metadata.md — read it first and follow it. This skill is only the procedure for applying it; when a rule is unclear, read the doc, do not invent a convention. For a legislature/parliament PEP dataset, take the title, description and coverage.frequency patterns from the /legislature-metadata skill; the rest of this checklist applies as written.

1. Read

The .yml, then the crawler (entry_point, usually crawler.py) for scope facts: what the source covers, what is deliberately skipped, which lookups are not plain type lookups and what they do.

2. Check each field against the doc

  • title — the doc's Title rules.
  • summary — length and style per the doc's Summary rules.
  • description — the doc's Description rules, including its keep-out list.
  • coverage.frequency — the doc's house default for the dataset type.
  • tags — present and plausible for list type and target countries.
  • publisher — all subfields per the doc's Publisher section.
  • data — url fetchable; format and lang per the doc's Source data section.

3. Maintainer comments

  • Top-of-file # block: known failure modes, source quirks, recurring-warning runbooks.
  • One short # comment above each non-type lookup: what it matches, why it exists.

Only write down evidence: things observed in the code, the source data, issues.log, git history, or stated by the user. Never speculate, and never delete existing comments that still hold. Keep user-facing prose (description) and maintainer notes (comments) strictly apart.

Verify

bash
python -c "from pathlib import Path; from zavod.meta import load_dataset_from_path; d = load_dataset_from_path(Path('<yml path>')); print(d.name, '-', d.model.title)"

Re-read the edited fields: every factual claim in title/description traces to the source or crawler, the summary length is in range, and no sourcing or mechanics language remains in the description.

© opensanctions, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/dataset-metadata of opensanctions/opensanctions.

Open the folder on GitHubat commit 4499adf

Compare with similar skills

Dataset Metadata next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Metadata compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Metadata this skillopensanctions/opensanctions831—~604Automated safety check: NotesMIT
Exploratory Data Analysisspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: PassMIT
MatplotlibzLanqing/codex-claude-academic-skills4.6k17 repos~2.9kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Chart Visualizationbytedance/deer-flow83k2 repos~840Automated safety check: PassMIT
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Matplotlib

    zLanqing/codex-claude-academic-skills

    Low-level plotting library for full customization. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~2.9k tokens
    Data & AnalyticsAuto-check passed
  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Chart Visualization

    bytedance/deer-flow

    Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.

    83k GitHub starsUsed in 2 repos~840 tokens
    Data & AnalyticsAuto-check passed
  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Jupyter Notebook

    microsoft/ai-agents-for-beginners

    Official

    A skill your agent uses when the user asks to create, scaffold, or edit Jupyter notebooks (.ipynb) for experiments, explorations, or tutorials; prefer the bundled templates and run the helper script…

    77k GitHub starsUsed in 8 repos~1k tokens
    Data & AnalyticsAuto-check passed

More from opensanctions/opensanctions

All 11 skills in this repo
  • Crawler Pep

    opensanctions/opensanctions

    Scaffold a new PEP (Politically Exposed Persons) crawler — members of a parliament, legislature, senate, chamber of deputies, cabinet, judiciary, or an asset-declaration register — from a source URL…

    831 GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • Crawler Constants To Yml

    opensanctions/opensanctions

    Move hardcoded lookup/config constants (gender maps, header dicts, value translations, column-label maps, date formats) out of a crawler and into the dataset .yml — as datapatch lookups wherever…

    831 GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Legislature Metadata

    opensanctions/opensanctions

    Refactor the title, description and coverage frequency of a legislature/parliament PEP dataset .yml into the house style.

    831 GitHub stars~953 tokensUpdated today
    Auto-check passed
  • Name Framework Migration First Step

    opensanctions/opensanctions

    Migrate ad-hoc name cleaning in a crawler to h.reviewnames (Step 1 of the name framework migration).

    831 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Refactor Crawler

    opensanctions/opensanctions

    Rewrite messy or AI-generated crawler code into clean, production-ready style that follows the zavod best practices.

    831 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Release Datasets

    opensanctions/opensanctions

    Release one or more datasets by adding them to a topical collection, bumping coverage.start, and verifying.

    831 GitHub stars~566 tokensUpdated today
    Auto-check passed

Questions about Dataset Metadata

What does Dataset Metadata do?

Bring a dataset .yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments). Dataset Metadata is an agent skill from opensanctions/opensanctions.yml's metadata in line with house conventions (title, summary, description, coverage, publisher, maintainer comments).

When should I use Dataset Metadata?

Dataset Metadata fits situations like: the user asks to fix; standardise a datasets metadata.

How do I install Dataset Metadata in Claude Code?

Run `npx skills add opensanctions/opensanctions --skill dataset-metadata -a claude-code`. Or copy the skill folder (.claude/skills/dataset-metadata in opensanctions/opensanctions) into .claude/skills/dataset-metadata in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Metadata in Codex?

Run `npx skills add opensanctions/opensanctions --skill dataset-metadata -a codex`. Or copy the skill folder (.claude/skills/dataset-metadata in opensanctions/opensanctions) into .agents/skills/dataset-metadata in your project. Codex loads it when a task matches its description.

Can I use Dataset Metadata in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opensanctions/opensanctions --skill dataset-metadata -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-metadata, .gemini/skills/dataset-metadata, .github/skills/dataset-metadata and .opencode/skills/dataset-metadata in your project.

What does Dataset Metadata need to run?

Going by SKILL.md and its folder, Dataset Metadata needs the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Edit, Glob, Grep, Bash.

Does Dataset Metadata access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dataset Metadata safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Dataset Metadata use?

Dataset Metadata is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Metadata use?

About 604 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dataset Metadata?

Skills that share tags, products or a category with Dataset Metadata: Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Matplotlib (zLanqing/codex-claude-academic-skills, 4.6k stars), Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars) and Chart Visualization (bytedance/deer-flow, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Metadata?

opensanctions (a GitHub organization) maintains it in opensanctions/opensanctions, which has 831 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.

Source: opensanctions/opensanctions on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.