Dinobase Business Data Queries
kappa90/dinobase
Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.
A skill your agent uses when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID…
$ npx skills add owid/etl --skill generate-metadata -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl generate-metadata --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/generate-metadata .claude/skills/generate-metadata && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "generate-metadata" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/generate-metadata into .claude/skills/generate-metadata/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "generate-metadata", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/generate-metadataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill generate-metadata -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl generate-metadata --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/generate-metadata .agents/skills/generate-metadata && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "generate-metadata" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/generate-metadata into .agents/skills/generate-metadata/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "generate-metadata", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill generate-metadata -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl generate-metadata --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/generate-metadata .cursor/skills/generate-metadata && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "generate-metadata" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/generate-metadata into .cursor/skills/generate-metadata/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "generate-metadata", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/generate-metadata--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill generate-metadata -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl generate-metadata --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/generate-metadata .gemini/skills/generate-metadata && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "generate-metadata" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/generate-metadata into .gemini/skills/generate-metadata/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "generate-metadata", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl generate-metadataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill generate-metadata -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/generate-metadata .github/skills/generate-metadata && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "generate-metadata" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/generate-metadata into .github/skills/generate-metadata/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "generate-metadata", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill generate-metadata -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl generate-metadata --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/generate-metadata .opencode/skills/generate-metadata && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "generate-metadata" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/generate-metadata into .opencode/skills/generate-metadata/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "generate-metadata", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
generate-metadataA skill your agent uses when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID…
Generate Metadata is an agent skill from owid/etl. Use when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID metadata standards. Trigger when writing or editing .meta.yml files, when a garden step has empty or minimal metadata, or when user asks to improve/add/enrich metadata.
Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL and Data analysis. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit bf5dc8e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Generate Metadata loads about 6.8k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 3,097 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from owid/etl at commit bf5dc8e, republished under its MIT licence (© owid). 3,097 words, ~6,782 tokens.
.claude/skills/generate-metadata/SKILL.md (or your agent's skills folder).Practical guidelines for writing high-quality metadata in *.meta.yml files.
Core principle: Metadata is for the public. Every field should help someone understand the data. Write in plain language -- if a layperson can't understand it, rewrite it.
Don't use for: Snapshot metadata (different process) or quick single-field edits.
definitions:
common:
processing_level: minor # or major
presentation:
topic_tags:
- <Topic>
attribution_short: <Short Source Name>
display:
numDecimalPlaces: 1
# Reusable text blocks
my_note: &my_note |-
Reusable text.
dataset:
update_period_days: 365 # Must be accurate: 0 for datasets that will never update
tables:
<table_name>:
variables:
<variable_name>:
title: <Human-readable title>
unit: <Full unit name>
short_unit: <Symbol>
description_short: |-
<1-2 sentences>
description_key: |-
<Key information as free-form markdown: paragraphs, with sub-lists where helpful>
description_from_producer: |-
<Original text from source>
description_processing: |-
<What OWID did to process this>
display:
name: <Legend label>
presentation:
title_public: <Public-facing title>
title_variant: <Disambiguating label>Key rule from CLAUDE.md: The dataset: block carries only update_period_days and owners -- everything else is inherited from origin. Always set owners when it's missing (canonical name from the schemas/dataset-schema.json enum, resolved via etl.owners.resolve_owner; first entry = accountable owner).
Full field reference: For a complete list of all supported metadata fields (beyond what's covered here), read schemas/dataset-schema.json.
| Field | Purpose | Length |
|---|---|---|
title | Primary identifier, always required | ~100 chars max |
presentation.title_public | Human-readable public title | Must be excellent |
presentation.title_variant | Disambiguator ("Historical data", "WHO") | Short phrase |
display.name | Chart legend label | ~30 chars max |
Rules:
display.name is set, also set title_public (see title hierarchy in docs/architecture/metadata/faqs.md)title_public unless the title has dimension breakdowns or codes (e.g. SDG indicator numbers). The data page will show the curated chart title instead, which is usually better. Use display.name for cleaner legend/table labels.title_variant disambiguates when multiple indicators share a similar title — use short phrases like "Historical data", "WHO", "Age-standardized", "Extrapolated". Watch for redundancy with attribution_short (avoid "V-Dem - V-Dem" duplication).# GOOD
title: Number of neutron star mergers in the Milky Way
display:
name: Neutron star mergers
# BAD
title: Number of neutron star mergers (NASA, 2023)| Field | Format | Examples |
|---|---|---|
unit | Lowercase, plural, "per" not "/" | tonnes per hectare, %, "" |
short_unit | SI abbreviation | t/ha, %, "" |
unit explicitly, even to "" for dimensionless indicators (scores, indexes)short_unit is only needed when there's an actual unit to abbreviate. Omit it for dimensionless indicators -- it defaults to None and grapher won't show a unit label.kilowatts per person)short_unit should use SI abbreviations (g not grams, % not pct)Decimal precision: 0 for counts, 1 for percentages, 2 for economic/per-capita values. Be consistent across related variables. Always set numDecimalPlaces explicitly -- it's a frequent source of review feedback.
description_short -- 1-2 sentences, ~200 characters ideally. Answers "What does this number measure?" Longer explanations belong in description_key.
|- block scalar for multi-line; inline strings are fine for single sentences[text](url) for links, [term](#dod:term) for OWID definition popups (e.g. [stunted](#dod:stunting))# GOOD
description_short: |-
The number of people living in extreme poverty, defined as living on less than $2.15 per day.
# BAD - repeats the title
description_short: |-
Manufactured cigarettes sold in this country in this year.
# BAD - mentions sources (belongs in description_key)
description_short: |-
The number of people living in extreme poverty, based on data and estimates from different sources.description_key -- Free-form markdown text for the "About this data" panel: prose paragraphs, with markdown sub-lists only where a list genuinely helps. (A YAML list of bullet points is still accepted and renders as a markdown list, but prefer prose — see grapher's descriptionKey-to-string migration.)
description_short and the chart's title/subtitle, so a sentence that repeats another at the same level of detail costs the reader attention and makes the panel look padded. Before adding one, read the whole rendered list and ask what it adds that isn't already there; if the answer is "it says an existing bullet more fully", edit that bullet rather than adding a second. Expanding description_short is not redundancy — the short line is a one-sentence summary, and unpacking it in the first bullet (the full definition, how it's measured, what's included) is exactly what the panel is for. What to avoid is a bullet that restates it and stops there. Two shapes to watch: a source or caveat sentence duplicating a caveat already made in different words, and a Jinja variant bullet restating the shared bullet for some dimension values only.# GOOD - prose paragraphs, sub-list only where it helps
description_key: |-
Extreme poverty is measured using the International Poverty Line of $2.15 per day in 2017 international dollars.
This metric uses household survey data adjusted for purchasing power parity (PPP).
# BAD - unexpanded acronyms, jargon
description_key: |-
Uses IPL of $2.15/day (2017 PPP).description_from_producer -- Exact producer text, verbatim or minimally edited. Only if producer provides clear definitions. Can be set in definitions.common when the same producer description applies to all variables, with per-variable overrides as needed.
description_processing -- What OWID did. Only for major transformations (aggregations, per-capita calculations, combining sources). Don't document routine operations (country harmonization, dropping nulls).
description_short or description_key too.minor: data largely unchanged (reformatting, unit conversion, harmonization) -> use most restrictive origin licensemajor: significant transformations (calculations, combinations, imputations) -> use CC BY 4.0major.Set in definitions.common, override per-variable as needed.
.venv/bin/python -c "import json; tags=json.load(open('schemas/dataset-schema.json'))['properties']['tables']['additionalProperties']['properties']['variables']['additionalProperties']['properties']['presentation']['properties']['topic_tags']['items']['enum']; print('\n'.join(tags))"definitions.common.presentation.topic_tags.presentation.attribution_short: Short source name ("WHO", "World Bank"). Set in common when uniform.presentation.grapher_config: Only set when you want a specific default chart view. Common sub-fields:note: Chart footnotes — methodology caveats, sample sizes, inflation adjustments. Keep to 1-2 sentences.selectedEntityNames: Pre-select countries/regions for the default view (e.g. ["United States", "China", "Europe"])selectedEntityColors: Map entity names to hex colors (e.g. {"Africa": "#A2559C", "Asia": "#00847E"})map: Map tab settings — colorScale with baseColorScheme (e.g. "YlOrRd"), binningStrategy ("manual"), and customNumericValues for bin thresholdsdefinitions.common.presentation when all variables share the same chart defaultsdisplay.numDecimalPlaces: Set explicitly. Use metadata-export --decimals auto to auto-detect.display.tolerance: Number of years to allow gap-bridging on line charts (default 0). Set higher (e.g. 5-10) for sparse historical data where connecting distant points is acceptable.display.roundingMode: Use "significantFigures" with numSignificantFigures instead of numDecimalPlaces when values span many orders of magnitude.Use definitions.common when 3+ variables share the same field values. Remember: common does NOT merge -- it completely overrides. Use <<: *anchor for partial overrides.
Use anchors/aliases for identical blocks shared by 2+ variables. Define in definitions: at the top. Name anchors to indicate their target field (e.g. description_producer_refugee not description_refugee) so reviewers can tell which metadata field the text will end up in.
Use <<: *anchor merge to extend a shared mapping while overriding specific keys (e.g. <<: *common_display then numDecimalPlaces: 0).
Adding text to an existing file: follow the pattern it already uses. If its text lives in definitions: and variables reference {definitions.<key>}, add your text as new definitions at the top next to the related ones — inline prose under the variable renders fine but leaves the file with two authoring styles. In big datasets prefer definitions-at-top regardless of reuse: with hundreds of variables it is what keeps the file readable, since all the prose sits in one place and the variable blocks stay skimmable. Grep the existing definitions first: new text often duplicates a bullet already defined under another name, and reusing (or replacing that key's text, after checking which other variables reference it) beats a near-duplicate. Then widen the grep past the file — boilerplate travels, so the same caveat usually sits in other datasets' .meta.yml too (search distinctive 5–8 word fragments across etl/steps/, since near-duplicates differ by a word or two). When your wording supersedes theirs, propose the same fix there in a separate PR rather than editing another dataset's text inside yours; when a sibling already words it better, adopt that wording instead of minting a third variant. Details and the text-neutrality proof for such refactors are in .claude/skills/edit-faust-metadata/SKILL.md ("Writing new text into a garden .meta.yml").
Use Jinja templates for dimensional datasets (age, sex, cause breakdowns). Custom delimiters: <% %> for blocks, << >> for expressions.
# Jinja example: dimensional variable with conditional descriptions
definitions:
description_short_time_spent: |-
<% if who_category == "Alone" %>
Time spent alone, by gender and age.
<%- else %>
Time spent with <<who_category.lower()>>, by gender and age.
<%- endif %>
variables:
time_spent:
title: Time spent with <<who_category>> throughout life
unit: hours per day
short_unit: h
description_short: "{definitions.description_short_time_spent}"
display:
name: With <<who_category>>Write Jinja conditionals in definitions:, with the condition first and each branch on its own line, and call them by name from the variable. This applies whenever the conditional picks a whole value: a full field, or a description_key bullet. Keep the <% if %> out of the variable block itself, so the variable stays a short list of {definitions.…} references. Put each <% if %> / <%- elif %> / <%- else %> / <%- endif %> tag on its own line, with what it renders underneath, so a reviewer reads condition → text without scanning one long line:
definitions:
description_key_gni_per_capita: GNI per capita is GNI divided by population, …
description_key_gni_per_capita_by_sex: GNI is not measured separately by sex, …
description_key_gni_per_capita_total_or_by_sex: |-
<% if sex == 'total' %>
{definitions.description_key_gni_per_capita}
<%- else %>
{definitions.description_key_gni_per_capita_by_sex}
<%- endif %>
tables:
undp_hdr_sex:
variables:
gni_pc:
description_key:
- "{definitions.description_key_gni}"
- "{definitions.description_key_gni_per_capita_total_or_by_sex}"❌ - "<% if sex == 'total' %>{definitions.a}<% else %>{definitions.b}<% endif %>", a one-line conditional inside the variable.
Name the conditional definition after what it chooses between (…_total_or_by_sex), not after one of the branches. Splitting the tags over several lines renders exactly like the one-line form. The Jinja environment sets trim_blocks and lstrip_blocks, and the rendered text is stripped (owid.catalog.core.jinja), so the tag lines and their newlines disappear. {definitions.…} references resolve before Jinja runs, so a definition can wrap other definitions. Each branch must still be a single line, because a wrapped branch puts a real newline into the text.
Dash the closing tags: <%- elif %>, <%- else %>, <%- endif %>. trim_blocks removes the newline after a tag but not the one that ends the branch line before it. As a whole value that newline is stripped anyway. But once the definition is spliced into a longer string (Intro {definitions.x} outro.), the plain form renders Intro A\n outro., and the - is what gives Intro A outro.. Since a definition can be reused anywhere, dash them by default. Leave the opening <% if %> undashed, because <%- there would also eat the space before it and glue the text together. If text follows the conditional, keep it on the <%- endif %> line (<%- endif %> Outro.): on the next line it gets glued on as AOutro..
The exception is a fragment inside a sentence or a title, such as title: GNI per capita<% if sex != "total" %> (<<sex>>)<% endif %>. Splitting that one leaves a line break or stray indentation in the rendered text, so keep it inline. Short fragments like this can still live in a definition (among_sex: <% if sex == "males" %> among men<% elif … %><% endif %>) when several fields reuse them. When you restructure an existing conditional, render every dimension value before and after to confirm nothing changed (the text-neutrality recipe in .claude/skills/edit-faust-metadata/SKILL.md).
A sentence written with one breakdown in mind renders on all the others, where the view's own filtering can make it false. Sweep every dimension value before shipping — check 6 of the quality suite below.
Use {definitions.xxx} string interpolation for reusing text fragments inline (e.g. '{definitions.methodology}' in a description_key bullet). Unlike YAML anchors which substitute entire nodes, this inserts text within strings. Use anchors for whole fields/blocks, interpolation for composing text.
Use shared.meta.yml when multiple .meta.yml files in the same directory share definitions or macros. These files contain only Jinja macros and reusable definitions — no actual variable metadata. Step-level .meta.yml files then import and call these macros. Used in large multi-file datasets like IHME GBD.
For full syntax details, see docs/architecture/metadata/structuring-yaml.md.
processing_level: major, document the calculation in description_processing, use international-$ per person not "per capita"presentation.title_variant: Age-standardized, explain standardization method in description_keyunit: "%", absolute gets the count unitvery_worried, not_worried_at_all), put all shared metadata in definitions.common (question text, methodology, unit, display) and give individual variables only a title. This avoids repetition and keeps the file compact..venv/bin/etl metadata-export data/garden/<ns>/<ver>/<ds> --output /tmp/<ds>.meta.yml (never run without --show or --output to avoid overwriting).dvc for origin URLs, visit source docs for methodology and definitionsdefinitions: and common: to identify shared patterns before individual variablesINSTANT=1 .venv/bin/etlr data://grapher/<ns>/<ver>/<ds> --grapher --onlyThe canonical check suite for metadata text, shared by /update-dataset (§6b/§6c), /edit-faust-metadata, and this skill. Keep the three in sync: if a check is added or changed here, check whether those skills need a matching edit in the same commit.
Run all of these after the metadata is written and the steps are built, so every issue surfaces together:
/check-metadata-typos on each edited .meta.yml (garden first, then grapher). Accept or skip each suggested fix./check-metadata-style on the grapher step. A mechanical pass first catches template artifacts (doubled spaces, stray newlines, leading or trailing whitespace) that only appear after Jinja rendering; then it audits user-facing fields (title, subtitle, description_short, display.name, presentation.*) against OWID's Writing and Style Guide (.claude/skills/check-metadata-style/STYLE_GUIDE.md).<agency>…") — open the cited link and confirm it actually says that; agencies revise methodologydescription_short, or the title at the same level of detail. Expanding the short line is fine and expected; restating it is padding. See the description_key guidance above.[term](#dod:term) in the text: URLs must resolve (curl as the batch primary; on a 4xx from an OWID link, double-check with WebFetch + Wayback before acting); dod slugs checked against the dods table via public Datasette (SELECT name FROM dods WHERE name LIKE ...); a missing dod → keep the link and list it as a "create in admin" follow-up in the PR body.definitions: key several variants reference), every sentence must hold at every value it renders on, not just the one it was written for. Render the text per dimension value and read each output as a reader of that chart, asking what the view already restricts: a caveat that the data doesn't control for X is wrong on the variant grouped by X; a scope word like "all employees" overclaims on a variant filtered to a subgroup; a sentence about a toggle is wrong on views that exist for only one choice of that dimension. Prefer qualifying the wording so it holds everywhere (often one word, nothing extra to maintain) over adding a Jinja branch or a view-level override. Automated reviewers catch this class reliably, so sweeping first saves a review round./fact-check-dataset scoped to the newly written or edited metadata text only: treat each added/changed sentence as a claim and verify it against the producer's documentation (read what's behind the links — check 4 only proves they resolve). Catches text that is well-formed but factually wrong: stale methodology attributions, scope overclaims, misread units in prose. Scope by context: mandatory in /edit-faust-metadata (claims-only, no data cross-checks — cheap); the full-dataset review including data-value cross-checks stays the opt-in step described in /update-dataset §6c-bis (token-heavy).If any check rewrites a .meta.yml, re-run the affected step so the built catalog reflects the edits (add --grapher when the step is on the grapher channel, otherwise staging keeps serving the old text), then re-run the check to confirm zero remaining violations.
The checks above read the text as authored. The Metadata Diff Wizard page reads it as rendered — indicator metadata merged with any view-level override, the way the site resolves it — and lists every chart, MDim view and explorer view each edit lands on:
http://staging-site-<container_branch>/etl/wizard/metadata-diffIt is not an eighth check: it needs a staging server and a human's judgment, and it runs after the checks, once the text is settled. What it shows that nothing above does is reach — reword one shared description_key and dozens of charts change while every chart-config diff stays empty, because that text is inherited, not configured. (Chart Diff compares configs; this compares rendered texts. It is also not Chart Diff's own per-chart "Metadata differences" modal.)
Two preconditions, both silent when unmet:
staging-site-master otherwise: the same baseline Chart Diff uses.STAGING=1 .venv/bin/etlr grapher://grapher/<path> --grapher). Its scope is the branch's changed step files intersected with the datasets actually rebuilt there, so text that never reached the server simply doesn't appear. A server that is behind gets a 🚧 banner naming each stale dataset and the command to rebuild it — read that banner before trusting any count, because a stale dataset reports its diffs backwards, showing its older text as this branch's change.Reviewing marks each chart / MDim view / explorer view ✅ or ❌ with a note; the Review section then exports metadata-rejections.md, the rejections written as instructions naming the edit, the garden .meta.yml it was authored in, and the dataset owner. Nothing here gates the merge — rejecting changes no text — so an unactioned rejection is an open item somebody has to carry (.claude/docs/open-items.md). Verdicts are bound to the wording they were made on: reword the text and the decision reopens rather than counting as done.
Items that are easy to miss (obvious rules like "set title" are omitted — see field guidelines above):
description_short adds value beyond the title — if it just repeats the title, delete itdescription_key includes concrete examples where scope is abstractdescription_key only describes data that actually exists in the indicatordescription_processing matches current code and doesn't reference internal dataset namesdescription_short or description_key, not buried in description_processingnumDecimalPlaces set and consistent across related variablesdisplay.name paired with title_public when set{definitions.xxx}; Jinja for 10+ similar variables’ “ ”), not straight (' ") — straight marks are an LLM tell; see check-metadata-style/STYLE_GUIDE.md© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/generate-metadata of owid/etl.
Open the folder on GitHubat commit bf5dc8e
Generate Metadata next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Generate Metadata this skillowid/etl | 158 | — | ~6.8k | Automated safety check: Pass | MIT | |
| Dinobase Business Data Querieskappa90/dinobase | 263 | — | ~1.5k | Automated safety check: Pass | Custom licence | |
| Analytics Engineerborghei/Claude-Skills | 881 | — | ~3.4k | Automated safety check: Pass | MIT | |
| Data Engineering Data Driven Featureaiskillstore/marketplace | 430 | 7 repos | ~3k | Automated safety check: Pass | None | |
| Profiling Tablesastronomer/agents | 451 | — | ~964 | Automated safety check: Pass | Apache-2.0 | |
| Data Researchermajiayu000/claude-skill-registry | 666 | 1 repos | ~4.6k | Automated safety check: Pass | MIT |
kappa90/dinobase
Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.
borghei/Claude-Skills
Analytics engineering across data modeling, dbt, transformation, and semantic layers.
aiskillstore/marketplace
Build features guided by data insights, A/B testing, and continuous measurement using specialized agents for analysis, implementation, and experimentation.
astronomer/agents
Deep-dive data profiling for a specific table. An agent skill from astronomer/agents.
majiayu000/claude-skill-registry
Data discovery and analysis specialist focused on extracting actionable insights from complex datasets, identifying patterns and anomalies, and transforming raw data into strategic intelligence.
astronomer/agents
Queries the data warehouse with SQL and answers business questions about data.
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Categories
A skill your agent uses when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID…. Generate Metadata is an agent skill from owid/etl. Use when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID metadata standards.
Generate Metadata fits situations like: enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection; data exploration; web research following OWID metadata standards; editing .meta.yml files.
Run `npx skills add owid/etl --skill generate-metadata -a claude-code`. Or copy the skill folder (.claude/skills/generate-metadata in owid/etl) into .claude/skills/generate-metadata in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill generate-metadata -a codex`. Or copy the skill folder (.claude/skills/generate-metadata in owid/etl) into .agents/skills/generate-metadata in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill generate-metadata -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generate-metadata, .gemini/skills/generate-metadata, .github/skills/generate-metadata and .opencode/skills/generate-metadata in your project.
Going by SKILL.md and its folder, Generate Metadata needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Generate Metadata is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Generate Metadata: Dinobase Business Data Queries (kappa90/dinobase, 263 stars), Analytics Engineer (borghei/Claude-Skills, 881 stars), Data Engineering Data Driven Feature (aiskillstore/marketplace, 430 stars) and Profiling Tables (astronomer/agents, 451 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.