Agent skill

Generate Metadata

by owid in owid/etl

A skill your agent uses when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID…

MITAuto-check passedData & Analytics

Install Generate Metadata

skills CLI
$ npx skills add owid/etl --skill generate-metadata -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install owid/etl generate-metadata --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/generate-metadata .claude/skills/generate-metadata && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generate-metadata
GitHub stars
158
Token cost
~6.8k tokens
SKILL.md length
3,097 words
Files
1
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID…

  • Works in 7 steps: Generate skeleton: .venv/bin/etl… → Inspect the data: Load dataset, check… → Research sources: Read snapshot .dvc for… → …
  • Enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection
  • SKILL.md covers When to Use, Canonical YAML Structure, Field-by-Field Guidelines and YAML Efficiency Patterns, plus 4 more sections
  • Calls python

What it does

Generate Metadata is an agent skill from owid/etl. Use when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID metadata standards. Trigger when writing or editing .meta.yml files, when a garden step has empty or minimal metadata, or when user asks to improve/add/enrich metadata.

Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL and Data analysis. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.

When your agent uses it

  • Enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection
  • Data exploration
  • Web research following OWID metadata standards
  • Editing .meta.yml files

Example prompts

  • “/generate-metadata”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Generate skeleton: .venv/bin/etl metadata-export data/garden/// --output /tmp/.meta.yml (never run without --show or --output to avoid…
  2. Inspect the data: Load dataset, check columns, value ranges, existing origin metadata
  3. Research sources: Read snapshot .dvc for origin URLs, visit source docs for methodology and definitions
  4. Write top-down: Start with definitions: and common: to identify shared patterns before individual variables
  5. Fill systematically: For each variable: title -> units -> description_short -> description_key -> processing_level -> display
  6. Preview with INSTANT mode: INSTANT=1 .venv/bin/etlr data://grapher/// --grapher --only
  7. Run the metadata quality checks (next section)

What it can do on your machine

Read from SKILL.md and the folder at commit bf5dc8e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generate Metadata loads about 6.8k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 3,097 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from owid/etl at commit bf5dc8e, republished under its MIT licence (© owid). 3,097 words, ~6,782 tokens.

Download SKILL.mdSave it as .claude/skills/generate-metadata/SKILL.md (or your agent's skills folder).
name
generate-metadata
description
Use when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID metadata standards. Trigger when writing or editing *.meta.yml files, when a garden step has empty or minimal metadata, or when user asks to improve/add/enrich metadata.
metadata.internal
true
metadata.owner
Marigold

OWID Metadata Generation

Practical guidelines for writing high-quality metadata in *.meta.yml files.

Core principle: Metadata is for the public. Every field should help someone understand the data. Write in plain language -- if a layperson can't understand it, rewrite it.

When to Use

  • New garden/grapher datasets needing metadata
  • Enriching incomplete or sparse metadata files
  • Updating existing metadata for dataset refreshes

Don't use for: Snapshot metadata (different process) or quick single-field edits.

Canonical YAML Structure

yaml
definitions:
  common:
    processing_level: minor  # or major
    presentation:
      topic_tags:
        - <Topic>
      attribution_short: <Short Source Name>
    display:
      numDecimalPlaces: 1
  # Reusable text blocks
  my_note: &my_note |-
    Reusable text.

dataset:
  update_period_days: 365  # Must be accurate: 0 for datasets that will never update

tables:
  <table_name>:
    variables:
      <variable_name>:
        title: <Human-readable title>
        unit: <Full unit name>
        short_unit: <Symbol>
        description_short: |-
          <1-2 sentences>
        description_key: |-
          <Key information as free-form markdown: paragraphs, with sub-lists where helpful>
        description_from_producer: |-
          <Original text from source>
        description_processing: |-
          <What OWID did to process this>
        display:
          name: <Legend label>
        presentation:
          title_public: <Public-facing title>
          title_variant: <Disambiguating label>

Key rule from CLAUDE.md: The dataset: block carries only update_period_days and owners -- everything else is inherited from origin. Always set owners when it's missing (canonical name from the schemas/dataset-schema.json enum, resolved via etl.owners.resolve_owner; first entry = accountable owner).

Full field reference: For a complete list of all supported metadata fields (beyond what's covered here), read schemas/dataset-schema.json.

Field-by-Field Guidelines

Titles
FieldPurposeLength
titlePrimary identifier, always required~100 chars max
presentation.title_publicHuman-readable public titleMust be excellent
presentation.title_variantDisambiguator ("Historical data", "WHO")Short phrase
display.nameChart legend label~30 chars max

Rules:

  • Sentence-case, no trailing period
  • Never include producer, year, or units in titles
  • When display.name is set, also set title_public (see title hierarchy in docs/architecture/metadata/faqs.md)
  • Prefer NOT setting title_public unless the title has dimension breakdowns or codes (e.g. SDG indicator numbers). The data page will show the curated chart title instead, which is usually better. Use display.name for cleaner legend/table labels.
  • Prefer clear, reader-friendly titles over technically precise ones (e.g. "Share of people who think vaccines are safe" over "Share who strongly agree that...")
  • title_variant disambiguates when multiple indicators share a similar title — use short phrases like "Historical data", "WHO", "Age-standardized", "Extrapolated". Watch for redundancy with attribution_short (avoid "V-Dem - V-Dem" duplication).
yaml
# GOOD
title: Number of neutron star mergers in the Milky Way
display:
  name: Neutron star mergers
# BAD
title: Number of neutron star mergers (NASA, 2023)
Units
FieldFormatExamples
unitLowercase, plural, "per" not "/"tonnes per hectare, %, ""
short_unitSI abbreviationt/ha, %, ""
  • Always set unit explicitly, even to "" for dimensionless indicators (scores, indexes)
  • short_unit is only needed when there's an actual unit to abbreviate. Omit it for dimensionless indicators -- it defaults to None and grapher won't show a unit label.
  • Use "person" not "capita" (kilowatts per person)
  • Choose human-friendly scales -- if most values are below 1 tonne per person, use "kilograms per person" instead
  • short_unit should use SI abbreviations (g not grams, % not pct)

Decimal precision: 0 for counts, 1 for percentages, 2 for economic/per-capita values. Be consistent across related variables. Always set numDecimalPlaces explicitly -- it's a frequent source of review feedback.

Descriptions

description_short -- 1-2 sentences, ~200 characters ideally. Answers "What does this number measure?" Longer explanations belong in description_key.

  • Use |- block scalar for multi-line; inline strings are fine for single sentences
  • Don't mention units, sources, or processing (redundant or belongs elsewhere)
  • Don't just repeat the title -- if it adds nothing beyond the title, omit it entirely
  • Remove filler phrases like "Emissions are..." at the start
  • Supports Markdown: [text](url) for links, [term](#dod:term) for OWID definition popups (e.g. [stunted](#dod:stunting))
yaml
# GOOD
description_short: |-
  The number of people living in extreme poverty, defined as living on less than $2.15 per day.
# BAD - repeats the title
description_short: |-
  Manufactured cigarettes sold in this country in this year.
# BAD - mentions sources (belongs in description_key)
description_short: |-
  The number of people living in extreme poverty, based on data and estimates from different sources.

description_key -- Free-form markdown text for the "About this data" panel: prose paragraphs, with markdown sub-lists only where a list genuinely helps. (A YAML list of bullet points is still accepted and renders as a markdown list, but prefer prose — see grapher's descriptionKey-to-string migration.)

  • Plain language, no jargon (e.g. "livestock digestive processes" not "enteric fermentation"). Expand acronyms on first use.
  • Add concrete examples for abstract scope descriptions (what countries are included, what events qualify)
  • Order: data-specific points first, methodology second, caveats last
  • Only describe data that actually exists in the indicator. State implications of limitations explicitly.
  • Separate distinct facts with blank lines (paragraphs), so line breaks render meaningfully.
  • No redundancy. The bullets are read as a set, directly under description_short and the chart's title/subtitle, so a sentence that repeats another at the same level of detail costs the reader attention and makes the panel look padded. Before adding one, read the whole rendered list and ask what it adds that isn't already there; if the answer is "it says an existing bullet more fully", edit that bullet rather than adding a second. Expanding description_short is not redundancy — the short line is a one-sentence summary, and unpacking it in the first bullet (the full definition, how it's measured, what's included) is exactly what the panel is for. What to avoid is a bullet that restates it and stops there. Two shapes to watch: a source or caveat sentence duplicating a caveat already made in different words, and a Jinja variant bullet restating the shared bullet for some dimension values only.
yaml
# GOOD - prose paragraphs, sub-list only where it helps
description_key: |-
  Extreme poverty is measured using the International Poverty Line of $2.15 per day in 2017 international dollars.

  This metric uses household survey data adjusted for purchasing power parity (PPP).
# BAD - unexpanded acronyms, jargon
description_key: |-
  Uses IPL of $2.15/day (2017 PPP).

description_from_producer -- Exact producer text, verbatim or minimally edited. Only if producer provides clear definitions. Can be set in definitions.common when the same producer description applies to all variables, with per-variable overrides as needed.

description_processing -- What OWID did. Only for major transformations (aggregations, per-capita calculations, combining sources). Don't document routine operations (country harmonization, dropping nulls).

  • Document dropped data and aggregation choices. Keep in sync with code.
  • Don't reference internal dataset names — describe what was done, not which datasets were used.
  • Low-visibility field — surface important caveats in description_short or description_key too.
Processing Level and License
  • minor: data largely unchanged (reformatting, unit conversion, harmonization) -> use most restrictive origin license
  • major: significant transformations (calculations, combinations, imputations) -> use CC BY 4.0
  • The test: is the data identical to the original? If not (e.g. processing of historical regions, per-capita calculations), it's major.

Set in definitions.common, override per-variable as needed.

Topic Tags
  • 1-3 tags, most relevant first. First tag = primary (used in citations).
  • Must match valid tags from the schema. To list them:
    bash
    .venv/bin/python -c "import json; tags=json.load(open('schemas/dataset-schema.json'))['properties']['tables']['additionalProperties']['properties']['variables']['additionalProperties']['properties']['presentation']['properties']['topic_tags']['items']['enum']; print('\n'.join(tags))"
  • Set in definitions.common.presentation.topic_tags.
Presentation & Display
  • presentation.attribution_short: Short source name ("WHO", "World Bank"). Set in common when uniform.
  • presentation.grapher_config: Only set when you want a specific default chart view. Common sub-fields:
    • note: Chart footnotes — methodology caveats, sample sizes, inflation adjustments. Keep to 1-2 sentences.
    • selectedEntityNames: Pre-select countries/regions for the default view (e.g. ["United States", "China", "Europe"])
    • selectedEntityColors: Map entity names to hex colors (e.g. {"Africa": "#A2559C", "Asia": "#00847E"})
    • map: Map tab settings — colorScale with baseColorScheme (e.g. "YlOrRd"), binningStrategy ("manual"), and customNumericValues for bin thresholds
    • Set at variable level, or in definitions.common.presentation when all variables share the same chart defaults
  • display.numDecimalPlaces: Set explicitly. Use metadata-export --decimals auto to auto-detect.
  • display.tolerance: Number of years to allow gap-bridging on line charts (default 0). Set higher (e.g. 5-10) for sparse historical data where connecting distant points is acceptable.
  • display.roundingMode: Use "significantFigures" with numSignificantFigures instead of numDecimalPlaces when values span many orders of magnitude.

YAML Efficiency Patterns

Use definitions.common when 3+ variables share the same field values. Remember: common does NOT merge -- it completely overrides. Use <<: *anchor for partial overrides.

Use anchors/aliases for identical blocks shared by 2+ variables. Define in definitions: at the top. Name anchors to indicate their target field (e.g. description_producer_refugee not description_refugee) so reviewers can tell which metadata field the text will end up in.

Use <<: *anchor merge to extend a shared mapping while overriding specific keys (e.g. <<: *common_display then numDecimalPlaces: 0).

Adding text to an existing file: follow the pattern it already uses. If its text lives in definitions: and variables reference {definitions.<key>}, add your text as new definitions at the top next to the related ones — inline prose under the variable renders fine but leaves the file with two authoring styles. In big datasets prefer definitions-at-top regardless of reuse: with hundreds of variables it is what keeps the file readable, since all the prose sits in one place and the variable blocks stay skimmable. Grep the existing definitions first: new text often duplicates a bullet already defined under another name, and reusing (or replacing that key's text, after checking which other variables reference it) beats a near-duplicate. Then widen the grep past the file — boilerplate travels, so the same caveat usually sits in other datasets' .meta.yml too (search distinctive 5–8 word fragments across etl/steps/, since near-duplicates differ by a word or two). When your wording supersedes theirs, propose the same fix there in a separate PR rather than editing another dataset's text inside yours; when a sibling already words it better, adopt that wording instead of minting a third variant. Details and the text-neutrality proof for such refactors are in .claude/skills/edit-faust-metadata/SKILL.md ("Writing new text into a garden .meta.yml").

Use Jinja templates for dimensional datasets (age, sex, cause breakdowns). Custom delimiters: <% %> for blocks, << >> for expressions.

yaml
# Jinja example: dimensional variable with conditional descriptions
definitions:
  description_short_time_spent: |-
    <% if who_category == "Alone" %>
    Time spent alone, by gender and age.
    <%- else %>
    Time spent with <<who_category.lower()>>, by gender and age.
    <%- endif %>

variables:
  time_spent:
    title: Time spent with <<who_category>> throughout life
    unit: hours per day
    short_unit: h
    description_short: "{definitions.description_short_time_spent}"
    display:
      name: With <<who_category>>

Write Jinja conditionals in definitions:, with the condition first and each branch on its own line, and call them by name from the variable. This applies whenever the conditional picks a whole value: a full field, or a description_key bullet. Keep the <% if %> out of the variable block itself, so the variable stays a short list of {definitions.…} references. Put each <% if %> / <%- elif %> / <%- else %> / <%- endif %> tag on its own line, with what it renders underneath, so a reviewer reads condition → text without scanning one long line:

yaml
definitions:
  description_key_gni_per_capita: GNI per capita is GNI divided by population, …
  description_key_gni_per_capita_by_sex: GNI is not measured separately by sex, …
  description_key_gni_per_capita_total_or_by_sex: |-
    <% if sex == 'total' %>
    {definitions.description_key_gni_per_capita}
    <%- else %>
    {definitions.description_key_gni_per_capita_by_sex}
    <%- endif %>

tables:
  undp_hdr_sex:
    variables:
      gni_pc:
        description_key:
          - "{definitions.description_key_gni}"
          - "{definitions.description_key_gni_per_capita_total_or_by_sex}"

❌ - "<% if sex == 'total' %>{definitions.a}<% else %>{definitions.b}<% endif %>", a one-line conditional inside the variable.

Name the conditional definition after what it chooses between (…_total_or_by_sex), not after one of the branches. Splitting the tags over several lines renders exactly like the one-line form. The Jinja environment sets trim_blocks and lstrip_blocks, and the rendered text is stripped (owid.catalog.core.jinja), so the tag lines and their newlines disappear. {definitions.…} references resolve before Jinja runs, so a definition can wrap other definitions. Each branch must still be a single line, because a wrapped branch puts a real newline into the text.

Dash the closing tags: <%- elif %>, <%- else %>, <%- endif %>. trim_blocks removes the newline after a tag but not the one that ends the branch line before it. As a whole value that newline is stripped anyway. But once the definition is spliced into a longer string (Intro {definitions.x} outro.), the plain form renders Intro A\n outro., and the - is what gives Intro A outro.. Since a definition can be reused anywhere, dash them by default. Leave the opening <% if %> undashed, because <%- there would also eat the space before it and glue the text together. If text follows the conditional, keep it on the <%- endif %> line (<%- endif %> Outro.): on the next line it gets glued on as AOutro..

The exception is a fragment inside a sentence or a title, such as title: GNI per capita<% if sex != "total" %> (<<sex>>)<% endif %>. Splitting that one leaves a line break or stray indentation in the rendered text, so keep it inline. Short fragments like this can still live in a definition (among_sex: <% if sex == "males" %> among men<% elif … %><% endif %>) when several fields reuse them. When you restructure an existing conditional, render every dimension value before and after to confirm nothing changed (the text-neutrality recipe in .claude/skills/edit-faust-metadata/SKILL.md).

A sentence written with one breakdown in mind renders on all the others, where the view's own filtering can make it false. Sweep every dimension value before shipping — check 6 of the quality suite below.

Use {definitions.xxx} string interpolation for reusing text fragments inline (e.g. '{definitions.methodology}' in a description_key bullet). Unlike YAML anchors which substitute entire nodes, this inserts text within strings. Use anchors for whole fields/blocks, interpolation for composing text.

Use shared.meta.yml when multiple .meta.yml files in the same directory share definitions or macros. These files contain only Jinja macros and reusable definitions — no actual variable metadata. Step-level .meta.yml files then import and call these macros. Used in large multi-file datasets like IHME GBD.

For full syntax details, see docs/architecture/metadata/structuring-yaml.md.

Show full SKILL.md (1,251 more words)Show less

Common Variable Patterns

  • Per-capita variables: processing_level: major, document the calculation in description_processing, use international-$ per person not "per capita"
  • Age-standardized variables: Use presentation.title_variant: Age-standardized, explain standardization method in description_key
  • Absolute + share pairs: Keep consistent naming; the share variable gets unit: "%", absolute gets the count unit
  • Survey response breakdowns: When variables represent response options (e.g. very_worried, not_worried_at_all), put all shared metadata in definitions.common (question text, methodology, unit, display) and give individual variables only a title. This avoids repetition and keeps the file compact.

Workflow

  1. Generate skeleton: .venv/bin/etl metadata-export data/garden/<ns>/<ver>/<ds> --output /tmp/<ds>.meta.yml (never run without --show or --output to avoid overwriting)
  2. Inspect the data: Load dataset, check columns, value ranges, existing origin metadata
  3. Research sources: Read snapshot .dvc for origin URLs, visit source docs for methodology and definitions
  4. Write top-down: Start with definitions: and common: to identify shared patterns before individual variables
  5. Fill systematically: For each variable: title -> units -> description_short -> description_key -> processing_level -> display
  6. Preview with INSTANT mode: INSTANT=1 .venv/bin/etlr data://grapher/<ns>/<ver>/<ds> --grapher --only
  7. Run the metadata quality checks (next section)

Metadata quality checks

The canonical check suite for metadata text, shared by /update-dataset (§6b/§6c), /edit-faust-metadata, and this skill. Keep the three in sync: if a check is added or changed here, check whether those skills need a matching edit in the same commit.

Run all of these after the metadata is written and the steps are built, so every issue surfaces together:

  1. Typos — /check-metadata-typos on each edited .meta.yml (garden first, then grapher). Accept or skip each suggested fix.
  2. Jinja spacing and style guide — /check-metadata-style on the grapher step. A mechanical pass first catches template artifacts (doubled spaces, stray newlines, leading or trailing whitespace) that only appear after Jinja rendering; then it audits user-facing fields (title, subtitle, description_short, display.name, presentation.*) against OWID's Writing and Style Guide (.claude/skills/check-metadata-style/STYLE_GUIDE.md).
  3. Clarity for a general audience — read every user-facing field with non-specialist eyes; the other checks enforce structure and style, this one judges whether the text is understandable. Flag and rewrite (propose concrete rewrites, don't just flag):
    • Acronyms or technical terms not expanded on first use (skip GDP; expand GWIS, MFI, SDG, IHME)
    • Sentences that only make sense if you already know the data source
    • Quantitative claims with no unit context surfacing anywhere in the user-facing text
    • Inconsistent terminology between indicators in the same dataset
    • Domain phrases with a plain-English equivalent ("anthropogenic emissions" → "human-caused emissions")
    • Methodology-attribution claims ("following guidance from <agency>…") — open the cited link and confirm it actually says that; agencies revise methodology
    • Scope qualifiers present in the origin title but absent from user-facing text (private-only, adults-only, market-exchange-rate-only)
    • Text that adds nothing to what the reader has already read — a bullet repeating another bullet, description_short, or the title at the same level of detail. Expanding the short line is fine and expected; restating it is padding. See the description_key guidance above.
  4. Link verification — every URL and [term](#dod:term) in the text: URLs must resolve (curl as the batch primary; on a 4xx from an OWID link, double-check with WebFetch + Wayback before acting); dod slugs checked against the dods table via public Datasette (SELECT name FROM dods WHERE name LIKE ...); a missing dod → keep the link and list it as a "create in admin" follow-up in the PR body.
  5. Dimension sweep — for dimensional indicators (Jinja over a dimension, or a definitions: key several variants reference), every sentence must hold at every value it renders on, not just the one it was written for. Render the text per dimension value and read each output as a reader of that chart, asking what the view already restricts: a caveat that the data doesn't control for X is wrong on the variant grouped by X; a scope word like "all employees" overclaims on a variant filtered to a subgroup; a sentence about a toggle is wrong on views that exist for only one choice of that dimension. Prefer qualifying the wording so it holds everywhere (often one word, nothing extra to maintain) over adding a Jinja branch or a view-level override. Automated reviewers catch this class reliably, so sweeping first saves a review round.
  6. Adversarial claims verification — /fact-check-dataset scoped to the newly written or edited metadata text only: treat each added/changed sentence as a claim and verify it against the producer's documentation (read what's behind the links — check 4 only proves they resolve). Catches text that is well-formed but factually wrong: stale methodology attributions, scope overclaims, misread units in prose. Scope by context: mandatory in /edit-faust-metadata (claims-only, no data cross-checks — cheap); the full-dataset review including data-value cross-checks stays the opt-in step described in /update-dataset §6c-bis (token-heavy).

If any check rewrites a .meta.yml, re-run the affected step so the built catalog reflects the edits (add --grapher when the step is on the grapher channel, otherwise staging keeps serving the old text), then re-run the check to confirm zero remaining violations.

Reviewing the change on staging (Metadata Diff)

The checks above read the text as authored. The Metadata Diff Wizard page reads it as rendered — indicator metadata merged with any view-level override, the way the site resolves it — and lists every chart, MDim view and explorer view each edit lands on:

http://staging-site-<container_branch>/etl/wizard/metadata-diff

It is not an eighth check: it needs a staging server and a human's judgment, and it runs after the checks, once the text is settled. What it shows that nothing above does is reach — reword one shared description_key and dozens of charts change while every chart-config diff stays empty, because that text is inherited, not configured. (Chart Diff compares configs; this compares rendered texts. It is also not Chart Diff's own per-chart "Metadata differences" modal.)

Two preconditions, both silent when unmet:

  • Staging only — the page refuses to run against production. Its baseline is production where the server has production credentials, staging-site-master otherwise: the same baseline Chart Diff uses.
  • The grapher step must be on that server (STAGING=1 .venv/bin/etlr grapher://grapher/<path> --grapher). Its scope is the branch's changed step files intersected with the datasets actually rebuilt there, so text that never reached the server simply doesn't appear. A server that is behind gets a 🚧 banner naming each stale dataset and the command to rebuild it — read that banner before trusting any count, because a stale dataset reports its diffs backwards, showing its older text as this branch's change.

Reviewing marks each chart / MDim view / explorer view ✅ or ❌ with a note; the Review section then exports metadata-rejections.md, the rejections written as instructions naming the edit, the garden .meta.yml it was authored in, and the dataset owner. Nothing here gates the merge — rejecting changes no text — so an unactioned rejection is an open item somebody has to carry (.claude/docs/open-items.md). Verdicts are bound to the wording they were made on: reword the text and the decision reopens rather than counting as done.

Quality Checklist

Items that are easy to miss (obvious rules like "set title" are omitted — see field guidelines above):

  • description_short adds value beyond the title — if it just repeats the title, delete it
  • description_key includes concrete examples where scope is abstract
  • description_key only describes data that actually exists in the indicator
  • description_processing matches current code and doesn't reference internal dataset names
  • Important caveats surfaced in description_short or description_key, not buried in description_processing
  • numDecimalPlaces set and consistent across related variables
  • display.name paired with title_public when set
  • No TODO/FIXME, no copy-paste errors, no producer names/years in titles
  • Repeated text uses anchors/aliases or {definitions.xxx}; Jinja for 10+ similar variables
  • Prose uses curly apostrophes/quotes (’ “ ”), not straight (' ") — straight marks are an LLM tell; see check-metadata-style/STYLE_GUIDE.md
  • Typo check passed

© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/generate-metadata of owid/etl.

Open the folder on GitHubat commit bf5dc8e

Compare with similar skills

Generate Metadata next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generate Metadata compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generate Metadata this skillowid/etl158—~6.8kAutomated safety check: PassMIT
Dinobase Business Data Querieskappa90/dinobase263—~1.5kAutomated safety check: PassCustom licence
Analytics Engineerborghei/Claude-Skills881—~3.4kAutomated safety check: PassMIT
Data Engineering Data Driven Featureaiskillstore/marketplace4307 repos~3kAutomated safety check: PassNone
Profiling Tablesastronomer/agents451—~964Automated safety check: PassApache-2.0
Data Researchermajiayu000/claude-skill-registry6661 repos~4.6kAutomated safety check: PassMIT

Similar skills

  • Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.

    263 GitHub stars~1.5k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Engineering Data Driven Feature

    aiskillstore/marketplace

    Build features guided by data insights, A/B testing, and continuous measurement using specialized agents for analysis, implementation, and experimentation.

    430 GitHub starsUsed in 7 repos~3k tokens
    Data & AnalyticsAuto-check passed
  • Profiling Tables

    astronomer/agents

    Deep-dive data profiling for a specific table. An agent skill from astronomer/agents.

    451 GitHub stars~964 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Data Researcher

    majiayu000/claude-skill-registry

    Data discovery and analysis specialist focused on extracting actionable insights from complex datasets, identifying patterns and anomalies, and transforming raw data into strategic intelligence.

    666 GitHub starsUsed in 1 repo~4.6k tokens
    Data & AnalyticsAuto-check passed
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    451 GitHub stars~1.3k tokensUpdated today
    DatabasesAuto-check passed

More from owid/etl

All 35 skills in this repo
  • Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.

    158 GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.

    158 GitHub stars~19k tokensUpdated today
    Auto-check passed
  • Add new survey question codes (e.g. An agent skill from owid/etl.

    158 GitHub stars~11k tokensUpdated today
    Auto-check: notes
  • Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…

    158 GitHub stars~8.3k tokensUpdated today
    Auto-check passed
  • Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.

    158 GitHub stars~9.6k tokensUpdated today
    Auto-check: notes
  • Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.

    158 GitHub stars~7.3k tokensUpdated today
    Auto-check: notes

Questions about Generate Metadata

What does Generate Metadata do?

A skill your agent uses when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID…. Generate Metadata is an agent skill from owid/etl. Use when creating or enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection, data exploration, and web research following OWID metadata standards.

When should I use Generate Metadata?

Generate Metadata fits situations like: enriching metadata for OWID ETL datasets - generates comprehensive YAML metadata from dataset inspection; data exploration; web research following OWID metadata standards; editing .meta.yml files.

How do I install Generate Metadata in Claude Code?

Run `npx skills add owid/etl --skill generate-metadata -a claude-code`. Or copy the skill folder (.claude/skills/generate-metadata in owid/etl) into .claude/skills/generate-metadata in your project. Claude Code loads it when a task matches its description.

How do I install Generate Metadata in Codex?

Run `npx skills add owid/etl --skill generate-metadata -a codex`. Or copy the skill folder (.claude/skills/generate-metadata in owid/etl) into .agents/skills/generate-metadata in your project. Codex loads it when a task matches its description.

Can I use Generate Metadata in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill generate-metadata -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generate-metadata, .gemini/skills/generate-metadata, .github/skills/generate-metadata and .opencode/skills/generate-metadata in your project.

What does Generate Metadata need to run?

Going by SKILL.md and its folder, Generate Metadata needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Generate Metadata access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Generate Metadata safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Generate Metadata use?

Generate Metadata is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generate Metadata use?

About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Generate Metadata?

Skills that share tags, products or a category with Generate Metadata: Dinobase Business Data Queries (kappa90/dinobase, 263 stars), Analytics Engineer (borghei/Claude-Skills, 881 stars), Data Engineering Data Driven Feature (aiskillstore/marketplace, 430 stars) and Profiling Tables (astronomer/agents, 451 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generate Metadata?

owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.