Veomni Review
ByteDance-Seed/VeOmni
Pre-PR code review gate. An agent skill from ByteDance-Seed/VeOmni.
Review an OWID ETL data update PR end-to-end — runs the pipeline, compares snapshot fields against the previous version, verifies links, audits indicator metadata coverage, and cross-checks workflow…
$ npx skills add owid/etl --skill review-data-pr -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl review-data-pr --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/review-data-pr .claude/skills/review-data-pr && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "review-data-pr" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/review-data-pr into .claude/skills/review-data-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-data-pr", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/review-data-prType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill review-data-pr -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl review-data-pr --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/review-data-pr .agents/skills/review-data-pr && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "review-data-pr" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/review-data-pr into .agents/skills/review-data-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-data-pr", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill review-data-pr -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl review-data-pr --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/review-data-pr .cursor/skills/review-data-pr && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "review-data-pr" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/review-data-pr into .cursor/skills/review-data-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-data-pr", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/review-data-pr--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill review-data-pr -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl review-data-pr --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/review-data-pr .gemini/skills/review-data-pr && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "review-data-pr" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/review-data-pr into .gemini/skills/review-data-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-data-pr", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl review-data-prInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill review-data-pr -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/review-data-pr .github/skills/review-data-pr && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "review-data-pr" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/review-data-pr into .github/skills/review-data-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-data-pr", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill review-data-pr -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl review-data-pr --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/review-data-pr .opencode/skills/review-data-pr && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "review-data-pr" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/review-data-pr into .opencode/skills/review-data-pr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-data-pr", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
review-data-prReview an OWID ETL data update PR end-to-end — runs the pipeline, compares snapshot fields against the previous version, verifies links, audits indicator metadata coverage, and cross-checks workflow…
Review Data PR is an agent skill from owid/etl. Review an OWID ETL data update PR end-to-end — runs the pipeline, compares snapshot fields against the previous version, verifies links, audits indicator metadata coverage, and cross-checks workflow items from /update-dataset. Trigger when the user asks to "review this PR", "review the data PR", or invokes this on an open dataset-update branch.
Its SKILL.md is about 15k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Pull requests, Data pipelines and ETL and End-to-end testing. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bf5dc8e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghrgmakegitpythonmysqlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
catalog.ourworldindata.orgAlso links to:
ourworldindata.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Review Data PR loads about 15k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 8,216 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
ES_DIR/.github/workflows/` from the etl `.env`, asking the user where their checkout is if the variable isn't set — a fiAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from owid/etl at commit bf5dc8e, republished under its MIT licence (© owid). 8,216 words, ~15,250 tokens.
.claude/skills/review-data-pr/SKILL.md (or your agent's skills folder).End-to-end review of a dataset-update PR. Goes deeper than /review: actually runs the steps, compares to the previous version, audits metadata coverage against a fixed checklist, and reports on /update-dataset workflow status (Slack draft, Codex review, indicator upgrade, downstream deps).
Paired skill — keep in sync.
/update-datasetis the author-side counterpart of this skill: the steps it defines are the outcomes verified here. Whenever you add, remove, or change a check in this file, check whetherupdate-dataset/SKILL.mdneeds a matching author-side step (and add it in the same commit if so). The reverse also holds — see the mirror note there. The creation-side skills/create-datasetand/create-snapshotbelong to the same family: the checks here (§5 snapshot fields, §6 links, §7 code clarity, §9 metadata coverage, §10 quality) also gate PRs produced by/create-dataset, so when one of them changes, check whether the create skills need a matching edit in the same commit too.
gh pr list --head <branch>.gh pr view <num> --json title,body,isDraft,mergeable,statusCheckRollup,comments,reviewsFlag if PR description is empty (per user's standing rule: keep PR body in sync with substantial changes).
Flag 🟡 if the Summary doesn't open with a tracking-issue link (Tracks: owid/owid-issues#NNNN) — /update-dataset requires it as the first line; most data updates have a corresponding owid-issues ticket.
gh pr view <num> --json files --jq '.files[] | "\(.additions)+ \(.deletions)- \(.path)"'For very large diffs (>1MB) skip gh pr diff and read the changed files directly with Read.
From the changed files, identify:
snapshots/<namespace>/<new_version>/<short_name>.<ext>.dvc (a .py upload script is optional — .dvc + url_download or a local file path is enough)etl/steps/data/{meadow,garden,grapher}/<namespace>/<new_version>/<short_name>.{py,meta.yml}dag/archive/*.yml or by grepping for the same <short_name>)Open with a short pipeline brief. The reviewer is often seeing this dataset for the first time. Run the same overview that /update-dataset step 0b uses on the new version. Run it on the old version too while it's still active in the PR's DAG (the script only reads active steps), and compare the two chains:
.venv/bin/python .claude/skills/update-dataset/scripts/pipeline_overview.py <namespace>/<new_version>/<short_name>
.venv/bin/python .claude/skills/update-dataset/scripts/pipeline_overview.py <namespace>/<old_version>/<short_name> # if still activeGive the user about 5–8 lines: the chain, what each non-trivial step does, external inputs, consumers, and who did the previous update (git log --diff-filter=A on the old garden script, if the script can't reach it). Then compare against the author's Pipeline structure: … line in the PR Summary. If the line is missing, flag 🟡. If it says "unchanged" but any trigger below holds, flag 🔴.
Before running the pipeline, classify the PR. If any of the following are true, you're reviewing a restructure, not a version bump, and several downstream checks apply differently:
short_name changed (old version uses one name, new version uses another).When it's a restructure:
.py step copy from the old version. Step files should be authored from scratch, not produced by etl update rename. If the new step files look mechanically renamed (same logic, just version-bumped strings), flag 🟡 — the author may have skipped restructure-specific decisions.selectedEntityNames exist in the successor's data (v1 regional aggregates often don't — expect the garden step to rebuild them, mirroring the retired step's method), that pinned yAxis bounds don't clip the new range, and that the subtitle doesn't still describe the old construction. Any of the three broken: 🔴 (the default view renders empty, clipped, or mislabeled)./latest drafts are not expected in the PR body at all. /update-dataset keeps them in the author's workbench/ (steps 9 / 9b, owned by /draft-data-update-slack-post and /owid-staff:draft-data-update-post), so their absence from the PR is correct — don't flag it..venv/bin/etlr data://grapher/<namespace>/<new_version>/<short_name>
.venv/bin/etlr grapher://grapher/<namespace>/<new_version>/<short_name> --grapher --force --onlyThe --grapher upload is required to verify MySQL ingestion and to enable later checks (chart count, indicator upgrade verification). Confirm:
.dvc is committed, otherwise re-fetched)dataset id and shows variable upsertsShortcut: read DB checks off the populated staging server. OWID provisions a staging-site-<branch> server (via Buildkite) that runs the ETL chain and uploads to its MySQL. Once it's built, you can read the DB-dependent checks (chart count, attributionShort, rendered titles/Jinja coverage, indicator-upgrade, ghost variables) straight off staging instead of re-running --grapher locally — which also avoids re-triggering step side-effects (e.g. a grapher step that exports to Google Sheets). Confirm the staging ETL build actually ran and finished before trusting it: query staging-site-<branch> for the new dataset's variables (they exist) and check the owidbot PR comment shows a chart-diff block (✅) — that comment is produced after the staging build. ⚠️ Do not use the GitHub build-and-deploy check as that signal — it's the docs Cloudflare Pages deploy (.github/workflows/deploy-docs-cf.yml: make docs.build → deploys site/), with no ETL chain or Grapher upload, so a green build-and-deploy says nothing about pipeline correctness or the data DB. Reserve a local build for what the staging DB can't answer — chiefly entity-level canonicalization (§8c #2) (data lives outside MySQL). If you can't confirm staging is populated, run the pipeline locally per the steps above, and say in the report whether correctness rests on the staging build or a local run.
Review the actual PR head, not a stale local checkout. The local branch can lag origin (or carry an in-progress merge). Before reading step files locally, git fetch and confirm your tree matches the PR head — git diff HEAD origin/<branch> --stat should be empty, and gh pr view <num> --json headRefOid should match git rev-parse HEAD. gh pr view --files / gh pr diff and the staging DB always reflect origin; local Reads do not. If they diverge, sync (or review via gh pr diff) before trusting local files.
Read both .dvc files (old and new) and produce a side-by-side table for these fields:
| Field | Check |
|---|---|
title | Reasonable update if scope changed |
description | Updated to reflect new source / scope |
date_published | Should normally differ from date_accessed — source from url_main or the file. Equality is legitimate only as the documented fallback when no producer release date is discoverable (e.g. a scraped page carries fresh rows but no updated stamp — see /update-dataset Guardrails, "Scraped chart embeds"); expect a .dvc comment explaining it, and flag 🟡 for the author to confirm rather than 🔴. Bare equality with no rationale: ask. |
date_accessed | Updated to today (or run-date) |
producer / attribution_short | Same source, same values (unless changed deliberately) |
citation_full / attribution | Year bumped to the new release year — etl update copies both verbatim from the old .dvc, so a stale year ships silently. 🔴 if still the old version's year. |
citation_full year vs date_published year | Warn (🟡) if they differ. The year inside citation_full (and attribution) should normally match date_published's year. A mismatch is sometimes legitimate — the producer labels the release by edition rather than publish date (e.g. UN IGME's "2025 report" published 2026-03-17, so citation_full (2025) ≠ date_published 2026) — but it's just as often a stale citation the author forgot to bump. Surface it for the author to confirm; don't silently pass it. |
url_main | Status check — see step 6 |
url_download | Status check; OK to remove if data is now fetched via API |
license.url | Status check |
version_producer | Unchanged label + changed payload = in-place revision. If the producer's version label is the same as the old .dvc but the data changed, confirm the author verified the revision against the source's file-modification dates/hashes (not the label) and documented the behavior in a .dvc NOTE; date_published should be the replacement date. Missing NOTE on a known in-place reviser: 🟡. |
.py scrapes the producer's page or a chart platform's endpoint, re-fetch the producer's page and compare against the committed snapshot — the endpoint the script reads can lag the page (e.g. a Datawrapper chart CDN trailing the page's own <noscript> data tables by a full release, so the committed snapshot silently misses the newest wave). The committed data must match the page's current tables; a missing latest row/wave is a 🔴 (see /update-dataset Guardrails, "Scraped chart embeds").Run the HEAD-check loop from /update-dataset § 6c on every URL in the new .dvc and .meta.yml files. A curl non-2xx is a signal, not proof — Cloudflare-fronted hosts return false 404s to curl. Apply the same escalation as /update-dataset § 6c: re-check with WebFetch, then the Wayback availability API — remembering that no automated signal is decisive (hosts like BLS block both curl and WebFetch while serving browsers fine, a Wayback capture is historical evidence only, and a missing capture is non-evidence). A URL that fails all automated checks is a 🟡 — needs a human browser check (report the evidence trail: statuses, capture date or absence); escalate to 🔴 only once a browser check confirms the link is dead or the producer's site documents its retirement. A curl-only failure that WebFetch resolves is 🟢 informational.
docs.google.com 200 ≠ publicly viewable. Google Sheets/Docs links return HTTP 200 even when they're behind a permission wall (the 200 is the "request access"/sign-in page). When a user-facing description_key/description_processing links a Google Sheet, confirm real public access with WebFetch (ask whether the page shows data or a "you need access"/sign-in wall) — curl status alone will pass a private sheet.#fragment, run the anchor pass from /update-dataset § 6c: the fragment must match an id/name attribute in the page HTML (skip non-DOM fragments: any fragment containing = or / — text fragments, gid=, page=, hash routes like FAOSTAT's #data/FBS — plus #! hashbangs). Rule out client-side rendering and Cloudflare challenge bodies (WebFetch de-slugged-heading check) before flagging. A confirmed missing anchor on an otherwise-working page is 🟡 — the page loads, the reader just lands at the top; escalate to 🔴 only if the linked section genuinely no longer exists and the link's claim depends on it.For each step file, check:
--cli-flag parametrization (or at minimum a clear update comment). Note: a .py upload script is optional — many snapshots ship with only the .dvc and a url_download. Don't flag the absence of a script.paths.regions.harmonize_names(tb, ...) (the new API), not the legacy geo.harmonize_countries/update-dataset §5b-bis)unit == "Number of deaths") get summed while everything else gets population-weighted-averaged — a count series carrying a different unit label silently falls into the averaging branch and produces a meaningless regional "total". (Real case: IGME summed "Number of deaths" but "Number of stillbirths" fell through to the rates path, so regional stillbirth totals became population-weighted averages — Africa showed ~60k instead of ~1M.) Enumerate the distinct units, confirm each is routed correctly (a region's count value should be ≫ any member country's, not a mid-range average), and prefer a robust predicate (unit.startswith("Number of")) over an exact match. Catches a class of bug a green pipeline + Jinja-renders-fine review will miss.export_table_to_gsheet(...) / get_team_folder_id() in a garden/grapher step looks like a CI/deploy risk, but several OWID helpers early-return unless OWID_ENV.env_local == "dev" (so they no-op on staging/prod). Check the helper's guard before flagging — "unconditional call" ≠ "runs everywhere". If it is guarded, it's at most a 🟢/style note (intent could be made explicit at the call site; the dev-only side-effect can leave the exported artifact stale relative to prod), not a blocker.default_metadata=ds_garden.metadataRun the /check-outdated-practices skill on every new step file (snapshot, meadow, garden, and any helper modules like *_omms.py). It reads vscode_extensions/detect-outdated-practices/src/extension.ts as the single source of truth and greps the full pattern set — don't hand-maintain a copy of the patterns here, and don't eyeball helper calls and decide they look current (the geo.add_* family looks fine but is flagged). Report every hit it returns as 🟡.
Separately, the metadata/origin-stripping patterns from CLAUDE.md (pd.concat→pr.concat, pd.to_numeric/pd.to_datetime→pr.*, np.where, index.map(...), pd.DataFrame(tb) re-wrap) are not part of the extension — they're covered by the §7 code-clarity pass. Flag them there even when copy_metadata/fillna appears to mitigate.
/update-dataset steps 1c+6a (annotations) and 1d+5b (sanity_checks) define the catalog/resolve procedure. As reviewer, verify the outcome:
# NOTE: / # TODO: / # FIXME: / # HACK: / # XXX: that are unchanged from the old version. For each, confirm the PR body mentions whether the workaround is still needed, or that it was deleted with its code. Unresolved + undocumented = 🟡.SHOW_SANITY_CHECK_LOGS, DEBUG, LONG_FORMAT set to True. If a debug flag was left enabled, that's a 🔴 — must be reverted.sanity_checks function, scan for drop, filter, tb = tb[...] — row removals that the user might miss. Make sure the PR body lists them.# Sanity check block), the PR body should carry a "Sanity-check findings" section reporting what the checks said on the new data. A green pipeline run is not proof the invariants held — checks that paths.log.warning(...)/.critical(...) instead of assert/raise pass silently. If the new garden chain has logging-style checks and the PR body has no findings section, re-run the garden step (--private --force --only) and scan stdout/stderr for warning, dropped, outlier, AssertionError. Undocumented findings = 🟡; a check that newly raises on the new data = 🔴 (must be triaged with the author per /update-dataset §5b)./update-dataset §5c defines the full audit (validate .countries.json targets against the canonical regions + income-groups catalogs, audit .excluded_countries.json, scan the garden log for the three warnings, and confirm garden-output entities are canonical). As reviewer, verify the outcome — every entity reaching Grapher must be canonical, and any that isn't must be documented in the PR body.
Run after the §4 pipeline build. Three checks:
Garden log warnings. Re-run the garden step capturing output and scan for the three stable warning strings:
.venv/bin/etlr data://garden/<namespace>/<new_version>/<short_name> --force --only \
> /tmp/<short_name>_harmon.log 2>&1
rg -n "missing values in mapping\.|unused values in mapping\.|Unknown country names in excluded countries file:" /tmp/<short_name>_harmon.logmissing values in mapping (source countries not in .countries.json) is the actionable one — 🟡 unless the PR body documents the gap. unused values in mapping / Unknown … excluded are informational 🟢.
Garden-output entities are canonical. This is the check that catches inline tb["country"] = "…" assignments and post-harmonization mutations the .countries.json review can't see. This one needs a local build — entity lists aren't in MySQL (modern grapher stores indicator data outside the DB), so make query can't answer it; build the garden step and load it with owid.catalog.Dataset("data/garden/<ns>/<v>/<short>"). Build the canonical set (regions + latest income groups) and diff against the entities actually in the built garden tables — see the Python snippet in /update-dataset §5c (Python checks #3 + #5). Note geo.REGIONS already includes the four WB income groups, so an isin(REGIONS) filter de-dups them too. Any entity in the garden output that isn't in canonical regions or income groups is 🔴 unless it's a legitimately custom source aggregate (e.g. " (ILO)"/" (WB)"-suffixed regions, BRICS, G7) that the PR body explicitly notes lives outside the canonical system.
Over-exclusion. If .excluded_countries.json exists, flag any entry that is a canonical region/aggregate (/update-dataset §5c Python check #4) — dropping a real country/region silently is 🟡 unless the PR body says why (e.g. source double-counts "World").
If the garden step doesn't use the harmonizer at all (no .countries.json; country assigned inline), checks #2 and #3 still apply — #2 is the only thing that catches non-canonical inline values.
Cheap and worth doing on any PR more than a few days old. A branch serves its own snapshot of every other dataset, so if another update merged to master meanwhile, the branch's staging is behind on that dataset — and any chart combining both differs from production on two axes. Approving it syncs the stale config back and reverts the other update on a published chart. Nothing flags this: CI is green and the chart renders.
Check git log HEAD..origin/master --oneline for 📊 dataset commits; for each, look for charts carrying indicators from both datasets and compare every dimension's dataset version, staging vs production — not just the dimension this PR touches. The fix is merging master in and remapping the affected charts' foreign dimensions; charts using only the other dataset are out of chart-diff scope and never sync, so they're correctly left alone. 🔴 if a shared chart would regress.
The same revert hits a single chart edited on production after the staging server was built. The indicator upgrade saves the staging copy, which predates that edit, so approving the diff syncs the old config back and undoes it. For the upgraded charts, compare production's charts.updatedAt against the last pre-upgrade revision on staging (chart_revisions); a production edit newer than that must be restored on staging before approval. 🔴 if an approved diff would revert one. (Also why etl approve approves nothing after an upgrade, and how to compare configs fairly: /update-dataset, Final QA.)
The author-side audit is optional to run in /update-dataset (the check-empty-entities skill sweeps every chart/MDim/explorer/narrative/gdoc surface, which can consume many tokens) — so a missing audit is not a finding. If the author ran it, verify the outcome: a selection that had data on production but none on staging is a 🔴 regression from the update; a gap identical on production is 🟡 pre-existing — it still needs fixing (chart-config edit or content follow-up on the gdoc), just not necessarily in this PR, so confirm the PR body documents it and a fix is planned.
If the author didn't run it, you MUST offer it to the user as an optional add-on to this review (name the token cost) — surfacing this offer is mandatory, never silently skip it — and recommend accepting when the risk is real: many charts remapped, hand-curated (non-auto) mappings, a restructure, or indicators whose country coverage shrank. Run the full sweep on opt-in.
For a cheap version of the same question — which surfaces carry this dataset at all, without the per-view availability checks — run find-chart-references --dataset-id <id>. It's the surface list both step-7 audits are built on, and it answers "did the author miss a surface entirely" in one query.
Either way, do a cheap manual spot-check as part of the base review: open 2–3 of the most-viewed upgraded charts on staging (SVG render is enough) and confirm their pinned entity selections still draw lines — an empty published chart is a 🔴 however it's found, and a spot-check hit is itself a reason to recommend the full sweep.
/update-dataset step 7 runs the check-hardcoded-years skill after all remaps — it sweeps the same surfaces as 8d (charts, map tabs, MDim views, explorer views, narrative charts, article time= embeds/links) for numeric minTime/maxTime/timelineMinTime/timelineMaxTime/map.time pins and grades each against the new indicators' latest time. It's standard, so a missing audit is a finding (🟡). Verify the outcome either way with a cheap spot-check: query staging for the dataset's chart configs, filter numeric pins client-side ("latest"/"earliest"/absent are fine), and compare against the new data's latest time. When the release added a partial year (only some series reach it), also check the inverse: single-time discrete views on partially-published indicators should carry a deliberate timelineMaxTime clamp at the last complete year (see the chart-level incomplete-latest-year exception in check-hardcoded-years) — an unguarded one renders a tolerance-backfilled, mixed-vintage bar.
maxTime/map.time/timelineMaxTime pin below the new latest time = the update is invisible on that surface — 🟡 unless it's deliberate (pinned year in title/subtitle/slug, narrative charts, single-year comparisons); confirm the PR documents the fix or it's already applied on staging (charts ride Chart Diff; MDim/explorer fixes must be in the ETL YAML, not DB-only — a DB-side edit is overwritten at the next rebuild).The mandatory-fields checklist, the dataset.update_period_days requirement, and the presentation.attribution_short non-inheritance gotcha all live in /update-dataset § 6c. As reviewer, build the indicator × field matrix from that checklist and flag any missing field as 🔴.
Quick verification that presentation.attribution_short actually landed on the produced indicators (origin's value does NOT propagate):
make query SQL="SELECT shortName, attributionShort FROM variables WHERE catalogPath LIKE '%<ns>/<v>/<short_name>%'"Any NULL row is a 🔴.
Staging query mechanics. make query re-interprets % and single quotes via shell+make and breaks on the LIKE patterns / quoted strings these checks need. Connect directly instead and feed SQL via stdin or a .sql file: mysql -h staging-site-<normalized-branch> -u owid --port 3306 -D owid < /tmp/q.sql (host = branch lowercased, /._ → -, staging-site- prefix stripped, first 28 chars — see the query: target in the Makefile). Note datasets.catalogPath has no channel prefix — it's <ns>/<v>/<short> (e.g. un/2026-06-09/igme), not grapher/un/...; match catalogPath LIKE '<ns>/<v>/%'. Batch the metadata-gap checks in one file: counts of name='', name LIKE '%<\%%'/'%{definitions%'/'%<<%' (unrendered Jinja), attributionShort IS NULL, descriptionShort IS NULL, plus a double-space/leading-space scan on name and descriptionShort (Jinja whitespace artifacts).
Distinguish regressions from inherited gaps. Before flagging update_period_days / attributionShort / description_short gaps as 🔴, check the previous version's .meta.yml — if the gap was already there, it's pre-existing (carried over by etl update), not introduced by this PR. Still report it (the fix is cheap and the standing convention wants it), but say so: a pre-existing gap is a 🟡 "worth adding while you're here", not a regression that blocks the update. A field that was set and is now missing is the real 🔴.
Additional reviewer-side metadata checks:
dataset.owners lists the PR author. /update-dataset step 1a-bis requires the author — the person who ran the update / opened the PR — to append their canonical OWID name (an entry in the schemas/dataset-schema.json enum) to the garden .meta.yml owners: list, preserving the existing order and any # review / # backport / # fasttrack markers. Running an update makes you a contributor, so the author belongs there alongside the original owners (don't drop the originals). Verify the PR author is present: missing = 🟡; existing owners reordered or dropped = 🔴. Note the actor running this review skill is usually a different person (the reviewer) — the reviewer does not add themselves; the check is only that the author is listed. The exception is a self-review (author reviewing their own PR), where reviewer and author are the same person.processing_level: major must come with description_processing. Grep the new garden meta.yml for processing_level: major. Each occurrence (whether on definitions.common or per-indicator) requires a description_processing field on that same scope. 🟡 mismatch.description_processing should describe the indicator's own derivation, not just point at a shared generic note. When every aggregate indicator's description_processing is the exact same string (e.g. all four region indicators just reference {definitions.description_regions_processing} with no per-indicator detail), 🟡 flag — author should compose per-indicator sentences.proportion) with <% if <dim> == "X" %>...<% endif %> blocks for title, description_short, display.name, verify every active (dim1, dim2) cell renders a non-empty value. Easiest check: read every column from the grapher dataset and assert metadata.title is non-empty.paths.regions.add_population(tb) / paths.regions.add_aggregates(tb, regions=[...]) auto-resolve their DAG dependencies. If the garden step loads population (or income_groups) via paths.load_dataset(...) but never passes the dataset to anything, that's dead code — 🟡. The DAG dependency still needs to be declared either way.High-income countries, Upper-middle-income countries, Lower-middle-income countries, Low-income countries) are in the REGIONS list, the income_groups DAG dep is declared, and description_regions_processing references the income groups article. 🟢 informational if absent — not all datasets need this, but it's worth surfacing.sort: label order or a category map), compare the declared labels against the values that actually appear in the built grapher data. Labels declared in sort: (or in a category map) but never produced clutter chart legends with empty buckets. Load each categorical column from the grapher dataset, take its unique values, and diff against the sort: list:from owid.catalog import Dataset
ds = Dataset("data/grapher/<ns>/<v>/<short_name>")
tb = ds["<table>"]
present = set(tb["<col>"].dropna().astype(str).unique())
# compare `present` against the `sort:` labels in the .meta.ymlsort:/map label with no backing value is 🟡 — author should drop it from sort:/description_key (or from the map if it can never occur). Re-check on every refresh: phantoms reappear when a category drops out upstream.+ Column lines in the data-diff (the HTML report buries additions at severity 0; read the text/JSON output). For long-format tables a new series adds dimension combinations (rows), not columns — the grapher-level shortName anti-join still catches it; a column diff alone does not, and per-dimension value diffs miss new combinations of existing values. Check the meadow diff too: a column new in meadow that never reaches garden was dropped by the pipeline — confirm the author surfaced that choice rather than losing it silently (and when meadow hardcodes a column subset, additions won't appear in any diff — spot-check the snapshot's columns). If the update adds indicators, verify (a) the PR body lists them under "New indicators" (author side: /update-dataset step 5) and (b) they meet the same metadata-coverage bar as the rest. New indicators present but nowhere mentioned: 🟡. A +/− pair that is really a rename must have gone through the indicator-upgrade mapping, not the new-indicator list. Finally, check the source's file inventory, not just the snapshotted file: list the release's data files (release page, OSF/Zenodo file API, or the previous .dvc's companion-files # NOTE: when the host keeps no file history) and confirm any file added since the previous cycle was either ingested or documented as a deliberate skip in the PR body — every within-file diff is structurally blind to new companion files (a pre-built index, a summary panel). Undocumented new companion file: 🟡.Run /check-metadata-typos and /check-metadata-style against the new garden + grapher .meta.yml files. See /update-dataset § 6b for the full procedure (typos / whitespace + style / a manual clarity checklist for general-audience readability — apply that checklist here too). Report findings as 🟡 (or 🔴 if a violation breaks rendering or makes the text outright misleading).
Did the PR change reader-facing text? If so, open the Metadata Diff Wizard page on the PR's staging server (http://staging-site-<container_branch>/etl/wizard/metadata-diff) and grade the edits there. It lists every chart, MDim view and explorer view each edit lands on — the only view of a garden text change's actual reach, since inherited text produces no config diff at all. A wording you would reject is 🟡 (🔴 when the new text is misleading or contradicts what the view shows): reject it in the page and hand the author the exported metadata-rejections.md, which already names the edit, the garden .meta.yml and the dataset owner. Check two things before trusting the counts — the page's 🚧 stale-server banner (a dataset the server is behind on reports its diffs backwards) and whether the grapher step reached that server at all. Author side: the QA hand-off in /update-dataset.
Also grep the metadata prose for numbers carried over from the previous release (country counts, category counts, year ranges in description_key/descriptions): validated fields are covered by checks, prose numbers are not — a panel-composition change (dropped country, new category set) silently strands them (see /update-dataset Guardrails, "Grep metadata prose"). Stale prose count: 🟡.
Five further prose checks from /update-dataset § 6b that the skills don't automate:
Redundant user-facing text. Read each indicator's description_key as the reader gets it, together with description_short and the chart title: a bullet that restates another bullet, or the short line, at the same level of detail is padding — 🟡, with a proposed merge into the bullet that already covers it. Unpacking description_short (full definition, how it's measured, what's included) is not redundancy and must not be flagged. Watch new bullets added by this PR against the ones already there, and Jinja variants that restate the shared bullet for only some dimension values.
Dimensional text that only holds for one breakdown. For indicators templated over a dimension (or sharing a definitions: key across variants), read each rendered variant, not just one: a caveat that the data doesn't control for X is wrong on the variant grouped by X, a scope word like "all employees" overclaims on a variant filtered to a subgroup, and a sentence about a toggle is wrong on views that exist for only one choice of that dimension. Text that misdescribes the population on some variants: 🟡 (🔴 when it contradicts what the view actually shows). Author side: /update-dataset § 6b dimension sweep.
Methodology-attribution claims ("following guidance from <agency>…"): open the cited link and confirm it actually says that — agencies revise methodology, and a stale claim survives every link check (real case: metadata cited BEA guidance for an office-PPI deflator after BEA had switched that category to a different composite). A claim the cited page doesn't support: 🔴 (it's factually wrong reader-facing text).
Scope qualifiers in the origin title (private-only, adults-only, market-exchange-rates-only) must surface in description_short/description_key, not only in the citation. Missing: 🟡.
Change-magnitude claims in the PR body (revision medians, % changes): recompute at least the headline number independently from the raw old/new snapshots rather than trusting the author's diff — derived pct columns from assign/sort chains can silently misalign (a real update shipped "median +8.2%" where the true value was +14.3%). A wrong magnitude that feeds an announcement: 🔴.
/update-dataset § 6c-bis offers an adversarial factual review via /fact-check-dataset — optional to run because it's token-heavy (~25–45 web calls), so a missing report is not a finding; don't flag its absence. If the author didn't run it, you MUST offer it to the user as an optional add-on to this review (name the token cost; in review context it runs in that skill's spot-check scope, not the full author scope) — surfacing this offer is mandatory, never silently skip it — and recommend accepting when the update carries red flags: large unexplained value churn, an in-place source revision, a producer new to us, or editorial claims riding on specific values. Run it on opt-in.
When the PR body (or workbench/<short_name>/update-context.yml) does reference an ai/adversarial-review-<short_name>-<date>.md report, verify the outcome: its 🔴 findings must be resolved — metadata edited, or a <short_name>.corrections.yml added with reason/producer/status filled in. Then spot-check independently — don't take the report's word for it: re-verify 2–3 of its findings and 2–3 anchor values (World total + one major country, latest year) against an independent source yourself, following that skill's independence rules (a different producer measuring the same quantity; never OWID republishers or mirrors of the same producer). A value you confirm wrong that the report missed or waved through is 🔴.
/update-dataset step 7 asks the author to read the prose of every surface citing the dataset — articles, data
insights, and especially titles — for quantitative claims the update invalidates. As reviewer, check the PR body
records an outcome for it: either a clean verdict, or a named list of claims with what was decided (handed to
content, or deliberately left with a reason). A clean verdict has to say three things to be acceptable: what was
checked — no unbounded claim, and, on an update that revised any named period (a restatement, a corrected date),
no bounded claim touching a revised observation; the outcome must state which kind of update it was, since on an
append-only update bounded claims are out of scope; what was swept — the surfaces from find-chart-references;
and what was not — the sweep's coverage gaps (the script prints them and writes them with --gaps-json; a
chart nested in an article layout container, or a data insight holding its chart outside grapher-url, can be
missing from its list). A blanket "nothing stale" with none of that is the first failure mode below in mild form —
ask what was checked and swept.
Two ways this goes wrong, both worth a 🟡:
find-chart-references has run, and a headline multiple is the most-read number we publish.A claim deliberately left unchanged is a perfectly good outcome — a data insight whose chart is a static image
cannot have its text updated alone without desyncing it from the picture. Look for the reason, not for a fix.
The exception is a narrative chart's own title or subtitle: it renders live data, so with maxTime: "latest"
an unbounded claim in it is already wrong on staging. Spot-check each narrative chart carrying the dataset's
variables (AdminAPI.get_narrative_chart(id)["configFull"]); an unbounded claim the new data contradicts, with
neither a retitle nor a maxTime pin applied on staging, is a 🔴 — /update-dataset requires the fix in
this PR. If the narrative chart's parent is an MDim view, chart-sync won't carry the staging fix to
production, so the PR's open items must also list the same edit on production after merge; a staging-only
fix with no such item is a 🟡.
The remove-and-reorder procedure is in /update-dataset § "Removing the old version & reordering the DAG". Its two halves land at different times: the old dag entries are removed and archived during the update, but the old step files stay on disk until this review is signed off — the consecutive-version comparison (compare-previous-version VS Code extension) diffs same-named files across YYYY-MM-DD sibling folders on the filesystem and never reads the dag. Old step files already deleted at review time = 🟡 — ask the author to restore them until sign-off (deleting them is the final commit before merge). For the dag itself, verify the outcome (the old entries should be under dag/archive/ already, regenerated by etl archive-dag):
rg "<namespace>/<old_version>/<short_name>" dag/ -g "*.yml" | grep -v "^dag/archive" # should be empty (old version removed from active dag)
rg "<namespace>/<new_version>/<short_name>" dag/ -g "*.yml" | grep -v "^dag/archive" # should be in old slot, not at bottomInternal version consistency (the silent-stale-data bug). etl update occasionally leaves a new step depending on an old-version dep (e.g. new garden still pointing at old meadow or old snapshot), which silently loads stale data. Verify every dep inside the new chain's block is on the new version. Read the new chain's DAG block and confirm none of its dependency lines reference <old_version>:
# Print the new chain's block and eyeball the dependency lines — none should contain <old_version>.
rg -n -A8 "<namespace>/<new_version>/<short_name>" dag/ -g "*.yml" | rg "<old_version>" # should be emptyAny hit here is a 🔴 — the new step is wired to a stale dependency.
Visual inspection of the diff for:
# Source — dataset name.) preserved above the new entries # vs # is a frequent typo)Procedure in /update-dataset § "Downstream dependency check". One-liner:
rg "<namespace>/<old_version>/<short_name>" dag/ -g "*.yml" | grep -v "^dag/archive"After excluding the dataset's own chain, any remaining hits are downstream consumers — flag 🟡 unless the PR body already documents them under a "Downstream dependencies" section.
Silent-breakage check (when consumers were repointed in this PR). Mirrors /update-dataset § "Silent-breakage check". If the PR bumps downstream consumers to the new version (rather than deferring them), a consumer can still build green while quietly losing data — a region aggregate that goes NaN, a reclassified country that disappears, a join that stops matching. A green pipeline run does not prove coverage held. Verify with the two existing instruments:
buildkite/etl-automated-staging-environment status on the PR — the staging bake runs etl run ... --modified --continue-on-failure, which re-raises the first failure at the end, so any consumer crash turns the check red. A red check is a 🔴 in itself, and it also means the data-diff report under-reports (dependents of the failed step are skipped, stay stale in the catalog, and diff as unchanged) — don't accept report verdicts until the check is green. .venv/bin/etlr --modified --dry-run lists the affected scope locally when you need the list.https://catalog.ourworldindata.org/diffs/<sanitized_branch>/data-diff.html, easiest via the full report link in owidbot's PR comment — the path keeps the branch name's dots and underscores and replaces only characters outside [A-Za-z0-9._-] (e.g. /) with - — unlike the staging subdomain, which does replace ./_) — staging (new dep) vs production (old dep) — and read its verdicts rather than scanning the diff: every red "− lost N data point(s)" entry in the "Top changes" list and every dataset with a red coverage chip is a coverage loss to triage (legitimate churn vs. a silent drop); 🔴-tier datasets (median anomaly score ≥ 15%) get a review, 🟡 a skim, 🟢 is noise. The 📝 metadata-only filter separates pure metadata edits. Locally: .venv/bin/etl diff REMOTE data/ --changed --include garden --output-html data-diff.html. (Local build+diff only for small fan-outs; at foundational-dataset scale — hundreds of downstream steps — use owidbot's hosted report and treat any local rebuild as a one-time gate, ~35 min / ~7 GB.) Then run the full-report audit probes from /update-dataset § "Silent-breakage check" over all changed datasets: structural changes / "World" rows / raw-country rows / >30%-of-rows indicators / wipe-vs-edge for every coverage loss. An entity losing all its rows is the bug signature — classically a stale pinned-country requirement nulling a whole income-group aggregate after a reclassification (build-time ValueError since the geo guard, but verify) — while sparse-edge losses are membership churn.Any downstream build failure, dropped table/column/entity, or all-NaN series is a 🔴 (the update silently dropped data downstream) unless the author has already triaged it in the PR body. If consumers were deferred to a follow-up PR, these checks belong to that PR — here just confirm the "Downstream dependencies" list is complete.
Verify the author completed each post-step item from /update-dataset. The procedures live there — here we just confirm the outcomes:
| Item | Verify by |
|---|---|
| Indicator upgrade ran (§7) | make query SQL="SELECT COUNT(*) FROM chart_dimensions cd JOIN variables v ON cd.variableId=v.id WHERE v.catalogPath LIKE '%<ns>/<new_v>/%'" — non-zero |
| Explorers / MDims re-exported (§7) | Only if the DAG has viz://explorer/... or viz://chart/... steps for this dataset (rg -e "viz://explorer/.*/<short_name>" -e "viz://chart/.*/<short_name>" dag/ -g "*.yml"). The indicator-upgrader never touches these, so run the two staging queries from /update-dataset §7 (old-version references in explorer_variables / multi_dim_x_chart_configs) — both must return empty. A hit = 🔴, the viz step wasn't re-run. |
| Hardcoded-time-bounds audit ran (§7 / §8e) | Standard author-side step. Spot-check per §8e: numeric maxTime/map.time/timelineMaxTime pins among the dataset's staging chart configs vs the new indicators' latest time — non-deliberate pins below it must be fixed on staging or documented in the PR body. Skipped audit or an unfixed, undocumented pin = 🟡. |
| Chart-diff bot result | PR comments include <!--chart-diff-start--> block ✅. A diff that shifts every historical year/region by a tiny amount is usually upstream-dataset drift (the live data was built against an older population/regions/income_groups snapshot), not a regression — don't flag it as a 🔴; the real change should be isolable by rebuilding the old version on the current catalog (see /update-dataset §5). |
| Scheduled-issue workflow checked (§6d) | An update-*.yml in owid/owid-issues covers this dataset — exact name update-<namespace>-<short_name>.yml, fuzzy filename, or group workflow (check $OWID_ISSUES_DIR/.github/workflows/ from the etl .env, asking the user where their checkout is if the variable isn't set — a filename-only gh api …/contents listing can't verify cron/body or find group workflows); its cron is consistent with update_period_days + the release cadence evident from the PR (and fires after the producer's release window, not before); the issue body tells the next updater to run /update-dataset <short_name> (space-separated short names for group workflows) and carries no generic to-do checklist (a list written for that specific dataset is fine). Missing workflow for a ≥annual dataset, a cron that contradicts the observed cadence, a stale body or a generic checklist, or a workflow file missing its .yml extension (Actions never runs it) = 🟡. |
| Pinned chart FAUST not stale (§ Guardrails) | Only when the PR changed methodology/scope wording in indicator metadata (grapher_config.note, deflator/source claims): for each affected chart, read the authored layer (the chart_configs row named by patchConfigId) on staging — a chart overriding title/subtitle/note keeps the old text regardless of the metadata fix. Stale method/scope claim on a published chart's pinned FAUST = 🔴 (reader-facing contradiction with the data page). |
| Narrative-chart stale-FAUST sweep ran (§7) | Only when narrative charts carry the dataset's variables (query narrative_charts joined via parent chart_dimensions on staging). The upgrader warns only about charts visited in its own run, so the author must sweep all of them (_find_stale_faust_overrides + a plain grep of the patches for old unit/base-year strings — see /update-dataset §7). Cheap outcome check: grep the narrative patches on staging for the previous release's unit/base-year strings (e.g. "constant <old base year> US$") — a hit on a published narrative chart = 🔴; narrative charts present but no sweep evidence in the PR = 🟡. |
@codex review posted (§9) | gh pr view <num> --json comments shows the trigger comment + a Codex review |
| Codex threads resolved (§10) | Write the query to a file and pass -F query=@file.graphql (see GraphQL note below) — list reviewThreads(first:20){ nodes { isResolved } }; all isResolved: true. |
A clean Codex review has a different shape: no inline comments and zero review threads — just a single top-level "no issues" comment from chatgpt-codex-connector[bot] in the issue comments. That counts as reviewed (the threads row passes vacuously); don't flag the absence of inline threads as "review missing". Codex review state is COMMENTED (not APPROVED); fetch its inline notes via gh api repos/owid/etl/pulls/<num>/comments --paginate --jq '.[] | select(.user.login|test("codex";"i"))'. After @codex review, it usually responds in ~2–5 min — poll with a backgrounded until loop rather than blocking.
GitHub API gotchas (when posting/fetching review comments):
gh api graphql -f query='...' breaks on shell quoting (the inline query gets truncated → Expected NAME parse errors). Write the query/mutation to a file and use gh api graphql -F query=@/tmp/q.graphql with variables as separate -f name=val (string) / -F name=val (typed int/file). Read bodies from a file with -F body=@/tmp/comment.md.PATCH pulls/comments/{id} 404s on a draft, and POST pulls/{n}/comments errors one pending review per pull request. Add with addPullRequestReviewThread(input:{pullRequestReviewId, path, line, side:RIGHT, body}), edit with updatePullRequestReviewComment(input:{pullRequestReviewCommentId, body}). Get the pending review id via reviews(first:50, states:PENDING) and a draft comment's node id via that review's comments.> _Written by Claude <model name> — @<handle> at the wheel._ (model name = the model actually generating the content, e.g. "Sonnet 5"; handle = the human directing the work), same as PR descriptions. Skip it only on genuinely trivial one-liners (an @codex review ping).Out of scope for review: Slack announcement and the QA hand-off (Anomalist, Chart Diff, data-diff report) are author-side concerns, not reviewer checks.
Producer-docs vs. data consistency. If the PR description notes a discrepancy between the producer's documentation (codebook, methodology page, README, release notes — whatever is available) and the actual file shipped, that's a 🟢 informational item — the author has surfaced it for producer follow-up. Don't ask them to "fix" the data to match the docs; the PR should preserve what the source shipped and flag the discrepancy.
Structure the review with:
/latest drafts live in workbench/, not the PR — don't expect them here.).claude/docs/open-items.md plus the update workflow's fourth one (deferred to a follow-up PR — downstream repoints, old-version archiving). Unactioned Metadata Diff rejections belong here — nothing in the merge enforces them. Re-state the full list on every re-review, not just the delta, and mark what cleared since last time.Check the PR body doesn't leave pending work unmentioned. A PR whose description lists only what was done, while the session left content edits pending, audits unrun, or a follow-up PR's scope undefined, is missing the one artifact that survives after the chat is gone — flag it 🟡. Judge it on whether a reader can tell what's outstanding, not on whether it uses any particular headings or wording. Work that was deliberately handed off needs a locator in the body too, not just a description: an item the next person can't act on without redoing the analysis isn't handed off.
update_period_days, missing presentation.attribution_short, stale year in citation_full/attribution, outdated __main__ block in snapshot, DAG reference to old version that should be archived, new step wired to a stale (old-version) DAG dependency, explorer/MDim still referencing old-version variables on staging, non-canonical garden-output entity that isn't a documented custom aggregate, sanity check that newly raises on the new data, downstream consumer that fails to build or silently lost coverage after a same-PR repoint (a "− lost N data point(s)" entry / red coverage chip in the data-diff report) without triage in the PR body, upgraded view (chart, MDim view, narrative chart, or article link) whose pinned entity selection has data on production but none on staging (renders empty), narrative chart whose own title/subtitle states an unbounded claim the new data contradicts, with neither a retitle nor a maxTime pin applied on staging, metadata claim contradicted by the producer's own documentation, data value confirmed wrong by two independent sources without a corrections.yml entry or PR-body triagesort: labels with no backing value, over-exclusion of a canonical region, undocumented missing values in mapping countries, count series mis-routed into a rate/average aggregation branch, citation_full year ≠ date_published year (verify), pre-existing inherited metadata gap (update_period_days/attributionShort/description_short already missing in the prior version), PR author missing from dataset.owners, missing tracking-issue link in the PR body, owid-issues scheduled workflow missing, mis-timed, or not pointing at /update-dataset, pre-existing empty-entity gap surfaced by the optional check-empty-entities audit or a spot-check (identical on production — needs a chart-config or gdoc fix via follow-up, just not necessarily in this PR), hardcoded-time-bounds audit (check-hardcoded-years) skipped by the author or a non-deliberate maxTime/map.time/timelineMaxTime pin below the new data's latest time left unfixed and undocumented (the update is invisible on that surface), "verify manually" item from an adversarial-review report left untriaged (only applies when the author ran the optional review)/review skill is for general PR review — this skill is the dataset-specific superset.--grapher flag when running the grapher step end-to-end — without it, MySQL ingestion is not exercised and indicator metadata in the DB is not verified.© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/review-data-pr of owid/etl.
Open the folder on GitHubat commit bf5dc8e
Review Data PR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Review Data PR this skillowid/etl | 158 | — | ~15k | Automated safety check: Notes | MIT | |
| Veomni ReviewByteDance-Seed/VeOmni | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| source-mssql E2E Test Harnessairbytehq/airbyte | 22k | — | ~4.4k | Automated safety check: Pass | Custom licence | |
| Airbyte source-mysql E2E Testsairbytehq/airbyte | 22k | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 598 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Glue 09 10 Migrationaws-samples/aws-glue-samples | 1.5k | — | ~2.4k | Automated safety check: Pass | MIT-0 |
ByteDance-Seed/VeOmni
Pre-PR code review gate. An agent skill from ByteDance-Seed/VeOmni.
airbytehq/airbyte
Stands up a throwaway local SQL Server 2022 backend, applies SQL fixtures and runs Airbyte spec, check, discover and read against source-mssql images.
airbytehq/airbyte
Stands up a throwaway local MySQL 8.0 backend, applies SQL fixtures and sweeps the Airbyte spec, check, discover and read commands against a source-mysql image.
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.
retentioneering/retentioneering-tools
Help the user turn their Retentioneering ideas, friction reports, bug findings, or feature needs into high-quality upstream contributions: from capturing and validating the idea, through minimal…
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Categories
Review an OWID ETL data update PR end-to-end — runs the pipeline, compares snapshot fields against the previous version, verifies links, audits indicator metadata coverage, and cross-checks workflow…. Review Data PR is an agent skill from owid/etl. Review an OWID ETL data update PR end-to-end — runs the pipeline, compares snapshot fields against the previous version, verifies links, audits indicator metadata coverage, and cross-checks workflow items from /update-dataset.
Review Data PR fits situations like: the user asks to review this PR; review the data PR; invokes this on an open dataset-update branch.
Run `npx skills add owid/etl --skill review-data-pr -a claude-code`. Or copy the skill folder (.claude/skills/review-data-pr in owid/etl) into .claude/skills/review-data-pr in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill review-data-pr -a codex`. Or copy the skill folder (.claude/skills/review-data-pr in owid/etl) into .agents/skills/review-data-pr in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill review-data-pr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-data-pr, .gemini/skills/review-data-pr, .github/skills/review-data-pr and .opencode/skills/review-data-pr in your project.
Going by SKILL.md and its folder, Review Data PR needs the command-line tools its instructions call (gh, rg, make, git, python and mysql). Our summary lists: Python 3.
SKILL.md names 2 domains. In commands or code: catalog.ourworldindata.org; the agent is likely to contact it when it follows the instructions. As links in the text: ourworldindata.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Review Data PR is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 15k tokens (SKILL.md is roughly 61k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Review Data PR: Veomni Review (ByteDance-Seed/VeOmni, 2.2k stars), source-mssql E2E Test Harness (airbytehq/airbyte, 22k stars), Airbyte source-mysql E2E Tests (airbytehq/airbyte, 22k stars) and Crawl4AI Web Scraping (smallnest/goclaw, 598 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.