Agent skill

Fact Check Dataset

by owid in owid/etl

Adversarially review an ETL dataset's data and metadata for factual accuracy — verifies metadata claims against the producer's own documentation (fetched from the links in snapshot .dvc files and…

MITAuto-check passedResearch & Science

Install Fact Check Dataset

skills CLI
$ npx skills add owid/etl --skill fact-check-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install owid/etl fact-check-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fact-check-dataset .claude/skills/fact-check-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fact-check-dataset
GitHub stars
159
Token cost
~6.9k tokens
SKILL.md length
2,823 words
Files
1
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Adversarially review an ETL dataset's data and metadata for factual accuracy — verifies metadata claims against the producer's own documentation (fetched from the links in snapshot .dvc files and…

  • Works in 5 steps: Prioritize indicators by chart views… → Phase 0: source verification (mandatory… → Phase 1: internal anomaly scan (local,… → …
  • The user asks to adversarially review a dataset
  • SKILL.md covers Scope — what this skill is NOT, Inputs, Step 1 — Prioritize indicators… and Step 2 — Phase 0: source…, plus 6 more sections
  • Calls rg; reaches datasette-public.owid.io

What it does

Fact Check Dataset is an agent skill from owid/etl. Adversarially review an ETL dataset's data and metadata for factual accuracy — verifies metadata claims against the producer's own documentation (fetched from the links in snapshot .dvc files and metadata texts) and cross-checks anomalous plus anchor values against independent sources online, to catch unit errors, wrong-year values, and hard-to-detect mistakes made by the source itself. Use when the user asks to "adversarially review a dataset", "fact-check this dataset", "verify the data against the source"…

Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Fact-checking and source verification and Data pipelines and ETL. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.

When your agent uses it

  • The user asks to adversarially review a dataset
  • Fact-check this dataset
  • Verify the data against the source
  • Cross-check the values

Example prompts

  • “adversarially review a dataset”
  • “fact-check this dataset”
  • “verify the data against the source”
  • “/fact-check-dataset”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Prioritize indicators by chart views (heaviness control)
  2. Phase 0: source verification (mandatory before any critique)
  3. Phase 1: internal anomaly scan (local, cheap — runs before any web call)
  4. Phase 2: independent online cross-check (anomaly-led + anchors)
  5. Phase 3: adversarial metadata review (top-N + anomalous indicators only)

What it can do on your machine

Read from SKILL.md and the folder at commit 70c9705. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • datasette-public.owid.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fact Check Dataset loads about 6.9k tokens when it runs. Until then it costs about 163 tokens; SKILL.md has 2,823 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~163
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from owid/etl at commit 70c9705, republished under its MIT licence (© owid). 2,823 words, ~6,944 tokens.

Download SKILL.mdSave it as .claude/skills/fact-check-dataset/SKILL.md (or your agent's skills folder).
name
fact-check-dataset
description
Adversarially review an ETL dataset's data and metadata for factual accuracy — verifies metadata claims against the producer's own documentation (fetched from the links in snapshot .dvc files and metadata texts) and cross-checks anomalous plus anchor values against independent sources online, to catch unit errors, wrong-year values, and hard-to-detect mistakes made by the source itself. Use when the user asks to "adversarially review a dataset", "fact-check this dataset", "verify the data against the source", "cross-check the values", or as the factual-accuracy step inside /update-dataset, /create-dataset, and /review-data-pr.
metadata.internal
true
metadata.owner
paarriagadap

Adversarial data review

Attack the dataset the way a hostile referee would: treat every metadata sentence and every data value as a claim that must survive verification against (a) the producer's own documentation and (b) independent sources — because the producer itself can be wrong.

Two failure classes to hunt:

  1. Our mistakes — misread units, metadata content in the wrong field, stale producer text, scope overclaims, processing bugs that ship wrong numbers.
  2. The source's mistakes — unit slips, wrong-year values, transcription errors, stale pre-revision values. These are hard to detect precisely because our pipeline faithfully reproduces them; only an independent source can expose them.

Scope — what this skill is NOT

Factual accuracy only. Don't duplicate:

  • Style-guide compliance → /check-metadata-style
  • Spelling typos → /check-metadata-typos
  • Jinja whitespace artifacts → /check-metadata-style (its mechanical first pass)
  • Field-coverage / freshness / link-liveness audits → /update-dataset § 6c (that step checks the links resolve; this skill reads what's behind them)

Any rewrites you propose use American spelling.

Inputs

  • A step path garden/<namespace>/<version>/<short_name> (the data:// URI form also works), a bare <short_name> — resolve via the DAG like /update-dataset does (rg "/<short_name>:?$" dag/ -g "*.yml" | grep -v "^dag/archive", latest active version, ask the user if ambiguous).
  • Optional: --top N (deep-review cap, default 10) or --full (deep-review every indicator regardless of chart usage).
  • Also accepts a single chart (slug or id) as the target: scope the whole review to that chart's indicator(s) — verify every displayed country-category/value for the latest year, compare against the previous dataset version, and review the chart's own FAUST text as claims. Drafts count (find them in the staging DB; they have no production row).
  • Precondition: the garden dataset is built locally (.venv/bin/etlr data://garden/<ns>/<version>/<short_name>; PREFER_DOWNLOAD=1 is fine for already-published upstream deps). Exception: the /create-snapshot context runs Phase 0 only against the .dvc and the fetched docs — no meadow/garden step exists yet, so skip this precondition (and Steps 1, 3–5) there.

Where the table says mandatory, the calling skill always runs it, in the background and with effort scaled to the dataset; elsewhere it is optional (it can use many tokens, see Step 1). Scope by calling context:

ContextScope
/update-dataset § 6c-bis (mandatory, background)New/changed metadata text + newly added data values (latest wave/year); deep review = top-N + anomalies
/create-dataset Step 6b (mandatory, background)All indicators if there are few; otherwise the main ones, with a few value cross-checks each
/create-snapshot § 5 (optional)Phase 0 only — verify the .dvc claims against the fetched producer docs (no built dataset yet, so no data cross-checks)
/review-data-pr § 10bOnly if the author ran it: verify outcomes and independently spot-check 2–3 findings and 2–3 anchor values
/edit-faust-metadata (mandatory)Claims-only, on the added/edited metadata text exclusively — verify each new/changed sentence against the producer docs behind the links in the text and the snapshot .dvc. NO data-value cross-checks, anomaly scans, or indicator prioritization (no data changed), and unedited metadata is out of scope — a handful of web calls, not the full review
StandaloneTop-N + anomalies (or --full)

Step 1 — Prioritize indicators by chart views (heaviness control)

The deep per-indicator work (metadata claim review + online cross-checks) costs real time (~25–45 web calls on a typical dataset), so scope it: deep-review the top N indicators (default 10) ranked by summed 365-day views of the charts that use them, plus every indicator the anomaly scan flags. For a brand-new dataset with no charts, review all indicators.

Never cap silently. The report's Part 2 must list every skipped indicator with its rank — a truncated review that reads as complete is itself a factual error.

Views rank prioritization; the DB defines coverage. Analytics only sees published charts — when the task is "all charted indicators" (or the chart under review is a draft), inventory usage from the grapher/staging DB instead (chart_dimensions JOIN variables on catalogPath, no publishedAt filter), then rank the published subset by views.

Map the ranking onto the NEW build before selecting. Pre-merge, the ranked catalog_paths are the old version's — extract a version-independent identity (<table>#<column>, i.e. everything after the version segment) and join it onto the new garden/grapher build's actual table/column list. Old-charted identities missing from the new build are renames — resolve them via the indicator-upgrade mapping before ranking, or list them explicitly. New-build indicators absent from the ranking are uncharted (including newly added ones) — they are candidates for the anomaly-driven track and must appear in the reviewed-or-SKIPPED inventory, never silently dropped by the chart join.

python
from etl.analytics.config import SEMANTIC_LAYER_SCHEMA as S  # tables must be schema-qualified for BigQuery
from etl.analytics.data import read_analytics  # Metabase; auto-falls back to analytics Datasette without creds

NS, SHORT = "<namespace>", "<short_name>"
ind = read_analytics(f"SELECT indicator_id, catalog_path FROM {S}.indicators WHERE catalog_path IS NOT NULL")
# Match namespace + short_name across ANY version — before the PR merges, live charts still point at the OLD version.
ind = ind[ind["catalog_path"].str.contains(f"/{NS}/") & ind["catalog_path"].str.contains(f"/{SHORT}/")]
cxi = read_analytics(f"SELECT chart_slug, indicator_id FROM {S}.charts_x_indicators")
charts = read_analytics(f"SELECT chart_slug, views_365d FROM {S}.charts").drop_duplicates("chart_slug")
usage = (
    # LEFT-join the views: a newly published or unvisited chart has no analytics row yet, and an
    # inner join would silently drop its indicators from the charted inventory (they rank at 0 instead).
    ind.merge(cxi, on="indicator_id").merge(charts, on="chart_slug", how="left")
    .groupby("catalog_path")
    .agg(n_charts=("chart_slug", "nunique"), views_365d=("views_365d", "sum"))
    .sort_values("views_365d", ascending=False)
)

If the analytics layer is entirely unreachable, fall back to the public grapher Datasette and rank by number of charts (it has no view counts):

python
from etl.http import session  # OWID infra → tagged User-Agent

sql = (
    # NOTE: variables.catalogPath DOES carry the channel prefix ("grapher/<ns>/<version>/<short>/<table>#<col>");
    # it's datasets.catalogPath that is channel-less — don't mix the two conventions up.
    "SELECT v.catalogPath, COUNT(DISTINCT cd.chartId) AS n_charts "
    "FROM variables v JOIN chart_dimensions cd ON cd.variableId = v.id "
    "WHERE v.catalogPath LIKE 'grapher/<namespace>/%/<short_name>/%' GROUP BY 1 ORDER BY 2 DESC"
)
rows = session.get("https://datasette-public.owid.io/owid.json", params={"sql": sql}).json()["rows"]

Step 2 — Phase 0: source verification (mandatory before any critique)

Establish what you're looking at before criticizing anything. Investigative, not adversarial.

  1. Collect the source's links. From the snapshot .dvc (url_main, url_download, license.url) and every URL embedded in metadata texts and step files:
    bash
    rg --no-filename -No "https?://[^\"' )>]+" snapshots/<ns>/<version>/ etl/steps/data/{meadow,garden,grapher}/<ns>/<version>/
  2. Compare the source's file-modification dates/hashes against our snapshot's date_accessed/md5 (e.g. the OSF API lists date_modified and hashes per file). Producers replace files in place without bumping version labels — an unchanged version string proves nothing, and an in-place revision is the single highest-yield thing this phase can find. While at it, diff the source's file inventory against what we snapshot: a new companion file (pre-built index, summary table, construction script) is a new-indicator candidate that no within-file diff can surface — route it to the update workflow's "Surface new indicators" step rather than just cataloguing it.
  3. Fetch and READ the producer's own documentation — methodology pages, indicator definitions, codebooks/data dictionaries, release notes. Not secondary commentary, not a blog post about the source: the source itself. Follow the links from url_main to the actual methodology document when the landing page is thin. Access escalation per repo convention: curl → WebFetch → Wayback Machine before treating a 4xx as real.
  4. Check for an existing <short_name>.corrections.yml next to the garden step. Its entries are known, already-handled source errors — acknowledge them in Part 0 of the report; never re-flag them as new findings.
  5. Establish the pipeline. For each metric under review: what does the source publish (exact indicator name, unit, definition, granularity, upstream data — official statistics, modeled, survey)? What does OWID add on top (read description_processing, the garden step code, and the corrections file)? Every later critique must state whose layer it targets. While reading, harvest the codebook's worked examples as test vectors: any country/value/date the documentation itself cites must match the data — a codebook example contradicting the shipped file is the strongest class of source error (provable entirely from the producer's own materials).
  6. Field-placement audit. The .dvc meta.origin.description must carry producer content only; garden description_processing must carry OWID content only; description_from_producer must be verbatim producer text — diff it against the fetched docs (typography-only drift is fine, paraphrase is not). Beyond placement, each field must be factually consistent with what the docs actually say (units, scope, coverage, method).

HARD RULE — proportionality. The severity of any provenance or factual critique must be proportional to the depth of verification you achieved. If you read the documentation and confirmed a gap, make a strong claim. If the docs were unreachable after the full escalation, cap the language at "I was unable to verify … — worth checking before merge" and the severity at 🟢. Never assert an error you couldn't check.

Step 3 — Phase 1: internal anomaly scan (local, cheap — runs before any web call)

Scan the built garden dataset for internal red flags. Paste-and-adapt sketch — adjust the key columns (extra dimensions like sex/age need to join the groupby keys), the thresholds, and which units are additive:

python
from owid.catalog import Dataset
import numpy as np
import pandas as pd

ds = Dataset("data/garden/<ns>/<version>/<short_name>")
findings = []
for tname in ds.table_names:
    tb = ds.read(tname, safe_types=False)  # read() resets the index, so key columns are regular columns
    if tb.index.names != [None]:  # defensive: if keys sit in the index (e.g. the table came via ds[tname]), restore them
        tb = tb.reset_index()
    year_col = "year" if "year" in tb.columns else ("date" if "date" in tb.columns else None)
    has_country = "country" in tb.columns  # year-only tables exist (gravitational-wave counts), as do static ones (GWP factors)
    vals = [c for c in tb.columns if c not in ("country", year_col) and pd.api.types.is_numeric_dtype(tb[c])]
    for col in vals:
        unit = (tb[col].metadata.unit or "").lower()
        s = tb[(["country"] if has_country else []) + ([year_col] if year_col else []) + [col]].dropna(subset=[col])
        if s.empty:
            findings.append((tname, col, "EMPTY", "all-NaN column")); continue
        mx, mn = s[col].max(), s[col].min()
        # 1. Unit/magnitude sniffs
        if any(w in unit for w in ("%", "percent", "share")):
            if mx <= 1.5: findings.append((tname, col, "UNIT", f"unit is % but max={mx:.3g} — fraction stored?"))
            if mx > 150: findings.append((tname, col, "UNIT", f"% column max={mx:.3g} — can it exceed 100?"))
        if mn < 0 and any(w in unit for w in ("people", "number", "deaths", "tonnes", "count")):
            findings.append((tname, col, "SIGN", f"negative min={mn:.3g} in count-like unit"))
        if not has_country or year_col is None:
            continue  # year-only or static table: only the unit/magnitude sniffs apply; checks 2-6 need country + time
        # 2. Robust per-country outliers (median/MAD z-score; skip short series — MAD is unstable under ~8 points)
        g = s.groupby("country")[col]
        mad = g.transform(lambda x: (x - x.median()).abs().median()).replace(0, np.nan)
        z = ((s[col] - g.transform("median")) / mad).where(g.transform("count") >= 8)
        for _, r in s[z.abs() > 6].head(20).iterrows():
            findings.append((tname, col, "OUTLIER", f"{r['country']} {r[year_col]}: {r[col]:.4g} (|z|>6)"))
        # 3. Trend breaks with a unit-error signature (~×10/×100/×1000 jumps)
        ss = s.sort_values(["country", year_col])
        ratio = ss.groupby("country")[col].pct_change().add(1).abs()
        for _, r in ss[np.log10(ratio.replace(0, np.nan)).abs() >= 1].head(20).iterrows():
            findings.append((tname, col, "BREAK", f"{r['country']} {r[year_col]}: ≥×10 year-over-year jump"))
        # 4. Coverage drop in the most recent period
        cov = s.groupby(year_col)["country"].nunique()
        if len(cov) > 1 and cov.iloc[-1] < 0.7 * cov.iloc[-2]:
            findings.append((tname, col, "COVERAGE", f"{cov.index[-1]}: {cov.iloc[-1]} countries vs {cov.iloc[-2]}"))
        # 5. Suspicious constants: long runs of identical NON-ZERO values (source forward-fill?) —
        #    repeated zeroes are normal for sparse count/event indicators and must not count.
        runs = ss.groupby("country")[col].apply(lambda x: ((x == x.shift()) & (x != 0)).mean())
        for ctry in runs[runs > 0.5].index[:10]:
            findings.append((tname, col, "CONSTANT", f"{ctry}: >50% of series identical to previous period"))
        # 6. World vs sum of countries — ONLY for additive units (counts, tonnes, deaths; never rates/shares/indices)
        #    flag when |World − Σ countries| / World > 5% in a spot-checked year

The scan is a candidate generator, not a verdict. Review the raw findings yourself and discard the obviously legitimate ones (wars, pandemics, currency redenominations, real policy shocks) with a stated reason each before spending any web calls in Phase 2.

Show full SKILL.md (1,549 more words)Show less

Step 4 — Phase 2: independent online cross-check (anomaly-led + anchors)

This is the half that catches the source's mistakes (failure class 2 above).

What to check:

  • Every anomaly that survived your Phase-1 triage. Cap the WebSearch effort at ~10 values; list anything beyond the cap as unchecked in Part 2.
  • Fixed anchors, regardless of anomalies: the World total (if the dataset has one), 2–3 major or topic-relevant countries, the latest year, and one mid-series historical year.
  • One-year shifts in newly added years, for sources that compile national-office figures (OECD, Eurostat, UN agencies). Compare the new years for 1–2 countries with the national office's own figures: our value for year Y matching its figure for Y−1 means the series is shifted.

Independence rules (anti-circularity — read before searching). An independent source is a different producer measuring the same quantity (WHO vs. IHME, IEA vs. Energy Institute, IMF vs. World Bank, UN WPP vs. a national statistics office), or the primary source the producer aggregates. Never count as independent: ourworldindata.org itself; sites that republish OWID (Wikipedia charts and infoboxes frequently cite us — check the citation); mirrors of the same producer (tradingeconomics and friends scrape WB/IMF); or the producer's own secondary pages.

Procedure per value: WebSearch the quantity + entity + year → open 1–2 authoritative hits with WebFetch → record source, value, and link in the Part 2 table. Never cite a number straight from the search-results summary — summaries blend several sources and lag living pages; every figure that reaches a finding, a PR body, or a producer question must be quoted from a page you actually opened.

Measurement-artifact scrutiny (per source, not per value): search "<producer> completeness bias", "<producer> coverage <region>", "<indicator> revision history" and read what comes back. When you flag a comparability problem, name the specific mechanism by which the data misleads (e.g. "death registration completeness below 60% in region X inflates apparent improvement"); a bare "comparisons should be made with care" is banned.

Internal accounting identities beat external sources. When a producer publishes components and their aggregate, check that they reconcile — a contradiction inside the producer's own release is arithmetic, needs no independent source, and can't be waved away as a methodology difference. It also survives the common case where every external source is bot-blocked. This is confirmation route (b) in the Tolerance gate below — it settles a source error on its own, provided the guards hold: every term of the identity comes from the same release and vintage (never mix an old download's share with a new download's rate), the terms are defined so the identity holds by construction (shares that sum to 1, urban+rural weighted by the same population split), and the contradiction is far beyond rounding. Two habits that make this reusable: sweep the identity across all entities, not just the suspicious one, to prove the error is isolated rather than systematic; and check the producer's footnote table (WDIfootnote.csv and equivalents) — an unqualified bad value is a stronger finding than a flagged one. One more guard before anything reaches the corrections route: the contradiction proves an error exists among the identity's terms — it does not by itself say which term is wrong. Identify the bad term with evidence beyond the identity: the entity's own adjacent years (the term solved from the identity matches its own prior-year value, while the other terms sit in line with their own histories), the all-entity sweep isolating a single term, or a footnote. If nothing singles out one term, the finding is still a confirmed producer error — report it and hand it to producer follow-up, but don't guess which value to overwrite in corrections.yml.

A "no charts use this indicator" clearance goes stale the moment charts are remapped. Blast-radius checks are version-scoped; a chart stranded on an older version won't match a query filtered to the version you're updating, and will silently come into scope once the stranded-chart sweep runs. Re-run any such check after the indicator upgrade before relying on it.

A mismatch that matches exactly under swapped names is a label error, and it can be on either side. When a comparison against the producer's own table fails for two related rows (two subregions, two sexes, two variants) and each row's values equal the other row's to the last decimal, don't treat it as our pipeline being wrong. Settle which side has the labels crossed with an independent reconstruction — rebuild both aggregates from their members, or check the populations each label implies against the table's own totals — before touching anything. If it is the producer's document, that is a confirmed producer error to report; if it is ours, it is an entity-mapping bug.

Tolerance: rounding, vintage/revision drift, and methodology gaps of a few percent are not findings. The targets are magnitude errors (×10/×100/×1000), wrong-year values, sign errors, entity mix-ups, and stale pre-revision values. Declaring a confirmed source error requires one of two routes: (a) ≥2 independent sources that agree with each other and disagree with ours beyond methodology tolerance, or (b) an accounting-identity contradiction internal to the producer's own release, under the guards in Internal accounting identities beat external sources above.

Attribution before routing. Before routing any confirmed bad value, read the raw snapshot (from etl.snapshot import Snapshot; Snapshot("<ns>/<version>/<file>").read() — read() picks the reader from the file's format; use the format-specific read_csv/read_excel/read_json only when auto-detection needs overriding) to determine where it entered: present in the source file → source error (corrections route); absent → our processing introduced it (trace snapshot → meadow → garden and fix the step).

Step 5 — Phase 3: adversarial metadata review (top-N + anomalous indicators only)

Treat each prioritized indicator's user-facing text as a set of claims and attack them against the Phase-0 documentation:

  • Does unit (and any display.unit/conversion) match the producer's stated unit? A (mils)/(000) marker in the source's column header demands a visible conversion in garden.
  • Does the title / description_short overclaim scope — "global" when the source covers reporting countries only, "countries" when it's high-income countries?
  • Does description_key state contested definitions as settled, or omit a caveat the producer's own docs (or your Phase-2 literature search) prominently state — coverage gaps, comparability breaks, denominator choices?
  • Are causal or certainty words ("shows", "proves", "leads to", "drives") backed by the source's methodology, or do they smuggle in an interpretation?
  • For categorical indicators built from label maps, list the distinct source labels and verify each maps explicitly — values routed to a fallback bucket ("unknown", "other") are silent misclassifications, because the fallback is an existing category and no validation fires. Recommend an observed labels ⊆ map keys assert where one is missing.
  • For Jinja-templated metadata, spot-check the rendered text readers actually see: Dataset("data/grapher/<ns>/<version>/<short_name>").read(t, load_data=False)[col].metadata.

Lead with the concrete rewrite, not the objection. "Add a link" is a valid fix. Match the register and length of the original — prefer a word swap over an added clause.

Routing findings

FindingAuthor flows (update/create/standalone)Review flow
Metadata contradicts producer docs (unit/definition/scope; content in the wrong field per the .dvc-vs-description_processing split)Edit .meta.yml/.dvc, re-run the step (--grapher for grapher channel)🔴/🟡 with quote + doc link
Value wrong in our output but correct in the raw snapshotFix the step code — never corrections.yml, never mask🔴
Value confirmed wrong at the source (raw snapshot carries it; confirmed via route (a) or route (b) of the Step 4 Tolerance gate — for (b), with the erroneous term identified)Add <short_name>.corrections.yml next to the garden step + tb = paths.apply_corrections(tb) (format: etl/data_corrections.py); fill reason/producer/status, add an expect guard; tell the user to notify the producer and record the reported: date🔴 if confirmed and uncorrected
Suspicious but unconfirmed (independent sources disagree with each other, or methodology plausibly explains the gap)"Verify manually" item in the report — do not add a correction🟡
Docs/data unreachable after curl → WebFetch → Wayback"Unable to verify — worth checking" (proportionality cap)🟢
Producer-doc vs. shipped-file discrepancyPreserve the data as shipped; flag for producer follow-up🟢

Output report

Write to ai/adversarial-review-<short_name>-<YYYY-MM-DD>.md:

# Adversarial data review — <ns>/<version>/<short_name>

## TL;DR
(max 3 sentences: overall verdict + the one thing to fix first)

## Part 0 — Source verification
- Source(s) + documents accessed (URL, access status: read / paywalled / unreachable)
- What the source publishes (indicators, units, definitions, granularity, upstream data)
- What OWID adds (from description_processing + step code + corrections.yml)
- Field-placement check (.dvc description / description_processing / description_from_producer)
- Data-quality flags stated in the source's own docs
- Flags the docs SHOULD state but don't (from the artifact-literature search)
- Existing corrections.yml entries acknowledged

## Part 1 — Findings (numbered, ordered 🔴 → 🟡 → 🟢)
N. 🔴|🟡|🟢 [data-level|text-level] <one-line defect>
   Evidence: <quote or value + link>
   Fix: <concrete action + routing per the table above>
   Why: <one tight sentence>

## Part 2 — Cross-check appendix
- Values checked: indicator | entity | year | our value | independent value(s) | source link | verdict ✓/✗/~
  (mark each row as anchor or anomaly)
- SKIPPED indicators (name, views_365d rank, why skipped) — mandatory, never silent
- Anomalies beyond the WebSearch cap, listed as unchecked

Severity rubric (aligned with /review-data-pr): 🔴 = confirmed factual error (metadata contradicted by the producer's own docs, or a value confirmed wrong via route (a) or (b) in Step 4, with snapshot-level attribution); 🟡 = likely issue needing confirmation; 🟢 = informational or unverifiable.

In author flows, apply the 🔴 fixes immediately (they're why the skill ran before commit); leave 🟡/🟢 as report items for the user to triage.

Key constraints

  • Always fetch and search — never judge data plausibility from memory or training data. The whole point is a fresh look at what the source and the wider literature actually say today.
  • Establish source-vs-OWID attribution (Phase 0) before critiquing either layer; every finding names whose layer it targets.
  • Severity proportional to verification depth — the HARD RULE in Step 2.
  • Don't manufacture objections. A clean report is a valid outcome; if uncertain whether something is wrong, say "verify this" rather than asserting it.
  • Separate data-level from text-level findings — a legitimate metric can sit under overclaiming text, and carefully hedged text can sit on top of broken data.
  • Methodology differences are not errors. Name the specific mechanism before calling a mismatch an error.
  • Factual accuracy only — no style/typo/spacing duplication (see Scope).
  • corrections.yml override values on categorical columns must come from the source's current vocabulary — assigning a retired label fails with Cannot setitem on a Categorical; choose the current-vocabulary value that yields the same published output.
  • No persistent files beyond the ai/ report — plus, in author flows, the metadata/corrections edits themselves. Ad-hoc analysis scripts run from the session and are not committed.

© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fact-check-dataset of owid/etl.

Open the folder on GitHubat commit 70c9705

Compare with similar skills

Fact Check Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fact Check Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fact Check Dataset this skillowid/etl159—~6.9kAutomated safety check: PassMIT
Flowfile Docs ReviewEdwardvaneechoud/Flowfile385—~5.8kAutomated safety check: PassMIT
Gene Infoaipoch/medical-research-skills1.9k—~2.3kAutomated safety check: PassMIT
G1brycewang-stanford/Auto-Empirical-Research-Skills4.6k—~3.7kAutomated safety check: PassCustom licence
Research Workflow Automationwentorai/research-plugins2981 repos~1.9kAutomated safety check: PassMIT
Icde Submissionbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Flowfile Docs Review

    Edwardvaneechoud/Flowfile

    The editorial standard for Flowfile's docs site — the house style (register-2 voice rules), the persona-based nav map (which tab serves which arriving audience), the claim-type→source-of-truth…

    385 GitHub stars~5.8k tokensUpdated today
    Research & ScienceAuto-check passed
  • Gene Info

    aipoch/medical-research-skills

    Retrieves comprehensive gene information including PubMed publication counts, NCBI summaries, and Ensembl transcript data.

    1.9k GitHub stars~2.3k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed
  • G1

    brycewang-stanford/Auto-Empirical-Research-Skills

    VS-Enhanced Journal Matcher with Journal Intelligence MCP — Real-time journal data pipeline with checkpoint-based human decisions.

    4.6k GitHub stars~3.7k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • Research Workflow Automation

    wentorai/research-plugins

    Automate repetitive research tasks with pipelines, schedulers, and scripting

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Research & ScienceAuto-check passed
  • Icde Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when auditing an IEEE ICDE research-track submission for round choice, Microsoft CMT setup, conflict declarations, the IEEE 12-page format, single-blind requirements…

    1.2k GitHub stars~1.1k tokensUpdated 13 days ago
    Research & ScienceAuto-check passed
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    33k GitHub starsUsed in 11 repos~3.5k tokens
    Research & ScienceAuto-check: notes

More from owid/etl

All 35 skills in this repo
  • Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.

    159 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.

    159 GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Add new survey question codes (e.g. An agent skill from owid/etl.

    159 GitHub stars~10k tokensUpdated today
    Auto-check: notes
  • Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…

    159 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.

    159 GitHub stars~9.2k tokensUpdated today
    Auto-check: notes
  • Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.

    159 GitHub stars~7.2k tokensUpdated today
    Auto-check: notes

Questions about Fact Check Dataset

What does Fact Check Dataset do?

Adversarially review an ETL dataset's data and metadata for factual accuracy — verifies metadata claims against the producer's own documentation (fetched from the links in snapshot .dvc files and…. Fact Check Dataset is an agent skill from owid/etl.dvc files and metadata texts) and cross-checks anomalous plus anchor values against independent sources online, to catch unit errors, wrong-year values, and hard-to-detect mistakes made by the source itself.

When should I use Fact Check Dataset?

Fact Check Dataset fits situations like: the user asks to adversarially review a dataset; fact-check this dataset; verify the data against the source; cross-check the values.

How do I install Fact Check Dataset in Claude Code?

Run `npx skills add owid/etl --skill fact-check-dataset -a claude-code`. Or copy the skill folder (.claude/skills/fact-check-dataset in owid/etl) into .claude/skills/fact-check-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Fact Check Dataset in Codex?

Run `npx skills add owid/etl --skill fact-check-dataset -a codex`. Or copy the skill folder (.claude/skills/fact-check-dataset in owid/etl) into .agents/skills/fact-check-dataset in your project. Codex loads it when a task matches its description.

Can I use Fact Check Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill fact-check-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fact-check-dataset, .gemini/skills/fact-check-dataset, .github/skills/fact-check-dataset and .opencode/skills/fact-check-dataset in your project.

What does Fact Check Dataset need to run?

Going by SKILL.md and its folder, Fact Check Dataset needs the command-line tools its instructions call (rg). Our summary lists: Python 3.

Does Fact Check Dataset access the network?

SKILL.md names 1 domain. In commands or code: datasette-public.owid.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Fact Check Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fact Check Dataset use?

Fact Check Dataset is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fact Check Dataset use?

About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fact Check Dataset?

Skills that share tags, products or a category with Fact Check Dataset: Flowfile Docs Review (Edwardvaneechoud/Flowfile, 385 stars), Gene Info (aipoch/medical-research-skills, 1.9k stars), G1 (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars) and Research Workflow Automation (wentorai/research-plugins, 298 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fact Check Dataset?

owid (a GitHub organization) maintains it in owid/etl, which has 159 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 10, 2026.

Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.