Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
$ npx skills add owid/etl --skill map-explorer-to-mdim -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl map-explorer-to-mdim --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/map-explorer-to-mdim .claude/skills/map-explorer-to-mdim && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "map-explorer-to-mdim" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/map-explorer-to-mdim into .claude/skills/map-explorer-to-mdim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "map-explorer-to-mdim", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/map-explorer-to-mdimType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill map-explorer-to-mdim -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl map-explorer-to-mdim --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/map-explorer-to-mdim .agents/skills/map-explorer-to-mdim && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "map-explorer-to-mdim" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/map-explorer-to-mdim into .agents/skills/map-explorer-to-mdim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "map-explorer-to-mdim", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill map-explorer-to-mdim -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl map-explorer-to-mdim --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/map-explorer-to-mdim .cursor/skills/map-explorer-to-mdim && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "map-explorer-to-mdim" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/map-explorer-to-mdim into .cursor/skills/map-explorer-to-mdim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "map-explorer-to-mdim", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/map-explorer-to-mdim--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill map-explorer-to-mdim -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl map-explorer-to-mdim --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/map-explorer-to-mdim .gemini/skills/map-explorer-to-mdim && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "map-explorer-to-mdim" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/map-explorer-to-mdim into .gemini/skills/map-explorer-to-mdim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "map-explorer-to-mdim", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl map-explorer-to-mdimInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill map-explorer-to-mdim -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/map-explorer-to-mdim .github/skills/map-explorer-to-mdim && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "map-explorer-to-mdim" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/map-explorer-to-mdim into .github/skills/map-explorer-to-mdim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "map-explorer-to-mdim", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill map-explorer-to-mdim -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl map-explorer-to-mdim --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/map-explorer-to-mdim .opencode/skills/map-explorer-to-mdim && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "map-explorer-to-mdim" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/map-explorer-to-mdim into .opencode/skills/map-explorer-to-mdim/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "map-explorer-to-mdim", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
map-explorer-to-mdimTake (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Map Explorer To Mdim is an agent skill from owid/etl. Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end. Maps each explorer's views to the views of one or more replacement MDIMs, writes ONE apply-ready JSON payload per explorer for the admin bulk-redirect endpoint, audits every article that links or embeds each explorer (with the view each link will land on) plus the featured metrics pointing at it, preflights every validation the endpoint performs — including site redirects that would block it — and covers retiring the explorer's ETL step…
Its SKILL.md is about 7.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts (for example `scripts/audit_references.py`, `scripts/build_mapping.py` and `scripts/build_review.py`).
It sits in Data & Analytics, covering Data pipelines and ETL. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d5ba5a6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythoncurlmakegitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ourworldindata.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Map Explorer To Mdim loads about 7.3k tokens when it runs. Until then it costs about 228 tokens; SKILL.md has 3,629 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
pplies before running, and don't assume `.env.prod` exists:**what exists (`ls -la .env*`); on some machines it is `.env.prod`, on others `.env.live`.ls -la .env* 2>/dev/null # which credentials files exist?Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from owid/etl at commit d5ba5a6, republished under its MIT licence (© owid). 3,629 words, ~7,322 tokens.
.claude/skills/map-explorer-to-mdim/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.When an explorer is being retired in favour of one or more MDIMs, every explorer view needs a redirect to the equivalent MDIM view. This skill produces the input for that: a CSV of explorer views, a CSV per target MDIM, and a joint proposal mapping each explorer view to a target MDIM view (the suggestion is for human review).
The mapping itself is explorer-specific (how the explorer's dimensions translate to MDIM dimension slugs, and — when there are multiple MDIMs — which MDIM each view routes to). The skill automates everything mechanical (pulling views, the join, the shared-target accounting, validation) and leaves only the per-explorer rules for you to write, seeded with auto-suggested matches.
explorers.slug in the grapher DB (e.g. natural-disasters).multi_dim_data_pages.catalogPath,
e.g. natural_disasters/latest/deaths#deaths. The MDIMs must be published in the
DB you connect to (their fully-expanded views are read from multi_dim_data_pages.config).Both the explorer and the MDIMs are read from the grapher DB via OWID_ENV, so the
scripts only work where that DB actually contains both the explorer and the published
MDIMs. There are three ways to point OWID_ENV at such a DB — figure out which one
applies before running, and don't assume .env.prod exists:
staging-site-<branch> branch,
OWID_ENV already points at that prod-clone DB — run the commands as-is, no prefix.ENV_FILE=<prod creds> DATA_API_ENV=production. Don't assume the file name — check
what exists (ls -la .env*); on some machines it is .env.prod, on others .env.live.ENV_FILE=<their file> [DATA_API_ENV=production].Preflight — check, then ask if needed:
ls -la .env* 2>/dev/null # which credentials files exist?
# Connectivity test (swap the ENV_FILE prefix for whatever applies; drop it on a staging branch):
ENV_FILE=<prod creds> DATA_API_ENV=production .venv/bin/python -c \
"from etl.config import OWID_ENV; print('DB OK:', OWID_ENV.read_sql('SELECT 1 AS x').iloc[0,0])"If no prod credentials file exists and you're not on a staging branch with the data,
stop and ask the user which credentials / env file to use (e.g. "which env file holds DB
credentials that can reach the explorer + MDIMs? Or should I run this from a staging
branch?"). Then use that file as the ENV_FILE= prefix for both
script invocations below. Don't hardcode credentials.
If the connection works but a query returns nothing, the scripts stop with a clear message (explorer slug not found, or MDIM not published in this DB) — that means the DB you reached doesn't have it, so re-check which DB you're pointed at.
.venv/bin/python .claude/skills/map-explorer-to-mdim/scripts/extract_views.py \
--explorer <slug> \
--mdim <ns/v/short#short> [--mdim <ns/v/short#short> ...] \
--out ai/<slug>-mdim-mappingWrites into the out folder:
explorer_views.csv — id (1..N) + dimension_1..M (explorer display values).multidim_<short>_views.csv — one per MDIM; id is letter-prefixed by --mdim order
(A1…, B1…, C1…) so ids are unique across MDIMs; columns are the MDIM dimension slugs._scaffold.md — the explorer dimension legend (which dimension_i is which name),
the distinct values per dimension, each MDIM's dims/choices, auto-suggested value
matches (where a slugified explorer value equals a real MDIM choice slug), and a
ready-to-edit mapping_rules.py template._sources.json — machine-readable record of the explorer slug + dimension names and
each MDIM's short/prefix/catalogPath/dim-slugs. Consumed by build_mapping.py to emit
mapping.json (below); don't hand-edit it.mapping_rules.pyOpen _scaffold.md, then write ai/<slug>-mdim-mapping/mapping_rules.py defining:
EXPLORER_DIMENSIONS — list naming dimension_1..N (copy from the scaffold; keep order).MDIMS — MDIM short names in the same order as --mdim (= prefixes A, B, C, …).route(dims) -> str — given a view's {dimension name: value}, return the target MDIM
short name. For a single MDIM this is just return "<short>". For several, it's a
decision on some explorer dimension (e.g. natural-disasters routes on Impact:
Deaths→deaths, Economic damages (% GDP)→economic_damages, the rest→affected).translate(dims, mdim) -> dict — return {mdim_dim_slug: choice_slug} for the target
MDIM view, built from the *_MAP dicts. Only include slugs the MDIM actually has
(e.g. economic_damages has no metric — single-choice dims are pruned from MDIM views).DEFAULT_MDIM = "<short>" — the catch-all target for the bare explorer URL
(see mapping.json → catchAll). Omit it and the best-fitting MDIM is chosen automatically
(the one receiving the most resolved views; tie-break = earliest in MDIMS). Set it only
when the automatic pick isn't the MDIM you'd want a param-less explorer link to land on.The scaffold seeds the *_MAP dicts with slugify(value) guesses. Verify every entry —
slugify won't catch label↔slug differences like Decadal average→decadal, Injuries→injured,
Volcanoes→volcanic_activity, or aggregate collapses like All disasters/All disasters (by type)→all_stacked.
.venv/bin/python .claude/skills/map-explorer-to-mdim/scripts/build_mapping.py --out ai/<slug>-mdim-mappingWrites mapping_proposal.csv, one row per explorer view:
| columns | meaning |
|---|---|
id, dimension_1..N | the explorer view (same as explorer_views.csv) |
target_mdim, target_view_id | the resolved target (target_view_id is the A*/B*/C* id) |
<mdim>_<dimslug> … | wide block; only the target MDIM's columns are filled with the translated slugs |
shared_target_explorer_ids | when >1 explorer view lands on the same MDIM view, the comma-joined list of all those explorer ids (e.g. 1,12); empty when the target is unique |
It also writes mapping.json — the machine record — and
admin_bulk_payload.json, which is what you actually apply. The API exists:
POST {admin}/api/multi-dim-redirects/bulk (handleBulkCreateMultiDimRedirects), also
reachable from the Bulk-create redirects from JSON button on /admin/multi-dim-redirects.
This endpoint is the only way to apply an explorer redirect: the CSV CLI that
map-charts-to-mdim hands over (createMultiDimRedirectsFromCsv) accepts an
/explorers/ source, but it writes no sourceQueryParams — so the row bakes as one
unconditional rule and every view of the explorer 302s to a single MDIM view,
per-view routing gone.
Its Zod schema deliberately mirrors this file's catchAll + redirects shape, ignores keys
it doesn't know (sourceViewId, viewId, mdim, stats, targets), and reports
target: null entries as skipped.
Post
admin_bulk_payload.json, nevermapping.json. The payload has empty-valued source dimensions stripped out. That is mandatory, not cosmetic: a condition is matched against the incoming URL's params, and an absent param is not an empty string — so a condition of{"Period": ""}can never match, and every view carrying one silently falls through to the catch-all instead of its intended target.mapping.jsonkeeps them because it is the faithful record of the view grid.
Unlike the CSV (positional dimension_N columns, meant for a spreadsheet), the JSON carries
every identifier a redirect needs:
{
"explorer": { "slug": "...", "dimensions": ["<name>", ...] },
"targets": [ { "mdim": "...", "catalogPath": "ns/v/short#short", "dimensions": ["<slug>", ...] } ],
"stats": { "total": N, "resolved": N, "unresolved": N },
"catchAll": { // bare explorer URL (no query params) fallback
"source": { "explorerSlug": "..." },
"target": { "mdim": "...", "catalogPath": "ns/v/short#short",
"viewId": null, "dimensions": {} } // no params → the MDIM's default view
},
"redirects": [
{
"sourceViewId": 1,
"source": { "explorerSlug": "...", "dimensions": { "<name>": "<value>", ... } },
"target": { // null when unresolved
"mdim": "...", "catalogPath": "ns/v/short#short",
"viewId": "A2", // internal id, cross-references the CSVs
"dimensions": { "<slug>": "<choiceSlug>", ... }
},
"sharedTargetSourceIds": [1, 29, 57], // present only when >1 source shares this target
"unresolvedReason": "..." // present only when target is null
}
]
}The source view is identified by the explorer slug + dimension name→display-value
(the explorer URL query params); the target view by the MDIM catalogPath + dimension
slug→choice-slug (the MDIM URL query params). Unresolved views are kept with
target: null so the API can see the full picture; a consumer typically skips them.
An unresolved view is not a view left alone. The endpoint reports
target: nullentries asskipped, so they get no rule of their own — and the catch-all constrains no params, so it matches them instead. Those URLs land on the target MDIM's default view, carrying the explorer's own params into the grapher URL, rather than on anything equivalent to the view that was asked for. Leaving views unresolved is a deliberate call (they may genuinely have no MDIM counterpart), never a no-op; preflight reports the count so it gets made on purpose.
catchAll is always present: it redirects the bare explorer URL (no query params) — and
serves as the sensible fallback for any view a consumer doesn't route individually — to the
best-fitting MDIM with no query params, which grapher renders as that MDIM's default view
(hence viewId: null, dimensions: {}). The best-fitting MDIM is the one most views resolve
to, or whatever DEFAULT_MDIM in mapping_rules.py overrides it to.
The catch-all's destination is the only one nothing pins. Its row stores
viewConfigId = NULL, so it resolves at request time to whichever view the MDIM renders for an empty selection — the first view in config order (grapher'sfilterToAvailableChoicestakes the first available view's choice at every dimension). Rebuilding the MDIM can reorder its views, or edit that view's config in place, and the bare explorer URL then lands somewhere nobody reviewed. Extraction records that view in_sources.json(not in the payload, which must stay byte-reproducible) and preflight warns when it moves. A warning, not a blocker: the destination is a moving reference by construction, so blocking would flip an approved catch-all to NOT READY on any unrelated MDIM rebuild. To pin it, give the catch-all an explicit target view instead.
The script prints a validation report: how many explorer views resolved, distinct MDIM views hit per MDIM, how many rows share a target, and FLAGS for any explorer view that didn't resolve to a real MDIM view (fix the rules and re-run until there are no flags).
Sanity-check the flagged rows and the judgment calls (approximate type matches, aggregate collapses, MDIM choices with no explorer source). For a topic owner's sign-off, build the side-by-side review page: each explorer view on the left, the MDIM view it will redirect to on the right, with approve/flag controls.
.venv/bin/python .claude/skills/map-explorer-to-mdim/scripts/build_review.py \
--mapping-dir ai/<slug>-mdim-mapping \
--explorer-slug <slug> \
--mdim-slug <short_1>=<grapher_slug_1> \
--mdim-slug <short_2>=<grapher_slug_2> \
[--host https://ourworldindata.org] \
[--output ai/<slug>_view_review.html] \
[--no-coverage]--mdim-slug maps each MDIM short name in mapping_rules.MDIMS to its published Grapher slug
(the /grapher/<slug> part of the URL, not the catalogPath). _sources.json records the slugs
extraction saw, and multi_dim_data_pages.slug has them if you have DB access; otherwise ask.
The default host is production; for MDIMs that only exist on a staging branch pass
--host https://staging-site-<branch>.
The script prints a coverage summary before writing the HTML: rows, distinct MDIM targets, many-to-one collapses, unresolved rows, MDIM views never targeted. Read it; it is the fastest way to spot a mapping gap the reviewer cannot see by eye.
What the reviewer gets: one pair at a time in iframes, the selection shown as chips above each
chart, Approve / Flag / Clear with an optional note, keyboard navigation (← →, a, f, c),
filters and live counts. Decisions auto-save to the browser's localStorage, can be mirrored to
a JSON file on disk (Chrome/Edge), and Import merges another reviewer's export. The file is
self-contained: send it directly, or publish it to vibe.owid.io with
/owid-staff:create-vibe-app (from the owid-staff plugin, auto-installed here) when the whole
team needs a link.
When the reviewer is done, ask for the exported JSON (or the auto-saved one): every row carries
status (approved / flagged / blank) and note. Approved rows are ready to wire up; flagged
rows need a second pass with the user. Corrections go through mapping_rules.py and a rebuild
(step 3), never through the HTML.
Gotchas:
localStorage is per browser and per file path, so switching machine loses the decisions
unless they were exported or mirrored to disk.mapping_proposal.csv; do not add them back by hand.Disaster Type=Floods), which is what the
dimension_1..N columns hold. hideControls=true is appended to both sides; "open ↗" shows
the view with its controls.mapping_proposal.csv plus mapping_rules.py first; a one-off adapter in ai/
is fine.ENV_FILE=<prod creds> DATA_API_ENV=production .venv/bin/python \
.claude/skills/map-explorer-to-mdim/scripts/audit_references.py \
--mapping ai/<a>-mdim-mapping --mapping ai/<b>-mdim-mapping \
--out ai/<combined>-redirectsWrites references.csv + references.md across all the explorers in one pass. It resolves
each referencing URL through the rules the payload would create, so every row says which view
the reader lands on — and separates the link had no params from the link names a choice the
explorer has since dropped, which lands on the MDIM's default view and needs authoring
attention rather than a URL swap.
The timeline differs from the chart skill's, and this is the thing to say out loud: an embedded explorer breaks the moment the redirect is created, not later at unpublish, because the embed renders by fetching the explorer page and parsing it. So the 🔴 rows are migrated before step 7, not after.
The report also carries a ⭐ Featured metrics section — the one surface where this skill's usual "a link survives the 302" reasoning inverts. A featured metric is a topic-page slot held by URL, resolved only when Algolia indexes, matching pathname and exact params against published records. So it does not survive: it empties silently, and cannot be re-added once the explorer is gone. Hence step 5b.
Work the ⭐ section of references.md, by hand at /admin/featured-metrics. These are editorial
slots, so ask whoever owns the topic before repointing their rail.
Per row — add the MDIM view under the same tag and income group, drag it to the old ranking,
delete the old row, then re-apply boost in search if it was on. One explorer view can hold
several rows: the key is (URL, tag, income group). Full procedure:
docs/guides/data-work/redirect-to-mdims.md.
Why here and not after step 7: creating the redirect darkens the explorer on the spot, and adding a featured metric requires a published slug. Once the redirect exists, no replacement is accepted and no record of the ranking survives. Unlike an embed, which breaks visibly, this fails quietly. The replacement is the bare MDIM view, not the redirect target — the admin strips reader params on paste and never validates the dimension params.
ENV_FILE=<prod creds> DATA_API_ENV=production .venv/bin/python \
.claude/skills/map-explorer-to-mdim/scripts/preflight.py \
--mapping ai/<a>-mdim-mapping --mapping ai/<b>-mdim-mapping \
--out ai/<combined>-redirects [--record ai/<combined>-redirects/bulk_redirects.json]Mirrors every validation the endpoint performs, per explorer — because it memoizes its
source-side checks and re-throws the cached rejection, so one source-level problem fails
every entry for that explorer. Blockers: the explorer path is already a site-redirect
source, or the target of one (the chain case); the slug collides with a chart's old slug; a
target MDIM is missing or unpublished; the target /grapher/<slug> is itself a redirect
source; two views share a source condition; the view fingerprint no longer matches the live
explorer; the payload was not built from the extraction sitting beside it; redirects already
exist that differ from the payload; or embedded references are outstanding. A missing
references.csv is a blocker too — "never looked" must not read like "looked and found
nothing".
Two of those checks exist because nothing else ties the payload to the run that produced it,
and an aborted or skipped rebuild leaves a stale one in place. build_mapping.py writes
mapping_proposal.csv and mapping.json before admin_bulk_payload.json, and aborts between
them on duplicate conditions — so the artifacts a human reviews can describe one build while the
payload about to be posted comes from another, with the fingerprint check none the wiser (it
validates the extraction, not the rules). Preflight therefore checks both ends:
explorer_views.csv + _sources.json, which
catches an extraction re-run with no rebuild at all;mapping.json beside it, which is
what catches a rebuild that aborted in between — including one where only the targets
changed, invisible to any source-side check.The reference gate is bound to live state the same way. audit_references.py records a
referenceDigests entry per explorer in references_manifest.json; preflight re-runs the sweep
and compares. Without that, the gate reads a CSV from an earlier run and cannot see a page that
added an embed since, or an audit folder carried over from another migration — both of which
would read as a clean audit while the redirect is about to break something live.
It also records mappingDigests, binding the audit to the mapping its advice came from.
Every replacement URL in references.md is derived from the mapping, so rebuilding the mapping
with different targets invalidates that advice while the explorer's own references — and so the
reference digest — stay identical. This is the one staleness with no second chance: an operator
who repoints an embed at the wrong view leaves it no longer naming the explorer, so no later
sweep can ever surface the mistake. Preflight blocks on it before the reference gate reports.
!!! note "An audit folder predating a surface reads as drifted, and should"
The digest hashes the findings, so adding a surface to find-chart-references invalidates
every recorded referenceDigests — most recently when featured metrics were added.
Preflight blocks on an audit folder from before that, which is correct: it really is missing
rows. Re-run audit_references.py rather than reading it as a bug.
Unverifiable is a blocker, not a warning, in all three cases — no extraction pair, no
mapping.json, no reference digest. A warning does not reach the exit code, so Ready would
print over a report stating in plain words that the payload's provenance is unknown. Re-running
extract_views.py / build_mapping.py / audit_references.py is the cheap way to clear any of
them; posting an artifact nothing backs is not.
A retired explorer whose redirects are all live reports DONE and exits 0. That is the
finished state, not an error. But retirement does not retire the target side: an MDIM that
has since been unpublished, rebuilt, re-slugged or edited in place breaks redirects that are
already in the DB, and no explorer row is left to notice it. So every target check runs for a
retired explorer too, and a slug carrying any blocker is never reported DONE.
Non-zero exit means do not post anything.
[!WARNING] Creating the redirect darkens the explorer immediately. It is checked on every
/explorers/*request, ahead ofenv.ASSETS.fetch, so it beats the baked explorer page and any_redirectsentry, and fires while the explorer is still published. There is no staged rollout and no bulk undo — removal is one row at a time.
Paste each admin_bulk_payload.json into Bulk-create redirects from JSON at
{admin}/multi-dim-redirects, one explorer at a time. Read the response
positionally: every results[i].source is the same /explorers/<slug> string, so index 0
is the catch-all when present and index i is redirects[i-1]. Expect
created + skipped + errors == entries.
Then wait for the bake plus ~2 minutes (the redirect map is fetched with a 2-minute edge TTL) and verify:
curl -sI "https://ourworldindata.org/explorers/<slug>?<one view's params>" | grep -i "^HTTP/\|^location:"
# expect 302 + location: /grapher/<mdim>?<view dims>The explorer is unpublished or deleted BY HAND in the admin. Removing the ETL step does
not unpublish it — the explorers row survives, so it keeps showing up in listings and
search. Do not flip isPublished in the step's .config.yml and re-run hoping to achieve it;
that is not the route (and no explorer config in the repo carries isPublished: false).
Then remove the explorer's ETL footprint — and never delete a step without archiving it
(see CLAUDE.md). Delete the step's .py and its sibling .config.yml (the periodic archive sweep only knows .py, which is why orphaned configs
exist on disk today), remove the dag/*.yml entry, make check, commit; then run
.venv/bin/etl archive-dag and commit dag/archive/*.yml separately, since it reads committed
history. git checkout anything unrelated it sweeps in. Archive anything now orphaned
upstream — a garden step that existed only to feed this explorer — in a second round. If any of
those is a migrated/backport dataset, delete its now-orphaned
snapshots/backport/latest/dataset_<id>_* mirror files too: archiving the DAG entry leaves them
on disk, and nothing will point at them again.
Grep before deleting a shared explorer data step. Some are consumed off-DAG, by scripts fetching their published catalog CSVs by URL. Those consumers are invisible to
.venv/bin/etl archive-dag, so the step looks like a safe leaf when it is not. Search for its catalog URL first and keep it until every consumer is retired.
Track these steps with TodoWrite in the chat. Do not generate a checklist file, and
there is no HANDOFF.md for this skill — unlike the chart path there is no cross-team
handoff, since the same operator pastes the payload.
| file | written by | contents |
|---|---|---|
explorer_views.csv | extract | id 1..N + dimension_1..M (display values) |
multidim_<short>_views.csv | extract | one per MDIM, ids A1…/B1… |
_scaffold.md | extract | dimension legend, distinct values, auto-matches, mapping_rules.py template |
_sources.json | extract | slugs, catalogPaths, MDIM ids/slugs/published, viewsFingerprint, configMd5 — don't hand-edit |
mapping_rules.py | you | routing + value translation |
mapping_proposal.csv | build | one row per explorer view, wide target block |
mapping.json | build | faithful machine record (empty source dims kept) |
admin_bulk_payload.json | build | the apply unit — one per explorer, paste into the admin modal |
references.csv / references.md | audit | combined across explorers, in --out |
bulk_redirects.json | preflight --record | combined record; not postable |
ai/<slug>_view_review.html | review | self-contained side-by-side HTML with approve/flag controls, for the topic owner |
multi_dim_data_pages.config (published, fully expanded). This
already reflects code-generated views (group_views aggregates) and pruned single-choice
dimensions — so e.g. a metric that has only one active choice won't appear as a column.all_stacked MDIM view). shared_target_explorer_ids
surfaces these so the reviewer sees the collisions.…_excluding_extreme_temperature
aggregate). Nothing redirects to those — fine, just confirm.dimension_1..N (compact, and joinable to the explorer CSV);
the name legend lives in _scaffold.md and in EXPLORER_DIMENSIONS.extract_views.py overwrites the CSVs and _sources.json but not your mapping_rules.py.mapping.json needs _sources.json; if you extracted before this output existed, just
re-run extract_views.py once (it preserves mapping_rules.py), then build_mapping.py.catchAll, so a merged file would silently drop all but one.DELETE /api/multi-dims/:id/redirects/:redirectId, one
row at a time. A 460-row batch posted wrongly is expensive to undo, which is why the
preflight exists._redirects (getRecentMultiDimRedirects excludes
non-/grapher/ sources), so there is no static-redirect fallback and no one-week window —
they live only in the baked redirect map.country=, time=, tab= ride through untouched, which is
what a reader following an old link wants.mapping_proposal.csv, the review HTML and sourceViewId key on. That is what
viewsFingerprint detects; configMd5 is only a secondary signal, since it flips on any
FAUST edit and gating on it would force a needless re-review.linkedChart unresolvable, so Chart blocks render nothing and tiles vanish. The five
WB/WID inequality explorers migrated in July are the worked example.© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts) in .claude/skills/map-explorer-to-mdim of owid/etl.
Open the folder on GitHubat commit d5ba5a6
Map Explorer To Mdim next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Map Explorer To Mdim this skillowid/etl | 159 | — | ~7.3k | Automated safety check: Notes | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 599 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Glue 09 10 Migrationaws-samples/aws-glue-samples | 1.5k | — | ~2.4k | Automated safety check: Pass | MIT-0 | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Dbt Databricks PR Readydatabricks/dbt-databricks | 380 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Apache Spark EngineerJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
databricks/dbt-databricks
A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
MaterializeInc/materialize
Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Check chart or multidim preview on the staging server using a browser.
Categories
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end. Map Explorer To Mdim is an agent skill from owid/etl. Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Map Explorer To Mdim fits situations like: the user says map explorer <slug to mdim(s) <..; suggest explorer-MDIM redirects; were sunsetting the <slug explorer; map its views to the new multidims.
Run `npx skills add owid/etl --skill map-explorer-to-mdim -a claude-code`. Or copy the skill folder (.claude/skills/map-explorer-to-mdim in owid/etl) into .claude/skills/map-explorer-to-mdim in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill map-explorer-to-mdim -a codex`. Or copy the skill folder (.claude/skills/map-explorer-to-mdim in owid/etl) into .agents/skills/map-explorer-to-mdim in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill map-explorer-to-mdim -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/map-explorer-to-mdim, .gemini/skills/map-explorer-to-mdim, .github/skills/map-explorer-to-mdim and .opencode/skills/map-explorer-to-mdim in your project.
Going by SKILL.md and its folder, Map Explorer To Mdim needs Python for the scripts in its folder and the command-line tools its instructions call (python, curl, make and git). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: ourworldindata.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Map Explorer To Mdim is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.3k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Map Explorer To Mdim: Crawl4AI Web Scraping (smallnest/goclaw, 599 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 159 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.