Scikit Learn
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Create a new snapshot (DVC file, plus a Python script only when one is needed) from a urlmain and optional urldownload.
$ npx skills add owid/etl --skill create-snapshot -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl create-snapshot --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/create-snapshot .claude/skills/create-snapshot && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "create-snapshot" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-snapshot into .claude/skills/create-snapshot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-snapshot", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/create-snapshotType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill create-snapshot -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl create-snapshot --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/create-snapshot .agents/skills/create-snapshot && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "create-snapshot" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-snapshot into .agents/skills/create-snapshot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-snapshot", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-snapshot -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl create-snapshot --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/create-snapshot .cursor/skills/create-snapshot && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "create-snapshot" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-snapshot into .cursor/skills/create-snapshot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-snapshot", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/create-snapshot--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill create-snapshot -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl create-snapshot --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/create-snapshot .gemini/skills/create-snapshot && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "create-snapshot" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-snapshot into .gemini/skills/create-snapshot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-snapshot", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl create-snapshotInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill create-snapshot -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/create-snapshot .github/skills/create-snapshot && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "create-snapshot" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-snapshot into .github/skills/create-snapshot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-snapshot", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-snapshot -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl create-snapshot --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/create-snapshot .opencode/skills/create-snapshot && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "create-snapshot" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-snapshot into .opencode/skills/create-snapshot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-snapshot", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
create-snapshotCreate a new snapshot (DVC file, plus a Python script only when one is needed) from a urlmain and optional urldownload.
Create Snapshot is an agent skill from owid/etl. Create a new snapshot (DVC file, plus a Python script only when one is needed) from a urlmain and optional urldownload. Fetches the page, extracts metadata with AI, confirms with user, writes files, and runs the snapshot. Use when the user wants to add a new data source or create a snapshot from a URL.
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics. It works with Python. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 69ab20e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ourworldindata.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Create Snapshot loads about 5.9k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 3,018 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from owid/etl at commit 69ab20e, republished under its MIT licence (© owid). 3,018 words, ~5,935 tokens.
.claude/skills/create-snapshot/SKILL.md (or your agent's skills folder).Create a new ETL snapshot from a source URL: fetch the page, infer metadata, confirm with the user, write the .dvc file (plus a .py script only when one is genuinely needed), then run the snapshot.
Paired skill — keep in sync.
/create-datasetconsumes the conventions defined here (its Step 4 builds the snapshot for a full dataset chain, reusing this skill's step 3 generator call withdataset_manual_import: True): whenever you change the.dvcfield guidance, the cookiecutter context, or the workflow in this file, check whethercreate-dataset/SKILL.mdneeds a matching edit (and make it in the same commit if so). The reverse also holds — see the mirror note there. The update-side skills are part of the same family: the fields written here are exactly what/update-dataset§6c re-checks on every version bump and what/review-data-pr§5 compares old-vs-new at review time — keep the field guidance consistent across all of them.
Required:
url_main — the dataset landing page URLOptional:
url_download — direct download URL for the data file (if available)Use WebFetch to fetch url_main. From the page content, extract as much metadata as possible:
| Field | Where to look |
|---|---|
title | Page <title>, main heading, dataset title |
description | Page description, abstract, or "about" section. Per the metadata reference: if the producer provides a good, factual description, use it exactly or conveniently rephrased. Academic abstracts and methodology summaries usually qualify — take them as-is. Never paste promotional or first-person copy ("our nation", "vital information", funding pitches): rewrite it as a neutral, factual description from OWID's point of view (who runs it, what it measures, coverage, cadence) |
producer | Organisation name, data owner, author |
citation_full | The producer's recommended citation, copied verbatim (light copyedits only — fixing a typo, stray spacing, or encoding artifacts; never rephrasing) — look for "cite as" / "suggested citation" / "please make the following reference" blocks on the page, an "Original citation" on repository pages (e.g. WRAP), NBER's suggested citation, or "Reference:" headers inside the data files themselves. Sites often have a dedicated "how to cite" / citation page separate from the dataset landing page — browse the site's navigation for one. Only construct a citation in standard format when the producer provides none |
attribution_short | Short org name / acronym |
date_published | Publication date or last-updated date. Not on the landing page? Look on other pages of the site (news/release notes, documentation), in the paper itself (title page and abstract often carry the exact date, e.g. "27 September 2013"), or in the files (file names like 10sd_jan15_2014.xlsx or shna2025tablesiii.xlsx encode the release, and notes/readme sheets or PDF metadata may state it). On a fully JS-rendered page that states no date, the download URL's HTTP Last-Modified header is a defensible source — corroborate it against a release-named filename (see /update-dataset §6c). Use a partial date ("YYYY" or "YYYY-MM") if that's all the source supports; never silently default it to date_accessed — ask the user if still unsure, and only if they confirm the producer publishes no date anywhere, fall back to date_accessed with a # comment in the .dvc documenting why (/review-data-pr §5 looks for exactly that comment) |
license_name | License section (e.g. "CC BY 4.0", "Open Government Licence"). Not on the landing page? Also check the documentation (sources & methods documents, notes/readme sheets inside the data files, repository cover sheets like WRAP) and other pages within the same website — producers often have a dedicated licensing / terms-of-use / "about the data" section reachable from the site navigation. If no license is stated anywhere, warn the user explicitly and fall back to rights-reserved © <producer> (<year>) — never invent permissive terms like "Free to use" |
license_url | Link to the license, or to the page/document where the terms are stated |
file_extension | Infer from url_download if provided (csv, xlsx, xls, zip, json…); default csv |
Leave fields blank if they cannot be inferred — the user will fill them in.
Present the inferred metadata and ask the user to fill in or correct:
Required fields the user must provide:
namespace — e.g. who, worldbank, un_igme (suggest based on producer)short_name — snake_case file stem, e.g. child_mortality_ratesversion — YYYY-MM-DD (default: today's date from date -u +"%Y-%m-%d")Pre-filled fields to confirm or correct:
title — dataset titleproducer — organisation namecitation_full — full citation stringattribution_short — short name / acronym (optional)date_published — YYYY-MM-DD or YYYY or YYYY-MMdescription — brief dataset description (optional)license_name — e.g. CC BY 4.0license_url — license page URL (optional)file_extension — inferred from download URL or csvis_private — default falsedataset_manual_import — default false (set to true if there's no url_download)Present this as a summary block so the user can quickly scan and correct individual fields. Wait for confirmation before proceeding.
Once the user confirms, generate both files from the wizard's snapshot cookiecutter — the same templates the wizard's snapshot page uses.
Never hand-write these files, and never copy the templates into this file.
apps/wizard/etl_steps/cookiecutter/snapshot/is the single source of truth. Hand-copied templates drift: an earlier version of this skill carried its own copies, and the manual-import one had already lost the canonical docstring. The template also gets details right that are easy to fluff by hand — most importantlylicensenested insideorigin(CLAUDE.md's most-repeated snapshot mistake, and a schemanot-constraint plustest_snapshot_license_lives_under_originexist because of it), andis_public: falsefor a private snapshot.
Files produced:
snapshots/<namespace>/<version>/<short_name>.<file_extension>.dvc
snapshots/<namespace>/<version>/<short_name>.py # removed again when no script is needed — see belowDecide whether a .py script is needed (CLAUDE.md, "No .py for simple downloads"):
url_download, no custom logic → set dvc_only: True, which deletes the generated script so only the .dvc remains. etls <namespace>/<version>/<short_name> runs it straight from the .dvc. This is the default case — /review-data-pr §3/§7 treat the script as optional, so don't keep one "for completeness".dataset_manual_import: True; the template emits the path_to_file variant of run().run(), between paths.init_snapshot() and snap.create_snapshot(...).Either way the script is a plain run() — no click decorators and no if __name__ == "__main__": block. The etls CLI imports the module and invokes run itself, supplying --path-to-file for manual imports. The template already gets this right; the warning matters because many old scripts in the repo still carry the boilerplate — don't copy one of those as a model.
Generate the files. Every field confirmed in step 2 goes in as cookiecutter context. Pass all the keys below — there is no committed cookiecutter.json supplying defaults, so a missing key is a Jinja UndefinedError, and an empty string is how you say "omit this field" (the template's {%- if %} guards drop it from the output).
Write the context as JSON with the Write tool, then run the generator against that file. Do not interpolate the field values into a python -c program: citation_full, title and description are producer prose, and an apostrophe (World Bank's), a double quote, a backslash, or a newline in any of them either breaks the program or silently changes the value before cookiecutter sees it. Producer citations contain apostrophes routinely, so this is the normal case, not an edge case. JSON keeps the prose out of shell and Python literals entirely.
First write /tmp/snapshot_context.json (booleans are real JSON true/false, and "" means "omit this field"):
{
"channel": "snapshots",
"namespace": "<namespace>",
"snapshot_version": "<version>",
"short_name": "<short_name>",
"file_extension": "<file_extension>",
"is_private": false,
"dataset_manual_import": false,
"dvc_only": false,
"title": "<title>",
"description": "<description>",
"title_snapshot": "",
"description_snapshot": "",
"origin_version": "",
"date_published": "<date_published>",
"producer": "<producer>",
"citation_full": "<citation_full>",
"attribution": "",
"attribution_short": "<attribution_short>",
"url_main": "<url_main>",
"url_download": "<url_download>",
"date_accessed": "<version>",
"license_name": "<license_name>",
"license_url": "<license_url>"
}Then generate. The only string literal in this program is a fixed path, so no field value can break it:
.venv/bin/python -c "
import json
from apps.utils.files import generate_step
from apps.wizard.etl_steps.utils import COOKIE_SNAPSHOT
from etl.paths import SNAPSHOTS_DIR
with open('/tmp/snapshot_context.json') as f:
data = json.load(f)
generate_step(cookiecutter_path=COOKIE_SNAPSHOT, data=data, target_dir=SNAPSHOTS_DIR)
# dvc_only: drop the script the template always writes.
if data['dvc_only']:
py = SNAPSHOTS_DIR / data['namespace'] / data['snapshot_version'] / (data['short_name'] + '.py')
py.unlink(missing_ok=True)
"Note that license belongs to origin — the template nests license_name / license_url correctly, so there is nothing to move afterwards.
Which fields to fill, and when to leave them empty:
| Context key | Value |
|---|---|
title / description | The data product. description is factual — the producer's own text when it is factual, never promotional copy |
title_snapshot / description_snapshot | Both empty by default. Set them only when the file is one table/extract of a broader product; description_snapshot then becomes required, and carries the file specifics (table number, variables, units, years) plus any OWID-side context such as manual transcription or an archived copy |
attribution | Empty unless producer (year) is genuinely uninformative |
attribution_short / origin_version | Empty when the producer gives none |
url_download | Empty for a manual import |
license_url | Empty when the producer states no license anywhere — don't fall back to the landing page |
date_accessed | The snapshot version date |
After generating, three things to do:
Verify the .dvc parses. The template escapes the single-line fields it quotes (title, producer, title_snapshot, attribution, attribution_short, version_producer, license_name) and uses block scalars for the multi-line prose (description, description_snapshot, citation_full), so ordinary producer text — quotes, colons, #, backslashes — round-trips correctly. This check is therefore a cheap guard, not a workaround for a known gap: if it ever fails, the bug is in the template's escaping and belongs upstream in apps/wizard/etl_steps/cookiecutter/snapshot/, not in a hand-fix to the generated file.
.venv/bin/python -c "
import sys, yaml
p = 'snapshots/<namespace>/<version>/<short_name>.<file_extension>.dvc'
try:
yaml.safe_load(open(p))
except yaml.YAMLError as e:
sys.exit(f'{p} is not valid YAML — re-quote the offending field:\n{e}')
print(f'{p}: valid YAML')
"Tidy the end of the .dvc. The template's final {%- endif -%} leaves a whitespace-only line (" ") after the license block, and for a private snapshot the file also ends without a final newline. Both parse fine as YAML, but committed files shouldn't carry either: drop the whitespace-only line and make sure the file ends with exactly one \n.
Add any # comments this skill calls for: the companion-files # NOTE: above url_download (step 6), and a one-line note when citation_full's year deliberately differs from date_published (step 5).
There is deliberately no outs: block — snap.create_snapshot() writes it with the real md5 and size when step 4 runs. Don't add a placeholder.
After writing the files, run:
.venv/bin/etls <namespace>/<version>/<short_name>dataset_manual_import is true, tell the user to download the file manually and re-run with --path-to-file <path>.file_extension — check what the download URL actually servesurl_download — verify with the userdataset_manual_import = trueLinks: run the HEAD-check loop from /update-dataset §6c ("Link verification") on every URL in the new .dvc (url_main, url_download, license.url, and any URL inside description). A curl non-2xx is a signal, not proof — Cloudflare-fronted hosts return false 404s to curl. Escalate with WebFetch, then the Wayback availability API, but per §6c no automated signal is decisive (a missing Wayback capture is non-evidence, and bot-blocked hosts fail curl and WebFetch while serving browsers fine): a link that fails every automated check gets reported with its evidence trail for the user to confirm in a browser — never mark it broken or swap it for an alternative on automated failures alone. URLs carrying a #fragment also need §6c's anchor pass — HTTP status alone can't validate a fragment.
Citation year vs date_published year: the year inside citation_full should normally match date_published's year. A deliberate mismatch is fine when the producer labels the release by edition rather than publish date (e.g. a "2025 report" published 2026-03-17) — leave a one-line # comment in the .dvc so the next reviewer doesn't re-flag it (/review-data-pr §5 checks exactly this pair).
Typos in the .dvc: run /check-metadata-typos on the new file, passing the .dvc path itself (or the extensionless snapshots/<namespace>/<version>/<short_name> stem) to its "current step only" scope, which normalizes both to the .dvc. Confirm it reports a target count of 1 before trusting the result — a scope that matched no file reports no typos. The prose fields written in step 3 — description, title, citation_full, attribution — are user-facing, and nothing downstream spell-checks them: /update-dataset §6c only re-checks them on a later version bump. citation_full is the exception to fixing what it reports: it is verbatim producer text, so only correct a typo there if the producer's own page has it right (see the citation_full note above).
Outdated practices in the .py, when step 3 wrote one: run /check-outdated-practices on it. Run the skill rather than eyeballing the file; it reads the detector extension as its source of truth and catches helper calls that look current but aren't. /review-data-pr treats a leftover __main__ block in a snapshot as a 🔴 blocker, so this is cheaper to fix here than at review.
Know what it does and does not cover, and don't read a clean result as more than it is. For snapshot paths the detector's main pattern is the if __name__ == "__main__": guard, plus the metadata-preserving patterns, which only bite when the script parses before storing. Don't assume it covers every convention in step 3 — check the pattern list in vscode_extensions/detect-outdated-practices/src/extension.ts rather than inferring coverage from a clean run, and confirm the @click convention yourself with the grep below (cheap, and correct whether or not the installed build carries a click pattern).
grep -n "@click" snapshots/<namespace>/<version>/<short_name>.py # expect no outputRoute the finding by where the pattern came from. The script templates in step 3 are variants of apps/wizard/etl_steps/cookiecutter/snapshot/, which is where the current practices are supposed to live — so a hit on a line that came from the template means the template is stale, and fixing only the generated file leaves every future snapshot carrying it. Fix it upstream in the cookiecutter (and in this skill's copy of it) as well as in the file you just wrote.
Optional deeper pass — adversarial source verification: /fact-check-dataset goes beyond "the links resolve" and reads the producer's documentation behind them, checking every .dvc claim (description accuracy, counts, date_published, license, citation) against what the docs actually say — its Phase 0 is the slice that applies at snapshot stage (the data cross-checks need a built garden dataset, e.g. via /create-dataset). Don't run it by default — it fetches docs and runs web searches, so it can consume many tokens; offer it when the source looks unreliable (no version labels, self-published, or the page and file seem to disagree).
Show:
.dvc, plus the .py if one was needed).dvc, not just in chat — a # NOTE: comment above url_download listing the release's other data files as of date_accessed (e.g. # NOTE: the release also ships Foo_Index.csv and codebook.pdf — not snapshotted (2026-07-20)). That comment is the baseline the next /update-dataset cycle diffs against when the host has no file-history API; a new companion file is invisible to every within-file check (see /update-dataset, "Surface new indicators").<namespace>/<version>/<short_name>"date_accessed in the DVC file should always equal the snapshot version date (the date you ran the snapshot).url_download is not provided and cannot be inferred, always set dataset_manual_import = true.<noscript> table — parse that instead: whole-page pd.read_html(io.StringIO(resp.text)), select the table whose columns exactly match the expected header, assert exactly one match. See /update-dataset Guardrails, "Scraped chart embeds".url_download in the .dvc and pass user_agent="owid-etl/1.0 (https://ourworldindata.org)" to snap.create_snapshot(...) — and keep the .py in that case (the script-less path can't set a UA), saying so in the docstring. Also re-check the producer's download page/API for a stable direct endpoint before carrying a manual flow forward. See /update-dataset Guardrails.outs block md5 and size fields are filled in automatically by DVC when the snapshot runs — just set them to empty/zero in the template.<TO BE FILLED> and ask the user.© <producer> (<year>).citation_full should be the producer's recommended citation verbatim whenever one exists (page "cite as" blocks, repository "Original citation", NBER suggested citations, "Reference:" lines inside the data files); slight modifications are fine to fix typos or spacing issues in the source (e.g. "mirror : a" → "mirror: a"), but never rephrase or reformat the citation style. If the recommended citation is for a working-paper version of a published work, keep it and append the published version (e.g. "Published as: …"). Construct a standard-format citation only when the producer recommends none.description describes the data product factually. When the producer's own text is factual (abstracts, methodology summaries), prefer it, exactly or lightly rephrased — don't rewrite for the sake of rewriting. When the producer's page only offers promotional or first-person copy ("our nation", "vital information", how the data guides funding), do NOT paste it — write a neutral description from OWID's point of view instead: who produces it, what it measures, coverage, cadence. (Schema guideline: "If the producer provides a good description, use that, either exactly or conveniently rephrased.")title/description describe the data product, while title_snapshot/description_snapshot describe the specific file. Whenever title_snapshot is set, also write a description_snapshot — that one doesn't need to be verbatim. OWID-side context (hand-transcription notes, "retrieved from the Internet Archive", etc.) belongs in description_snapshot, never inside the producer's description.© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/create-snapshot of owid/etl.
Open the folder on GitHubat commit 69ab20e
Create Snapshot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Create Snapshot this skillowid/etl | 158 | — | ~5.9k | Automated safety check: Pass | MIT | |
| Scikit LearnzLanqing/codex-claude-academic-skills | 4.6k | 17 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| TimesFM Forecastinggoogle-research/timesfm | 34k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| StatsmodelszLanqing/codex-claude-academic-skills | 4.6k | 16 repos | ~4.9k | Automated safety check: Pass | BSD-3-Clause | |
| Scientific Figure MakingChenLiu-1996/figures4papers | 8.1k | — | ~557 | Automated safety check: Pass | Custom licence |
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
google-research/timesfm
Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
zLanqing/codex-claude-academic-skills
Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.
ChenLiu-1996/figures4papers
Covers publication-ready matplotlib figures for academic papers, slides, and reports—bars, trends, scatter, heatmaps, and multi-panel layouts—with this…
Nuitka/Nuitka
Diagnose and fix ModuleNotFoundError in Nuitka standalone binaries caused by missing implicit imports.
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Works with
Categories
Create a new snapshot (DVC file, plus a Python script only when one is needed) from a urlmain and optional urldownload. Create Snapshot is an agent skill from owid/etl. Create a new snapshot (DVC file, plus a Python script only when one is needed) from a urlmain and optional urldownload.
Create Snapshot fits situations like: the user wants to add a new data source; create a snapshot from a URL.
Run `npx skills add owid/etl --skill create-snapshot -a claude-code`. Or copy the skill folder (.claude/skills/create-snapshot in owid/etl) into .claude/skills/create-snapshot in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill create-snapshot -a codex`. Or copy the skill folder (.claude/skills/create-snapshot in owid/etl) into .agents/skills/create-snapshot in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill create-snapshot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-snapshot, .gemini/skills/create-snapshot, .github/skills/create-snapshot and .opencode/skills/create-snapshot in your project.
Going by SKILL.md and its folder, Create Snapshot needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: ourworldindata.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Create Snapshot is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Create Snapshot: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), TimesFM Forecasting (google-research/timesfm, 34k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars) and Statsmodels (zLanqing/codex-claude-academic-skills, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 7, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.