Excel and CSV Data Analysis
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
Create a brand-new OWID dataset in ETL from a data file the user provides — a local CSV/Excel, a downloaded file, or a web link to data (snapshot → meadow → garden → grapher → PR → staging server).
$ npx skills add owid/etl --skill create-dataset -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl create-dataset --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/create-dataset .claude/skills/create-dataset && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "create-dataset" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-dataset into .claude/skills/create-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-dataset", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/create-datasetType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill create-dataset -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl create-dataset --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/create-dataset .agents/skills/create-dataset && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "create-dataset" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-dataset into .agents/skills/create-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-dataset", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-dataset -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl create-dataset --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/create-dataset .cursor/skills/create-dataset && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "create-dataset" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-dataset into .cursor/skills/create-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-dataset", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/create-dataset--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill create-dataset -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl create-dataset --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/create-dataset .gemini/skills/create-dataset && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "create-dataset" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-dataset into .gemini/skills/create-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-dataset", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl create-datasetInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill create-dataset -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/create-dataset .github/skills/create-dataset && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "create-dataset" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-dataset into .github/skills/create-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-dataset", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-dataset -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl create-dataset --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/create-dataset .opencode/skills/create-dataset && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "create-dataset" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-dataset into .opencode/skills/create-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-dataset", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
create-datasetCreate a brand-new OWID dataset in ETL from a data file the user provides — a local CSV/Excel, a downloaded file, or a web link to data (snapshot → meadow → garden → grapher → PR → staging server).
Create Dataset is an agent skill from owid/etl. Create a brand-new OWID dataset in ETL from a data file the user provides — a local CSV/Excel, a downloaded file, or a web link to data (snapshot → meadow → garden → grapher → PR → staging server). Use when someone has data they want in ETL so they can build charts. Designed for non-technical "Cloud co-work" users: infer aggressively, build a working dataset first, then ask the person to review and correct.
Its SKILL.md is about 7.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL, Excel spreadsheets and CSV and tabular files. It works with Microsoft Excel. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 69ab20e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitghpythonmakecurlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, gh and curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Create Dataset loads about 7.1k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 3,987 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from owid/etl at commit 69ab20e, republished under its MIT licence (© owid). 3,987 words, ~7,139 tokens.
.claude/skills/create-dataset/SKILL.md (or your agent's skills folder).Turn data the user provides into a fully wired-up OWID dataset — snapshot, meadow, garden, grapher, DAG entries, a draft PR, and a staging server where the user can build charts.
The input can be anything: a local CSV or Excel file, a file they just downloaded, or a link to a page/file on the web. Whatever the form, the job is the same — get the tabular data into a snapshot, then build the chain on top of it.
This skill is for people who are not ETL experts (typically working in Claude Code on the web / "Cloud co-work"). They may know some metadata (units, what the columns mean) or they may just have a link to the source page. They should not be quizzed field-by-field.
Paired skill — keep in sync.
/create-snapshotis the canonical owner of the snapshot-creation conventions this skill consumes in Step 4: whenever snapshot conventions change here or there, mirror the change in the other file in the same commit — see the mirror note there. The update-side skills are part of the same family: this skill points into/update-dataset's canonical sections (§5b-bis sanity bounds, §5c harmonization audit, §6b metadata quality, §6c metadata checklist + link verification, §6d scheduled issues), whose outcomes/review-data-prverifies — when those sections change, check whether this file needs a matching edit.
attribution_short: TBD), note it in the review, and keep going. A dataset that's 80% right and on staging beats a perfect one that never ships.Required (one of):
Optional:
metadata_url — a link to the source page / documentation (if different from a data link). If given, fetch it for metadata (producer, citation, license, definitions).master firstUsers are often not experienced with git and may be sitting on a stale checkout. Before doing anything, get the repo onto a fresh master so the new branch and PR are based on current code:
git switch master && git pull --ff-only origin masterIf they have uncommitted local changes or a detached HEAD that blocks this, surface it plainly and ask how they want to proceed — don't force it. (When testing/iterating in an unusual git state, this step may be intentionally skipped — but for a real user, always start here.)
If the input is a web link, first obtain the actual data file: download a direct CSV/Excel/JSON URL (use etl.http's session for OWID hosts; plain download otherwise), or if it's a landing page, WebFetch it to find the download link. Save it locally so the snapshot can ingest it. If the input is already a local file, use it directly.
Then read the file and figure out its shape. Do not ask the user anything in this step.
duckdb skill or a short .venv/bin/python snippet — whatever's quickest.entity, country, nation, location, geo, area, or region. OWID-exported CSVs usually call it Entity. The right column holds country-like names, not codes. If there's both a name column and an ISO code column, use the name column and drop the code.year, date, time, period. A column of 4-digit integers in ~1500–2100 is a year; ISO-date strings are a date. If the time dimension is a date (not a plain year), the garden step must keep it as date and .format(["country", "date"]).sex, age, variant, fuel_type — it's a dimension, not an indicator. The table is then long-format keyed by ["country", "year", <dim>...]. Most simple CSVs have none; handle them only if present (see the long-format note in /update-dataset).| Signal in column name / values | Inferred unit / short_unit |
|---|---|
share, pct, percent, _rate, and values mostly in 0–100 | % / % |
share / proportion with values in 0–1 | flag: is this a fraction (×100 → %) or already %? |
usd, gdp, price, cost, $ | US dollars / $ |
population, count, number, _n, integer counts | leave unit descriptive (e.g. cars), short_unit empty |
tonnes, kg, kwh, gwh, co2, emissions | physical unit from the name |
| can't tell | leave unit: "", flag in review |
Electric car sales (IEA, 2026) - data.csv implies: title Electric car sales, producer IEA, year 2026. Propose:short_name — snake_case of the core title (electric_car_sales).namespace — the producer's slug if one already exists under snapshots/ (check ls snapshots/), else a sensible new one.version — today's date (date -u +"%Y-%m-%d").Write a one-paragraph internal summary of what you found before moving on.
Now ask the user once, with all your best guesses pre-filled, using AskUserQuestion where it fits. Frame it as "here's what I figured out — correct anything that's wrong, otherwise I'll build it." Before listing the specifics, set expectations in one plain sentence so the end isn't a surprise — e.g. "I'll build the dataset and put it on a staging server for you to review; once you're happy you merge the PR and it goes live." Then keep the questions to the few things that genuinely can't be guessed or that would be expensive to get wrong:
metadata_url (the source page). If they give one, fetch it now for citation, license, and column definitions.CC BY 4.0 for academic/IGO sources if unknown) and let them correct.# NOTE: in the snapshot .dvc (Step 4 follows /create-snapshot's convention) — that NOTE, not the PR body, is the baseline the update and review workflows diff against; mention them in the PR body as well for the reviewer.ev_sales_share a percentage 0–100 or a fraction 0–1?"). Don't ask about columns you're confident on.Everything else (units you inferred, descriptions, topic tags) you'll fill in and surface for review later — don't ask now.
If a metadata_url was provided, fetch it (WebFetch) and extract producer, citation_full, attribution_short, date_published, license, and any column definitions — same fields as /create-snapshot Step 1.
On PR permissions: opening the draft PR and pushing to staging is this skill's deliverable. If a session-level rule says "don't open PRs unless explicitly asked," treat the user invoking this skill as that explicit request — go ahead and open the PR. The one exception is when you can't (e.g. the session pins you to a fixed pre-assigned branch you must not leave): in that case build on the current branch, open the PR from it if you can, and say clearly in the Step 7 handoff what you did and didn't do — don't silently skip the PR.
etl pr needs a branch (it fails on a detached HEAD). Create the PR scaffold up front so the rest of the work lands on a branch:
.venv/bin/etl pr "Add <dataset title> (<producer>)" dataThis creates the branch + a draft PR and does not commit. (If the user is on a detached HEAD or the working tree is dirty in a way that blocks this, create a branch with git checkout -b data-<short_name> first and tell them you'll open the PR after the build.)
This is a manual-import snapshot (you hand the snapshot a local file rather than relying on a stable download URL — even for web inputs, you've already saved the file locally in Step 1). Use the file's real extension in the .dvc name (.csv, .xlsx, …). Follow /create-snapshot conventions, with these specifics:
snapshots/<namespace>/<version>/<short_name>.<ext>.dvc — fill origin from Step 2 (title, producer, citation_full, attribution_short, date_published, url_main, license). description must describe the data product factually — use the producer's own text when it is factual (page prose or paper abstract), but rewrite promotional or first-person copy from OWID's point of view (see /create-snapshot) — and citation_full the producer's recommended citation verbatim when one exists (slight modifications only to fix typos or spacing in the source); for the license, check the documentation too and warn the user if none is stated anywhere (fall back to © <producer> (<year>)); if the file is one table/extract of a broader product, use title_snapshot + description_snapshot for the file specifics and any OWID-side notes (see /create-snapshot). Set date_accessed: <version>. Omit fields you don't have rather than leaving them blank; use TBD placeholders only where the review needs to flag them.snapshots/<namespace>/<version>/<short_name>.py — the manual-import script.Generate both files with the command in /create-snapshot step 3, passing dataset_manual_import: True and dvc_only: False. It runs the wizard's snapshot cookiecutter, which emits the path_to_file variant of run() and — the part that is easy to get wrong by hand — nests license inside origin rather than at the top level. Don't hand-write either file, and don't copy a template into this skill: the cookiecutter is the single source of truth, and copies drift.
Then run it against the user's file:
.venv/bin/etls <namespace>/<version>/<short_name> --path-to-file "<path_to_file>"Scaffold the three steps with /create-etl-steps (DAG file = the topic that best fits, e.g. energy, health; ask in Step 2 if unclear). Then adapt:
country if needed, cast low-cardinality string columns (country, dims) to category, tb.format(["country", "year"]) (or ["country", "date"], plus any dims). Keep it light.paths.regions.harmonize_names(tb, country_col="country", countries_file=paths.country_mapping_path), then tb.format(...). Add sanity_check_inputs / sanity_check_outputs if the step does more than load-and-format (see CLAUDE.md "Sanity checks"); ground every threshold in the built data and pick value bounds from the by-indicator-type table in /update-dataset §5b-bis, then negative-test the checks. Don't strip origins — follow the metadata-preserving patterns in CLAUDE.md (pr.concat, no np.where, etc.).<short_name>.meta.yml in garden) — generate it with /generate-metadata, then verify it against the mandatory-fields checklist in /update-dataset §6c. Fill title, unit, short_unit, description_short, description_key (free-form markdown prose; non-empty), display.name, display.numDecimalPlaces, display.tolerance per indicator; topic_tags and processing_level in definitions.common; presentation.attribution_short explicitly under definitions.common.presentation — it does not inherit from the origin's attribution_short. In the dataset block, set update_period_days plus owners: the user is the new dataset's first owner — resolve their canonical OWID name from git config user.name via etl.owners.resolve_owner (must match the schemas/dataset-schema.json enum; add the mapping in etl/owners.py + an enum row if missing), mirroring /update-dataset step 1a-bis. Use the units you inferred in Step 1; mark anything uncertain so it shows up in the review./check-outdated-practices on every new step file (including any helper modules, and the Step 4 snapshot .py if one was written) and fix findings before the first run, per /update-dataset step 1b — run the skill, don't eyeball the patterns./update-dataset "Removing the old version & reordering the DAG", step 4) and verify it parses: .venv/bin/python -c "from etl.dag_helpers import load_dag; load_dag()"..venv/bin/etlr <namespace>/<version>/<short_name>Fix whatever breaks (trace upstream, never mask). The most common task is country harmonization: the garden run logs unmatched country names that need mapping into <short_name>.countries.json.
Shortcut for OWID-exported CSVs: if the file came out of OWID (entity column literally named Entity), the country names are almost always already canonical. Don't sit through an interactive etl harmonize — auto-build the mapping by matching each entity against the canonical regions (and their aliases), and only fall back to manual mapping for the leftovers:
import json
from pathlib import Path
from owid.catalog import Dataset
import pandas as pd
# Needs the regions dataset built locally:
# .venv/bin/etlr data://garden/regions/2023-01-01/regions
tb_regions = Dataset(str(sorted(Path("data/garden/regions").glob("*/regions"))[-1]))["regions"]
canonical = set(tb_regions["name"].dropna().astype(str))
alias_map = {}
for name, al in tb_regions[["name", "aliases"]].dropna(subset=["aliases"]).itertuples(index=False):
for a in json.loads(al):
alias_map[a] = name
entities = sorted(pd.read_csv("<path_to_csv>")["Entity"].unique())
mapping, unmatched = {}, []
for e in entities:
if e in canonical: mapping[e] = e
elif e in alias_map: mapping[e] = alias_map[e]
else: unmatched.append(e) # decide these by hand: typo, alias, custom aggregate, or excludeResolve unmatched by hand against canonical regions; follow the harmonization-audit guidance in /update-dataset (Step 5c). Residual aggregates the producer defines (e.g. Rest of World, European Union (27)) can be kept as custom entities (map to themselves) — they won't join with population/region data, so note that in the review. Don't silently drop countries — if you exclude any, list them.
After the garden step builds, also run the garden-output entity check from /update-dataset §5c (Python check #5): load the built garden tables and confirm every country value is canonical (regions + latest income groups). It catches what the mapping file can't see — inline tb["country"] = ... assignments and post-harmonization mutations. List any non-canonical entity (including the custom aggregates you kept) in the Step 7 review and the PR body as living outside the canonical system.
Then build and upload the grapher step to staging — target the grapher/... path (no --only, so the grapher:// MySQL upsert step actually runs):
STAGING=<branch> .venv/bin/etlr grapher/<namespace>/<version>/<short_name> --grapherConfirm the upsert actually succeeded before moving on: it should print the dataset's admin URL / id (…/admin/datasets/<id>). Capture that <id> — you'll hand it to the user in Step 7. If the upsert errored or printed no dataset, fix it now rather than handing over a link that won't resolve.
Optionally run /fact-check-dataset on garden/<namespace>/<version>/<short_name>. It's not part of the default build because it can consume many tokens (~12–20 web calls even for a small dataset) — mention it as an offer in the Step 7 handoff ("I can also fact-check the data and metadata against the source's documentation and independent sources — say the word") and run it when the user opts in, or proactively when the build surfaced red flags (values that look implausible, a source page contradicting the file).
When it runs: since the dataset is brand-new (no charts yet, few indicators), review all indicators. This catches the two things the Step 7 review table can't: metadata you inferred that the source's own documentation contradicts (units, definitions, scope), and values the source itself got wrong (unit slips, wrong-year rows) — verified against independent sources online. Fold the findings into the handoff in plain language ("I double-checked the numbers against <independent source> — X and Y match; Z looks off, here's why"), and route confirmed source errors to <short_name>.corrections.yml per that skill's routing table rather than editing the data inline.
Quality pass before handoff. Run all four checks from /update-dataset §6b on the new garden + grapher .meta.yml files — /check-metadata-typos, /check-metadata-style (its first pass covers Jinja spacing), the general-audience clarity checklist, and the dimension sweep. Include the Step 4 .dvc in the /check-metadata-typos scope: its description / title / citation_full are user-facing too, and no other check spell-checks them (/create-snapshot step 5 makes the same pass for a standalone snapshot — keep the two consistent). Leave citation_full alone unless the producer's own page has the word right; it is verbatim producer text whenever a piece of user-facing text is shared across sibling variants — Jinja over a dimension (the long-format case Step 1 item 5 detects) or a definitions: key that several indicator blocks reuse, which a wide-format file with one column per subgroup produces without any extra-dimension column to detect. A sentence written for one variant renders on all its siblings, where what each view already restricts can make it false; render the text per variant and read each output as a reader of that chart. A brand-new dataset is exactly where such shared text gets written for the first time, so it needs the sweep as much as an update does. Then run the link-verification loop from /update-dataset §6c on every URL in the new .dvc and .meta.yml files, including its anchor-fragment pass for URLs with #… (a curl non-2xx is a signal, not proof — escalate WebFetch → Wayback availability API, and per §6c no automated signal is decisive: a link failing every automated check goes to the user for a browser confirmation, never auto-marked broken or replaced). After any .meta.yml edit, re-run the affected step (--grapher for grapher) so the built catalog and staging reflect it.
Run make check, confirm you're still on the work branch (git branch --show-current — a branch switch in the user's IDE silently moves your shell too), then commit and push:
git add .
git commit -m "📊🤖 Add <dataset title> (<producer>)"
git pushUpdate the PR description (attribution blockquote per CLAUDE.md "Team"; what the dataset is, source, coverage, indicator count). Then post @codex review as a separate PR comment; while the user reviews, address any valid findings and resolve the threads per /update-dataset step 10 — don't block the handoff on the review round.
Hand the user a review — this is the second and final checkpoint. Present a compact table they can scan and correct:
| Indicator (column) | Title | Unit | Decimals | Inferred? |
|---|
Plus a short list of:
TBD placeholders).Always tell them where the metadata file lives locally, so they (or you) can edit it directly:
etl/steps/data/garden/<namespace>/<version>/<short_name>.meta.yml — this is the file to edit for titles, units, descriptions, and decimals.etl/steps/data/garden/<namespace>/<version>/<short_name>.countries.jsonThen give them the staging links so they can build charts. Paste the actual URLs directly into the chat — don't tell them to "open the PR" or "go to the staging admin"; non-experts won't know where those are, and they're unlikely to open the PR at all:
<id> you captured in Step 6), https://staging-site-<branch>/admin/datasets/<id>. This is where they create charts from the new indicators.curl -so /dev/null -w "%{http_code}" https://staging-site-<branch>/admin/ stops returning a connection error / 404 "site not found") so you only give them a link that already works. Handing over a link before the build is ready — with no signal that waiting will fix it — reliably reads as "it's broken / I don't see my dataset."Ask for corrections in plain terms ("anything in the table look wrong? any column you'd describe differently?"). Apply their feedback by editing the .meta.yml / .countries.json and re-running the affected step (--grapher for grapher), then push again.
The dataset and any charts they build live on the staging server, not on ourworldindata.org. Getting them to live is a manual step the user has to take, and the workflow (approve charts in chart-diff, then merge the PR) is unfamiliar to non-experts — they will not discover it on their own. Spell it out explicitly in the handoff, with the real links pasted in:
Charts must be approved in chart-diff before they sync to live — and this is a human step, not yours. If the user creates any charts on the staging admin, those charts only reach production if they're approved in chart-diff first. Never approve charts yourself (no CLI, no script, no direct DB write) — chart-diff exists so a person eyeballs each chart before it goes live, and that review must stay human. Give them the direct link and tell them to click Approve on each chart themselves:
http://staging-site-<branch>/etl/wizard/chart-diffCreating the charts is also something the user does in the admin — assume they'll need to visit the admin at some point regardless; your job is to make that the only thing they have to do there.
Merging the PR is what publishes the dataset to live. No external review is required for a data PR like this. The user can click "Ready for review" then "Squash and merge" on the PR themselves (link the PR URL) — but you can also do the merge for them from the session, if they explicitly ask you to and the checks are green. This is the one outward-facing action where, given an explicit "merge it" from the user, you should go ahead — it saves a non-expert a trip to GitHub. Do it carefully:
gh pr checks <number> # confirm checks are green first
gh pr ready <number> # take it out of draft
gh pr merge <number> --squash # publishOffer a scheduled update issue if the data refreshes regularly. If update_period_days is roughly in [2, 366] (updated at least yearly but less than daily), the dataset qualifies for a scheduled "Data update" issue in owid/owid-issues, so future refreshes don't depend on anyone's memory. Offer to create the workflow per /update-dataset §6d — cron timed shortly after the producer's expected release window, issue body pointing the next updater at /update-dataset <short_name> — and only create it with the user's sign-off (the commit goes straight to owid-issues main).
Make this a short, plain-language checklist at the end of your handoff — e.g. "When you're happy: (1) approve your charts here «chart-diff link», (2) then either merge the PR yourself «PR link» or just tell me to merge it and I'll do it once the checks are green — it's live a few minutes later." Paste the real URLs, not placeholders.
Entity → country. OWID-exported files name the country column Entity; rename it in meadow.master (Step 0). Saves you from basing the branch/PR on a stale checkout — common with non-git-savvy users. etl pr and staging also need a real branch (not a detached HEAD).© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/create-dataset of owid/etl.
Open the folder on GitHubat commit 69ab20e
Create Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Create Dataset this skillowid/etl | 158 | — | ~7.1k | Automated safety check: Pass | MIT | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| CSV Data Analysis5zjk5/prompt-engineering | 127 | — | ~2.6k | Automated safety check: Pass | None | |
| Dataset Quality Auditzebbern/claude-code-guide | 4.6k | — | ~996 | Automated safety check: Pass | MIT | |
| Multi Source Data Integration ExtractionDrchronx/ai-agent-research-starter-kit | 134 | — | ~671 | Automated safety check: Pass | Custom licence | |
| Regression Modelerzebbern/claude-code-guide | 4.6k | — | ~909 | Automated safety check: Pass | MIT |
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
5zjk5/prompt-engineering
This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.
zebbern/claude-code-guide
Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…
Drchronx/ai-agent-research-starter-kit
Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.
zebbern/claude-code-guide
Run regression analysis (OLS or logistic) on uploaded CSV/Excel data, generating coefficients, R², p-values, VIF, and plain-language interpretation.
franklee16/academic-research-skills
Multi-domain data analysis router for uploaded datasets, spreadsheets, CSV/XLSX files, business metric tables, or user questions that require selecting the right business analysis framework before…
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Works with
Categories
Create a brand-new OWID dataset in ETL from a data file the user provides — a local CSV/Excel, a downloaded file, or a web link to data (snapshot → meadow → garden → grapher → PR → staging server). Create Dataset is an agent skill from owid/etl. Create a brand-new OWID dataset in ETL from a data file the user provides — a local CSV/Excel, a downloaded file, or a web link to data (snapshot → meadow → garden → grapher → PR → staging server).
Create Dataset fits situations like: someone has data they want in ETL so they can build charts; tasks that involve Data pipelines and ETL; tasks that involve Excel spreadsheets.
Run `npx skills add owid/etl --skill create-dataset -a claude-code`. Or copy the skill folder (.claude/skills/create-dataset in owid/etl) into .claude/skills/create-dataset in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill create-dataset -a codex`. Or copy the skill folder (.claude/skills/create-dataset in owid/etl) into .agents/skills/create-dataset in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill create-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-dataset, .gemini/skills/create-dataset, .github/skills/create-dataset and .opencode/skills/create-dataset in your project.
Going by SKILL.md and its folder, Create Dataset needs the command-line tools its instructions call (git, gh, python, make and curl).
SKILL.md contains no URLs. Its commands use git, gh and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Create Dataset is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.1k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Create Dataset: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), CSV Data Analysis (5zjk5/prompt-engineering, 127 stars), Dataset Quality Audit (zebbern/claude-code-guide, 4.6k stars) and Multi Source Data Integration Extraction (Drchronx/ai-agent-research-starter-kit, 134 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 7, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.