Agent skill

Create Etl Steps

by owid in owid/etl

Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.

MITAuto-check passedData & Analytics

Install Create Etl Steps

skills CLI
$ npx skills add owid/etl --skill create-etl-steps -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install owid/etl create-etl-steps --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/create-etl-steps .claude/skills/create-etl-steps && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
create-etl-steps
GitHub stars
158
Token cost
~2.8k tokens
SKILL.md length
1,070 words
Files
1
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.

  • Works in 7 steps: Parse the snapshot path → Find the snapshot file extension → Determine the DAG file → …
  • Tasks that involve Data pipelines and ETL
  • SKILL.md covers Inputs and Workflow
  • Calls python and git

What it does

Create Etl Steps is an agent skill from owid/etl. Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.

When your agent uses it

  • Tasks that involve Data pipelines and ETL

Example prompts

  • “/create-etl-steps”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Parse the snapshot path
  2. Find the snapshot file extension
  3. Determine the DAG file
  4. Generate the step files
  5. Add DAG entries
  6. Name the checks the filled-in steps will need
  7. Report to the user

What it can do on your machine

Read from SKILL.md and the folder at commit bf5dc8e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Create Etl Steps loads about 2.8k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 1,070 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from owid/etl at commit bf5dc8e, republished under its MIT licence (© owid). 1,070 words, ~2,790 tokens.

Download SKILL.mdSave it as .claude/skills/create-etl-steps/SKILL.md (or your agent's skills folder).
name
create-etl-steps
description
Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.
triggers
create etl steps, create meadow garden grapher, create pipeline steps, scaffold etl steps
metadata.internal
true
metadata.owner
antea04

Create ETL Steps

Create meadow, garden, and grapher step files for a given snapshot, by running the same cookiecutter templates the wizard runs.

Never copy the templates into this file. apps/wizard/etl_steps/cookiecutter/{meadow,garden,grapher}/ is the single source of truth, and this skill invokes it via generate_step_to_channel. An earlier version of this skill embedded hand-copied templates; they drifted (the multi-snapshot meadow branch and non_redistributable for private datasets both went missing) while the copies that hadn't drifted made the rest look current. If a template needs changing, change it in apps/wizard/etl_steps/cookiecutter/.

/create-dataset calls this skill at its Step 5. Keep the two consistent: if the inputs or generated files change here, check whether create-dataset/SKILL.md needs a matching edit, and make it in the same commit.

Inputs

Required:

  • snapshot_path — in the format namespace/version/short_name (e.g. washu/2026-04-22/pm25_air_pollution)

Optional:

  • dag_file — which DAG file to add entries to (e.g. environment, climate). If not provided, ask the user.
  • is_private — default False. Affects both the generated metadata and the DAG URIs (see steps 4 and 5).
  • update_period_days — default 365.
  • topic_tags — default none.

Workflow

1. Parse the snapshot path

Extract:

  • namespace — e.g. washu
  • version — e.g. 2026-04-22
  • short_name — e.g. pm25_air_pollution
2. Find the snapshot file extension

Look in snapshots/<namespace>/<version>/ for a .dvc file matching <short_name>.*. The part between <short_name>. and .dvc is the file_extension.

For example: pm25_air_pollution.csv.dvc → file_extension = csv

The full snapshot filename is <short_name>.<file_extension>. Collect all the snapshot filenames the meadow step should read — if the chain has more than one snapshot, pass them all in step 4 and the meadow template generates the multi-snapshot loop by itself.

3. Determine the DAG file

If the user has not specified a DAG file, list the available files in dag/ (excluding archive/) and ask the user which one to use.

4. Generate the step files

Call the wizard's generator once per channel. It creates the directories, renders the templates, runs ruff on the generated Python, and copies the result into etl/steps/data/<channel>/:

bash
.venv/bin/python -c "
from apps.utils.files import generate_step_to_channel
from apps.wizard.etl_steps.utils import COOKIE_STEPS, remove_playground_notebook
from etl.owners import resolve_owner
import subprocess

namespace, version, short_name = '<namespace>', '<version>', '<short_name>'
snapshot_names = ['<short_name>.<file_extension>']  # every snapshot the meadow step reads
dag_file, is_private = '<dag_file>.yml', False
update_period_days, topic_tags = 365, []

git_name = subprocess.check_output(['git', 'config', 'user.name'], text=True).strip()
owner = resolve_owner(git_name)
# The garden .meta.yml only renders an `owners:` block when `owner` is truthy, so an
# unresolved name would silently produce metadata with no accountable owner.
assert owner, f'resolve_owner() did not recognize git user.name={git_name!r} — ask for a canonical owner'

common = {
    'namespace': namespace,
    'short_name': short_name,
    'version': version,
    'add_to_dag': True,
    'dag_file': dag_file,
    'is_private': is_private,
}
per_channel = {
    'meadow': {'channel': 'meadow', 'snapshot_names_with_extension': snapshot_names},
    # topic_tags must be a pre-joined string, not a list: cookiecutter renders a list
    # as only its first element.
    'garden': {
        'channel': 'garden',
        'meadow_version': version,
        'update_period_days': update_period_days,
        'topic_tags': ('- ' + '\n- '.join(topic_tags)) if topic_tags else '',
        'owner': owner,
    },
    'grapher': {'channel': 'grapher', 'garden_version': version},
}

for channel, extra in per_channel.items():
    dataset_dir = generate_step_to_channel(cookiecutter_path=COOKIE_STEPS[channel], data={**common, **extra})
    # The meadow and garden cookiecutters both ship a playground.ipynb. The wizard keeps it only
    # for garden when the user asked for a notebook; this skill never does, so always drop it.
    remove_playground_notebook(dataset_dir)
    print(f'{channel}: {dataset_dir}')
"

Notes on this call:

  • Run the channels one at a time, never concurrently. generate_step writes a temporary cookiecutter.json into the template directory and deletes it afterwards, so two simultaneous runs corrupt each other's context.
  • Every variable a template references must be present in data. There is no committed cookiecutter.json supplying defaults, so a missing key is a Jinja UndefinedError, not a silent blank. The dicts above are what apps/wizard/etl_steps/forms.py:309 passes; if a template gains a variable, it has to be added here too.
  • If the owner assert fires, stop and ask which colleague is the accountable owner, then set owner to that canonical name and re-run. resolve_owner only recognizes the git identities in etl/owners.py, so it returns None on an unmapped git config user.name — a cloud sandbox, a fresh checkout, or a name spelled differently from the enum. The wizard form degrades to an empty string there, but it has a human in front of it who can see the missing block; a skill run does not, and CLAUDE.md requires owners on every dataset. Pick the name from the schemas/dataset-schema.json enum, and add the git identity to etl/owners.py if it is a colleague who is simply missing from the map.
  • generate_step prints the context dictionary to stdout, and importing apps.wizard logs a No runtime found, using MemoryCacheStorageManager warning from Streamlit. Both are expected noise, not errors.
  • Use /create-playground if the user does want a playground notebook, rather than keeping the cookiecutter's copy.

Files generated, after the playground removal: meadow .py; garden .py, .meta.yml, .countries.json, .excluded_countries.json; grapher .py. Verify the notebook is gone — leaving one behind is the easiest thing to get wrong here, since two of the three channels ship it.

Show full SKILL.md (472 more words)Show less
5. Add DAG entries

Append the following entries to dag/<dag_file>.yml under the steps: key, using ruamel_load / ruamel_dump to preserve comments. For a private dataset every data:// below becomes data-private:// (matching private_suffix in the wizard form):

yaml
  data://meadow/<namespace>/<version>/<short_name>:
    - snapshot://<namespace>/<version>/<short_name>.<file_extension>
  data://garden/<namespace>/<version>/<short_name>:
    - data://meadow/<namespace>/<version>/<short_name>
  data://grapher/<namespace>/<version>/<short_name>:
    - data://garden/<namespace>/<version>/<short_name>

The snapshot URI has its own prefix, driven by the snapshot's .dvc, not by is_private. A snapshot whose .dvc sets is_public: false is referenced as snapshot-private://; everything else as snapshot://. Read is_public out of each .dvc rather than assuming — a private dataset is normally built on private snapshots, but the two flags are independent, and a public snapshot can feed a private dataset. Getting this wrong is silent: snapshot-private:// builds a SnapshotStepPrivate, whose run() asserts is_public is False before pulling, and --private filtering keys off the prefix too, so a private snapshot mislabeled snapshot:// loses that assert and is no longer excluded from a public run. Every one of the 294 private snapshots in the active DAG uses snapshot-private://, with no exceptions — a plain snapshot:// on a private snapshot would be the first.

List every snapshot from step 2 as a dependency of the meadow step, not just the first.

python
from etl.files import ruamel_load, ruamel_dump
from etl.snapshot import Snapshot

base = "<namespace>/<version>/<short_name>"
snapshot_names = ["<short_name>.<file_extension>"]  # every snapshot from step 2
is_private = False
data_prefix = "data-private" if is_private else "data"

# Each snapshot's own is_public decides its prefix, independently of is_private.
snapshot_uris = []
for name in snapshot_names:
    path = f"<namespace>/<version>/{name}"
    prefix = "snapshot" if Snapshot(path).metadata.is_public else "snapshot-private"
    snapshot_uris.append(f"{prefix}://{path}")

dag_path = "dag/<dag_file>.yml"
with open(dag_path, "r") as f:
    data = ruamel_load(f)
data["steps"][f"{data_prefix}://meadow/{base}"] = snapshot_uris
data["steps"][f"{data_prefix}://garden/{base}"] = [f"{data_prefix}://meadow/{base}"]
data["steps"][f"{data_prefix}://grapher/{base}"] = [f"{data_prefix}://garden/{base}"]
with open(dag_path, "w") as f:
    f.write(ruamel_dump(data))
6. Name the checks the filled-in steps will need

Don't run /check-outdated-practices here. What this skill produces is untouched cookiecutter output, and the templates are verified clean against the detector's full pattern set — so running it on the scaffold is a guaranteed no-op. The patterns it looks for enter when the scaffold is adapted: real load logic, harmonization, aggregations, hand-copied helper modules like *_omms.py. That's why /create-dataset runs it at its Step 5, after adapting these files, and /update-dataset runs it at step 1b on files etl update carried over from the previous version — neither of which is scaffold output.

Template drift is covered by the detector itself rather than by a check here: apps/wizard/**/cookiecutter/** is in the scope of every pattern, so a stale practice in a template shows up in the editor as soon as someone opens it.

So report the checks the person will need once the steps do something, rather than running them on empty files. The metadata checks in particular have nothing to bite on yet — the scaffolded .meta.yml is entirely commented out:

  • /check-outdated-practices — after adapting the step .py files, and on any helper module copied in by hand
  • /check-metadata-style — user-facing text against the Writing and Style Guide, and Jinja rendering artifacts once the metadata uses templates
  • /check-metadata-typos — spelling

Also flag .claude/rules/sanity-checks.md if the garden step will do more than load-and-format: assertions are expected in the step, and the scaffold has none.

7. Report to the user

List all files created and the DAG entries added, and the deferred checks from step 6 — saying plainly that nothing has been checked yet because there is nothing to check, so the next person doesn't read silence as a clean bill of health. Suggest running:

bash
.venv/bin/etlr <namespace>/<version>/<short_name>

© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/create-etl-steps of owid/etl.

Open the folder on GitHubat commit bf5dc8e

Compare with similar skills

Create Etl Steps next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Create Etl Steps compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Create Etl Steps this skillowid/etl158—~2.8kAutomated safety check: PassMIT
Crawl4AI Web Scrapingsmallnest/goclaw5981 repos~2.5kAutomated safety check: PassMIT
Glue 09 10 Migrationaws-samples/aws-glue-samples1.5k—~2.4kAutomated safety check: PassMIT-0
Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples1.5k—~3.6kAutomated safety check: PassMIT-0
Dbt Databricks PR Readydatabricks/dbt-databricks380—~2.8kAutomated safety check: PassApache-2.0
Mz Dbt ReleaseMaterializeInc/materialize6.4k—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    598 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Glue 09 10 Migration

    aws-samples/aws-glue-samples

    Official

    Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.

    1.5k GitHub stars~2.4k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Official

    Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.

    1.5k GitHub stars~3.6k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Dbt Databricks PR Ready

    databricks/dbt-databricks

    Official

    A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.

    380 GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Mz Dbt Release

    MaterializeInc/materialize

    Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.

    6.4k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Erd Studio Setup

    liam-machine/erd-studio

    Friendly, step-by-step setup for ERD Studio in an existing dbt project, for people who may be new to dbt or data modelling.

    165 GitHub stars~8.5k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from owid/etl

All 35 skills in this repo
  • Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.

    158 GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.

    158 GitHub stars~19k tokensUpdated today
    Auto-check passed
  • Add new survey question codes (e.g. An agent skill from owid/etl.

    158 GitHub stars~11k tokensUpdated today
    Auto-check: notes
  • Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…

    158 GitHub stars~8.3k tokensUpdated today
    Auto-check passed
  • Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.

    158 GitHub stars~9.6k tokensUpdated today
    Auto-check: notes
  • Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.

    158 GitHub stars~7.3k tokensUpdated today
    Auto-check: notes

Questions about Create Etl Steps

What does Create Etl Steps do?

Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path. Create Etl Steps is an agent skill from owid/etl. Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.

When should I use Create Etl Steps?

Create Etl Steps fits situations like: tasks that involve Data pipelines and ETL.

How do I install Create Etl Steps in Claude Code?

Run `npx skills add owid/etl --skill create-etl-steps -a claude-code`. Or copy the skill folder (.claude/skills/create-etl-steps in owid/etl) into .claude/skills/create-etl-steps in your project. Claude Code loads it when a task matches its description.

How do I install Create Etl Steps in Codex?

Run `npx skills add owid/etl --skill create-etl-steps -a codex`. Or copy the skill folder (.claude/skills/create-etl-steps in owid/etl) into .agents/skills/create-etl-steps in your project. Codex loads it when a task matches its description.

Can I use Create Etl Steps in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill create-etl-steps -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-etl-steps, .gemini/skills/create-etl-steps, .github/skills/create-etl-steps and .opencode/skills/create-etl-steps in your project.

What does Create Etl Steps need to run?

Going by SKILL.md and its folder, Create Etl Steps needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Create Etl Steps access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Create Etl Steps safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Create Etl Steps use?

Create Etl Steps is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Create Etl Steps use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Create Etl Steps?

Skills that share tags, products or a category with Create Etl Steps: Crawl4AI Web Scraping (smallnest/goclaw, 598 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Create Etl Steps?

owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.