Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.
$ npx skills add owid/etl --skill create-etl-steps -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl create-etl-steps --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/create-etl-steps .claude/skills/create-etl-steps && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "create-etl-steps" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-etl-steps into .claude/skills/create-etl-steps/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-etl-steps", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/create-etl-stepsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill create-etl-steps -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl create-etl-steps --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/create-etl-steps .agents/skills/create-etl-steps && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "create-etl-steps" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-etl-steps into .agents/skills/create-etl-steps/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-etl-steps", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-etl-steps -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl create-etl-steps --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/create-etl-steps .cursor/skills/create-etl-steps && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "create-etl-steps" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-etl-steps into .cursor/skills/create-etl-steps/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-etl-steps", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/create-etl-steps--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill create-etl-steps -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl create-etl-steps --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/create-etl-steps .gemini/skills/create-etl-steps && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "create-etl-steps" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-etl-steps into .gemini/skills/create-etl-steps/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-etl-steps", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl create-etl-stepsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill create-etl-steps -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/create-etl-steps .github/skills/create-etl-steps && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "create-etl-steps" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-etl-steps into .github/skills/create-etl-steps/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-etl-steps", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-etl-steps -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl create-etl-steps --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/create-etl-steps .opencode/skills/create-etl-steps && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "create-etl-steps" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-etl-steps into .opencode/skills/create-etl-steps/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-etl-steps", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
create-etl-stepsCreate vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.
Create Etl Steps is an agent skill from owid/etl. Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bf5dc8e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythongitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Create Etl Steps loads about 2.8k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 1,070 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from owid/etl at commit bf5dc8e, republished under its MIT licence (© owid). 1,070 words, ~2,790 tokens.
.claude/skills/create-etl-steps/SKILL.md (or your agent's skills folder).Create meadow, garden, and grapher step files for a given snapshot, by running the same cookiecutter templates the wizard runs.
Never copy the templates into this file.
apps/wizard/etl_steps/cookiecutter/{meadow,garden,grapher}/is the single source of truth, and this skill invokes it viagenerate_step_to_channel. An earlier version of this skill embedded hand-copied templates; they drifted (the multi-snapshot meadow branch andnon_redistributablefor private datasets both went missing) while the copies that hadn't drifted made the rest look current. If a template needs changing, change it inapps/wizard/etl_steps/cookiecutter/.
/create-dataset calls this skill at its Step 5. Keep the two consistent: if the inputs or generated files change here, check whether create-dataset/SKILL.md needs a matching edit, and make it in the same commit.
Required:
snapshot_path — in the format namespace/version/short_name (e.g. washu/2026-04-22/pm25_air_pollution)Optional:
dag_file — which DAG file to add entries to (e.g. environment, climate). If not provided, ask the user.is_private — default False. Affects both the generated metadata and the DAG URIs (see steps 4 and 5).update_period_days — default 365.topic_tags — default none.Extract:
namespace — e.g. washuversion — e.g. 2026-04-22short_name — e.g. pm25_air_pollutionLook in snapshots/<namespace>/<version>/ for a .dvc file matching <short_name>.*. The part between <short_name>. and .dvc is the file_extension.
For example: pm25_air_pollution.csv.dvc → file_extension = csv
The full snapshot filename is <short_name>.<file_extension>. Collect all the snapshot filenames the meadow step should read — if the chain has more than one snapshot, pass them all in step 4 and the meadow template generates the multi-snapshot loop by itself.
If the user has not specified a DAG file, list the available files in dag/ (excluding archive/) and ask the user which one to use.
Call the wizard's generator once per channel. It creates the directories, renders the templates, runs ruff on the generated Python, and copies the result into etl/steps/data/<channel>/:
.venv/bin/python -c "
from apps.utils.files import generate_step_to_channel
from apps.wizard.etl_steps.utils import COOKIE_STEPS, remove_playground_notebook
from etl.owners import resolve_owner
import subprocess
namespace, version, short_name = '<namespace>', '<version>', '<short_name>'
snapshot_names = ['<short_name>.<file_extension>'] # every snapshot the meadow step reads
dag_file, is_private = '<dag_file>.yml', False
update_period_days, topic_tags = 365, []
git_name = subprocess.check_output(['git', 'config', 'user.name'], text=True).strip()
owner = resolve_owner(git_name)
# The garden .meta.yml only renders an `owners:` block when `owner` is truthy, so an
# unresolved name would silently produce metadata with no accountable owner.
assert owner, f'resolve_owner() did not recognize git user.name={git_name!r} — ask for a canonical owner'
common = {
'namespace': namespace,
'short_name': short_name,
'version': version,
'add_to_dag': True,
'dag_file': dag_file,
'is_private': is_private,
}
per_channel = {
'meadow': {'channel': 'meadow', 'snapshot_names_with_extension': snapshot_names},
# topic_tags must be a pre-joined string, not a list: cookiecutter renders a list
# as only its first element.
'garden': {
'channel': 'garden',
'meadow_version': version,
'update_period_days': update_period_days,
'topic_tags': ('- ' + '\n- '.join(topic_tags)) if topic_tags else '',
'owner': owner,
},
'grapher': {'channel': 'grapher', 'garden_version': version},
}
for channel, extra in per_channel.items():
dataset_dir = generate_step_to_channel(cookiecutter_path=COOKIE_STEPS[channel], data={**common, **extra})
# The meadow and garden cookiecutters both ship a playground.ipynb. The wizard keeps it only
# for garden when the user asked for a notebook; this skill never does, so always drop it.
remove_playground_notebook(dataset_dir)
print(f'{channel}: {dataset_dir}')
"Notes on this call:
generate_step writes a temporary cookiecutter.json into the template directory and deletes it afterwards, so two simultaneous runs corrupt each other's context.data. There is no committed cookiecutter.json supplying defaults, so a missing key is a Jinja UndefinedError, not a silent blank. The dicts above are what apps/wizard/etl_steps/forms.py:309 passes; if a template gains a variable, it has to be added here too.owner assert fires, stop and ask which colleague is the accountable owner, then set owner to that canonical name and re-run. resolve_owner only recognizes the git identities in etl/owners.py, so it returns None on an unmapped git config user.name — a cloud sandbox, a fresh checkout, or a name spelled differently from the enum. The wizard form degrades to an empty string there, but it has a human in front of it who can see the missing block; a skill run does not, and CLAUDE.md requires owners on every dataset. Pick the name from the schemas/dataset-schema.json enum, and add the git identity to etl/owners.py if it is a colleague who is simply missing from the map.generate_step prints the context dictionary to stdout, and importing apps.wizard logs a No runtime found, using MemoryCacheStorageManager warning from Streamlit. Both are expected noise, not errors./create-playground if the user does want a playground notebook, rather than keeping the cookiecutter's copy.Files generated, after the playground removal: meadow .py; garden .py, .meta.yml, .countries.json, .excluded_countries.json; grapher .py. Verify the notebook is gone — leaving one behind is the easiest thing to get wrong here, since two of the three channels ship it.
Append the following entries to dag/<dag_file>.yml under the steps: key, using ruamel_load / ruamel_dump to preserve comments. For a private dataset every data:// below becomes data-private:// (matching private_suffix in the wizard form):
data://meadow/<namespace>/<version>/<short_name>:
- snapshot://<namespace>/<version>/<short_name>.<file_extension>
data://garden/<namespace>/<version>/<short_name>:
- data://meadow/<namespace>/<version>/<short_name>
data://grapher/<namespace>/<version>/<short_name>:
- data://garden/<namespace>/<version>/<short_name>The snapshot URI has its own prefix, driven by the snapshot's .dvc, not by is_private. A snapshot whose .dvc sets is_public: false is referenced as snapshot-private://; everything else as snapshot://. Read is_public out of each .dvc rather than assuming — a private dataset is normally built on private snapshots, but the two flags are independent, and a public snapshot can feed a private dataset. Getting this wrong is silent: snapshot-private:// builds a SnapshotStepPrivate, whose run() asserts is_public is False before pulling, and --private filtering keys off the prefix too, so a private snapshot mislabeled snapshot:// loses that assert and is no longer excluded from a public run. Every one of the 294 private snapshots in the active DAG uses snapshot-private://, with no exceptions — a plain snapshot:// on a private snapshot would be the first.
List every snapshot from step 2 as a dependency of the meadow step, not just the first.
from etl.files import ruamel_load, ruamel_dump
from etl.snapshot import Snapshot
base = "<namespace>/<version>/<short_name>"
snapshot_names = ["<short_name>.<file_extension>"] # every snapshot from step 2
is_private = False
data_prefix = "data-private" if is_private else "data"
# Each snapshot's own is_public decides its prefix, independently of is_private.
snapshot_uris = []
for name in snapshot_names:
path = f"<namespace>/<version>/{name}"
prefix = "snapshot" if Snapshot(path).metadata.is_public else "snapshot-private"
snapshot_uris.append(f"{prefix}://{path}")
dag_path = "dag/<dag_file>.yml"
with open(dag_path, "r") as f:
data = ruamel_load(f)
data["steps"][f"{data_prefix}://meadow/{base}"] = snapshot_uris
data["steps"][f"{data_prefix}://garden/{base}"] = [f"{data_prefix}://meadow/{base}"]
data["steps"][f"{data_prefix}://grapher/{base}"] = [f"{data_prefix}://garden/{base}"]
with open(dag_path, "w") as f:
f.write(ruamel_dump(data))Don't run /check-outdated-practices here. What this skill produces is untouched cookiecutter output, and the templates are verified clean against the detector's full pattern set — so running it on the scaffold is a guaranteed no-op. The patterns it looks for enter when the scaffold is adapted: real load logic, harmonization, aggregations, hand-copied helper modules like *_omms.py. That's why /create-dataset runs it at its Step 5, after adapting these files, and /update-dataset runs it at step 1b on files etl update carried over from the previous version — neither of which is scaffold output.
Template drift is covered by the detector itself rather than by a check here: apps/wizard/**/cookiecutter/** is in the scope of every pattern, so a stale practice in a template shows up in the editor as soon as someone opens it.
So report the checks the person will need once the steps do something, rather than running them on empty files. The metadata checks in particular have nothing to bite on yet — the scaffolded .meta.yml is entirely commented out:
/check-outdated-practices — after adapting the step .py files, and on any helper module copied in by hand/check-metadata-style — user-facing text against the Writing and Style Guide, and Jinja rendering artifacts once the metadata uses templates/check-metadata-typos — spellingAlso flag .claude/rules/sanity-checks.md if the garden step will do more than load-and-format: assertions are expected in the step, and the scaffold has none.
List all files created and the DAG entries added, and the deferred checks from step 6 — saying plainly that nothing has been checked yet because there is nothing to check, so the next person doesn't read silence as a clean bill of health. Suggest running:
.venv/bin/etlr <namespace>/<version>/<short_name>© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/create-etl-steps of owid/etl.
Open the folder on GitHubat commit bf5dc8e
Create Etl Steps next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Create Etl Steps this skillowid/etl | 158 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 598 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Glue 09 10 Migrationaws-samples/aws-glue-samples | 1.5k | — | ~2.4k | Automated safety check: Pass | MIT-0 | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Dbt Databricks PR Readydatabricks/dbt-databricks | 380 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Mz Dbt ReleaseMaterializeInc/materialize | 6.4k | — | ~1.2k | Automated safety check: Pass | Custom licence |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
databricks/dbt-databricks
A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.
MaterializeInc/materialize
Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.
liam-machine/erd-studio
Friendly, step-by-step setup for ERD Studio in an existing dbt project, for people who may be new to dbt or data modelling.
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Categories
Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path. Create Etl Steps is an agent skill from owid/etl. Create vanilla meadow, garden, and grapher ETL step files by invoking the wizard's cookiecutter templates, given a snapshot path.
Create Etl Steps fits situations like: tasks that involve Data pipelines and ETL.
Run `npx skills add owid/etl --skill create-etl-steps -a claude-code`. Or copy the skill folder (.claude/skills/create-etl-steps in owid/etl) into .claude/skills/create-etl-steps in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill create-etl-steps -a codex`. Or copy the skill folder (.claude/skills/create-etl-steps in owid/etl) into .agents/skills/create-etl-steps in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill create-etl-steps -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-etl-steps, .gemini/skills/create-etl-steps, .github/skills/create-etl-steps and .opencode/skills/create-etl-steps in your project.
Going by SKILL.md and its folder, Create Etl Steps needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Create Etl Steps is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Create Etl Steps: Crawl4AI Web Scraping (smallnest/goclaw, 598 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 380 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.