Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
Sync upstream grapher schema changes (new chart types, config fields, enum values) into the ETL repo — vendored schema, multidim-schema, dataset-schema, and regenerated Python types.
$ npx skills add owid/etl --skill sync-grapher-schema -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl sync-grapher-schema --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/sync-grapher-schema .claude/skills/sync-grapher-schema && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "sync-grapher-schema" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/sync-grapher-schema into .claude/skills/sync-grapher-schema/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sync-grapher-schema", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/sync-grapher-schemaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill sync-grapher-schema -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl sync-grapher-schema --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/sync-grapher-schema .agents/skills/sync-grapher-schema && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "sync-grapher-schema" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/sync-grapher-schema into .agents/skills/sync-grapher-schema/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sync-grapher-schema", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill sync-grapher-schema -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl sync-grapher-schema --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/sync-grapher-schema .cursor/skills/sync-grapher-schema && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "sync-grapher-schema" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/sync-grapher-schema into .cursor/skills/sync-grapher-schema/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sync-grapher-schema", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/sync-grapher-schema--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill sync-grapher-schema -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl sync-grapher-schema --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/sync-grapher-schema .gemini/skills/sync-grapher-schema && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "sync-grapher-schema" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/sync-grapher-schema into .gemini/skills/sync-grapher-schema/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sync-grapher-schema", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl sync-grapher-schemaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill sync-grapher-schema -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/sync-grapher-schema .github/skills/sync-grapher-schema && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "sync-grapher-schema" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/sync-grapher-schema into .github/skills/sync-grapher-schema/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sync-grapher-schema", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill sync-grapher-schema -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl sync-grapher-schema --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/sync-grapher-schema .opencode/skills/sync-grapher-schema && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "sync-grapher-schema" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/sync-grapher-schema into .opencode/skills/sync-grapher-schema/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sync-grapher-schema", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
sync-grapher-schemaSync upstream grapher schema changes (new chart types, config fields, enum values) into the ETL repo — vendored schema, multidim-schema, dataset-schema, and regenerated Python types.
Sync Grapher Schema is an agent skill from owid/etl. Sync upstream grapher schema changes (new chart types, config fields, enum values) into the ETL repo — vendored schema, multidim-schema, dataset-schema, and regenerated Python types. Use when the scheduled sync workflow opened a draft PR or issue that needs completing, when the web team announces a grapher schema change ("new chart type in Grapher", "I added a field to the grapher config"), when someone asks to "sync the grapher schema", or when grapher configs fail ETL validation on fields that work fine in the…
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL. It works with Python. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bf5dc8e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitpythonpytestmakeghcurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
files.ourworldindata.orgAlso links to:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Sync Grapher Schema loads about 2.7k tokens when it runs. Until then it costs about 138 tokens; SKILL.md has 1,206 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from owid/etl at commit bf5dc8e, republished under its MIT licence (© owid). 1,206 words, ~2,690 tokens.
.claude/skills/sync-grapher-schema/SKILL.md (or your agent's skills folder).The grapher chart-config schema is owned by the web team in owid-grapher and published at https://files.ourworldindata.org/schemas/grapher-schema.NNN.json. It is mutated in place without version bumps (e.g. dumbbell plots landed in .010 directly), so when it changes upstream, four things in this repo need to follow:
| File | Role | Sync mechanism |
|---|---|---|
schemas/grapher-schema.NNN.json | Vendored copy of upstream, and the single source of truth for the version (DEFAULT_GRAPHER_SCHEMA is derived from its $id) | automatic (--refresh; --bump-version for a new version) |
schemas/multidim-schema.json + schemas/explorer-schema.json | View config $refs into the grapher schema | manual: add $ref for new properties |
schemas/dataset-schema.json | Embedded grapher_config block (validates garden .meta.yml) | manual: mirror changes, preserve deviations |
etl/viz/chart/model/schema_types.py | Generated Python TypedDicts | automatic (regenerate) |
Unit tests enforce consistency between all of these (tests/test_schema_types_generation.py, test_grapher_config_schema_sync in tests/test_metadata_schemas.py), so partial syncs fail CI. Full background: docs/guides/grapher-schema-sync.md.
A. Completing a bot PR (the common case). The scheduled workflow (.github/workflows/sync-grapher-schema.yml) detected an upstream change and opened a draft PR on the auto-sync-grapher-schema branch with the automatic part (refreshed vendored copy + regenerated types) already committed.
test_grapher_config_schema_sync output is the todo list — usually just steps 2-3 below.git checkout -b <new> + close the bot PR).B. Ad-hoc / from scratch. Someone announced a change and you're not waiting for the cron (alternatively, trigger the workflow manually: gh workflow run sync-grapher-schema.yml). Follow all steps below.
C. Version bump. The workflow opened a "New grapher schema version published upstream" issue → see the "Version bump" section at the bottom.
(Entry point B only.) Use the standard flow: .venv/bin/etl pr "sync grapher schema (<short summary>)" chore, unless the user wants the changes on the current branch.
(Entry point A: already committed by the workflow — just read the diff with git show on the bot commit, then continue at step 2.)
.venv/bin/python scripts/generate_schema_types.py --refresh
git diff schemas/grapher-schema.*.jsonschemas/multidim-schema.jsonOnly needed for new top-level properties (new chart-type config objects like dumbbell, new view-level fields). Existing $refs resolve against the live schema automatically.
For each new upstream property that makes sense in a multidim/explorer view, add a $ref entry to the view config properties block (search for "chartTypes" to find it). The same applies to schemas/explorer-schema.json. Refs are local relative refs to the vendored copy (resolved offline by Chart.validate_schema):
"<newProp>": {
"$ref": "grapher-schema.NNN.json#/properties/<newProp>"
},Lesson learned (#6196 → #6200): forgetting this step is how dumbbell went missing — the generated types were patched by hand instead, which regeneration would have destroyed. Never edit schema_types.py directly.
schemas/dataset-schema.jsonThe grapher config is embedded inline (not $ref'd) under ...variables.additionalProperties.properties.presentation.properties.grapher_config.properties. Mirror every change from the step-1 diff into that block — new properties, new enum values, updated descriptions.
Preserve these deliberate ETL-side deviations (do NOT "fix" them to match upstream):
data, includedEntities.chartTypes enum includes WorldMap (not upstream).oneOf with a Jinja escape hatch — keep the wrapper, edit only the enum branch:"oneOf": [
{ "enum": [...sync these values...] },
{ "type": "string", "pattern": "{definitions" }
]"pattern": "<%" instead — same idea.).venv/bin/python scripts/generate_schema_types.py
git diff etl/viz/chart/model/schema_types.pySanity-check the diff: it should reflect exactly the upstream changes (plus any multidim $ref additions). If a class or field unexpectedly disappears, a $ref is probably missing (step 2).
Hand-written types (e.g. GroupViewsConfig) live in etl/viz/chart/model/params.py — never add them to the generated file.
.venv/bin/pytest tests/test_schema_types_generation.py tests/test_metadata_schemas.py tests -k "chart or schema" -m "not integration" -q
make checktest_grapher_config_schema_sync pinpoints any enum value or property still missing from the embedded block (exact JSON path in the failure message) — iterate on step 3 until green.
Commit with ✨🤖. In the PR body, list the upstream changes synced (link the Slack announcement if there is one) and which of the four files each change touched. For entry point A, mark the bot PR ready for review instead of writing a new body — just add a comment summarizing the manual propagation you did.
Rarer case — when the web team publishes a new schema version instead of mutating in place. Detected by the integration test test_no_newer_grapher_schema_version (compares the $id of upstream grapher-schema.latest.json against DEFAULT_GRAPHER_SCHEMA).
.venv/bin/python scripts/generate_schema_types.py --bump-versionThat one command reads the new version from grapher-schema.latest.json's $id, vendors it as schemas/grapher-schema.MMM.json, deletes the old copy, repoints every $ref in schemas/multidim-schema.json + schemas/explorer-schema.json, and regenerates schema_types.py. Nothing in etl/config.py is hand-edited: DEFAULT_GRAPHER_SCHEMA is derived from whichever schema is vendored (vendored_grapher_schema_id()), so it follows automatically — and it always resolves to a concrete version, never latest, since grapher's config migrations are keyed on the version.
Then, by hand:
git show HEAD:schemas/grapher-schema.NNN.json | diff -u - schemas/grapher-schema.MMM.json.grapher_config block in schemas/dataset-schema.json, preserving the deliberate ETL-side deviations. test_grapher_config_schema_sync fails until this is done.$refs (the grapher_schema examples in multidim-schema.json, which show what a new config should pin).--bump-version is idempotent — it prints "already vendoring the newest published schema" and changes nothing when there is no new version, so it is safe to run blind.
grapher_schema pins in MDIM configsEvery multidim and single-chart config pins grapher_schema: "NNN" — required, with no fallback (required in schemas/multidim-schema.json, re-checked by Chart.validate_grapher_schema_pinned(), and swept offline by test_multidim_configs_pin_grapher_schema). Leave those pins at their old version. They record what each config was authored against, which is what lets grapher migrate them to MMM on upsert. Bumping them would tell grapher the configs are already current and skip the migration — the exact failure the pins exist to prevent.
The one thing to check: --bump-version repoints multidim view-config validation at the new version, so a config that is no longer valid under MMM will now fail Chart.validate_schema(). Fix the config and bump only that config's pin, since at that point it genuinely was re-authored against MMM.
Views can also carry their own $schema inside a config block, which overrides the chart-level pin (grapher spreads the view config last). As of #6705 follow-up no step does this any more, and ETL warns if one reappears — so treat a hit from grep -rn '\$schema' etl/steps/viz/chart as something to remove rather than to bump.
One caveat on "leave the pins alone": that holds for pins that are true. A pin that contradicts its own config body — pinned 005 while the config uses chartTypes, which only exists from 006 (the 005→006 migration creates it) — is stale, not a record, and leaving it makes grapher run migrations over a config they were never meant to touch. Check a suspicious pin against the properties of that schema version (curl https://files.ourworldindata.org/schemas/grapher-schema.NNN.json) and correct it to the version the config is actually written against.
© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/sync-grapher-schema of owid/etl.
Open the folder on GitHubat commit bf5dc8e
Sync Grapher Schema next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Sync Grapher Schema this skillowid/etl | 158 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 598 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Monitor With HaolemeHaolemeApp/Haoleme | 157 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | |
| Tushare Plugin BuilderYourdaylight/stock_datasource | 188 | — | ~2.5k | Automated safety check: Pass | MIT | |
| Credit Risk Data Cleaninggithub/awesome-copilot | 40k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Dbt Parser Refreshyu-iskw/dbt-artifacts-parser | 118 | — | ~716 | Automated safety check: Pass | Apache-2.0 |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
HaolemeApp/Haoleme
Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.
Yourdaylight/stock_datasource
Turns a Tushare API doc URL into a full data plugin for the stock_datasource repo: extractor, ClickHouse schema, query service, config and curl examples.
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
yu-iskw/dbt-artifacts-parser
Refreshes dbt artifact schemas from dbt-labs/dbt-core and regenerates Pydantic parser classes.
godatadriven/dbt-bouncer
Analyzes a dbt project and suggests dbt-bouncer checks that already pass (for existing projects) or a sensible starter config (for greenfield projects).
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Works with
Categories
Sync upstream grapher schema changes (new chart types, config fields, enum values) into the ETL repo — vendored schema, multidim-schema, dataset-schema, and regenerated Python types. Sync Grapher Schema is an agent skill from owid/etl. Sync upstream grapher schema changes (new chart types, config fields, enum values) into the ETL repo — vendored schema, multidim-schema, dataset-schema, and regenerated Python types.
Sync Grapher Schema fits situations like: the scheduled sync workflow opened a draft PR; issue that needs completing; the web team announces a grapher schema change (new chart type in Grapher; I added a field to the grapher config).
Run `npx skills add owid/etl --skill sync-grapher-schema -a claude-code`. Or copy the skill folder (.claude/skills/sync-grapher-schema in owid/etl) into .claude/skills/sync-grapher-schema in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill sync-grapher-schema -a codex`. Or copy the skill folder (.claude/skills/sync-grapher-schema in owid/etl) into .agents/skills/sync-grapher-schema in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill sync-grapher-schema -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sync-grapher-schema, .gemini/skills/sync-grapher-schema, .github/skills/sync-grapher-schema and .opencode/skills/sync-grapher-schema in your project.
Going by SKILL.md and its folder, Sync Grapher Schema needs the command-line tools its instructions call (git, python, pytest, make, gh and curl). Our summary lists: Python 3.
SKILL.md names 2 domains. In commands or code: files.ourworldindata.org; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Sync Grapher Schema is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Sync Grapher Schema: Crawl4AI Web Scraping (smallnest/goclaw, 598 stars), Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars), Tushare Plugin Builder (Yourdaylight/stock_datasource, 188 stars) and Credit Risk Data Cleaning (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.