Crawl4AI Web Scraping
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
Author or modify an Our World in Data explorer (multi-dimensional dashboard with dropdown selectors, published from ETL via viz://explorer/<ns/latest/<short).
$ npx skills add owid/etl --skill create-explorer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install owid/etl create-explorer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/create-explorer .claude/skills/create-explorer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "create-explorer" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-explorer into .claude/skills/create-explorer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-explorer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/owid/etl/tree/master/.claude/skills/create-explorerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add owid/etl --skill create-explorer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install owid/etl create-explorer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/create-explorer .agents/skills/create-explorer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "create-explorer" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-explorer into .agents/skills/create-explorer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-explorer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-explorer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install owid/etl create-explorer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/create-explorer .cursor/skills/create-explorer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "create-explorer" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-explorer into .cursor/skills/create-explorer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-explorer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/owid/etl.git --path .claude/skills/create-explorer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add owid/etl --skill create-explorer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install owid/etl create-explorer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/create-explorer .gemini/skills/create-explorer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "create-explorer" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-explorer into .gemini/skills/create-explorer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-explorer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install owid/etl create-explorerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add owid/etl --skill create-explorer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/create-explorer .github/skills/create-explorer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "create-explorer" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-explorer into .github/skills/create-explorer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-explorer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add owid/etl --skill create-explorer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install owid/etl create-explorer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/owid/etl.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/create-explorer .opencode/skills/create-explorer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "create-explorer" agent skill from https://github.com/owid/etl/tree/master/.claude/skills/create-explorer into .opencode/skills/create-explorer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "create-explorer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
create-explorerAuthor or modify an Our World in Data explorer (multi-dimensional dashboard with dropdown selectors, published from ETL via viz://explorer/<ns/latest/<short).
Create Explorer is an agent skill from owid/etl. Author or modify an Our World in Data explorer (multi-dimensional dashboard with dropdown selectors, published from ETL via viz://explorer/<ns/latest/<short). Trigger when the user wants to build a new explorer, add/remove views or dimensions on an existing one, change the explorer's chart text or selection defaults, or finish an explorer migration once the snapshot/garden/grapher chain is already in place.
Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL. The repository describes itself as: A compute graph for loading and transforming OWID's data. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 69ab20e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
assets.ourworldindata.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Create Explorer loads about 7.2k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 2,595 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from owid/etl at commit 69ab20e, republished under its MIT licence (© owid). 2,595 words, ~7,211 tokens.
.claude/skills/create-explorer/SKILL.md (or your agent's skills folder).Explorers are OWID's multi-dimensional dashboards (e.g. ourworldindata.org/explorers/food-prices). They're authored as YAML in this repo and published by ETL at viz://explorer/<ns>/latest/<short>.
This skill is the explorer-flavored sibling of /create-chart. They use the same engine (paths.create_chart for charts and multidims, paths.create_explorer for explorers) and the same YAML schema for dimensions / views / definitions.common_views. The differences are:
| Chart / multidim | Explorer | |
|---|---|---|
| Channel | viz://chart/... | viz://explorer/... |
| Step file location | etl/steps/viz/chart/<ns>/latest/ | etl/steps/viz/explorer/<ns>/latest/ |
| PathFinder method | paths.create_chart(...) | paths.create_explorer(...) |
| Top-level config block | title:, default_selection:, default_dimensions: | config: block carrying legacy explorer settings (explorerTitle, explorerSubtitle, selection, subNavId, entityType, …) |
| Slug convention | underscores in file paths and short_name | underscores in file path, hyphens in URL slug and short_name argument |
| Verification | preview URL on staging | preview URL on staging + diff against owid-grapher/explorers/<slug>.explorer.tsv |
| Save call | c.save() | c.save(tolerate_extra_indicators=True) (upstream grapher datasets usually have more indicators than the explorer references) |
If you're modifying an existing explorer (adjusting chart text, swapping a catalogPath, adding a dimension choice, reordering views), most of the deeper sections below don't apply — find the existing <short>.config.yml, edit, run etlr, done. The full structure is documented below for new explorers and substantial reshapes.
/migrate-explorer-to-etl skill has produced (or already located) the upstream snapshot/meadow/garden/grapher chain, and now needs the explorer step.Every explorer is exactly two files plus a DAG entry:
etl/steps/viz/explorer/<ns>/latest/
├── <short>.py # Python uses snake_case
└── <short>.config.ymlmkdir -p etl/steps/viz/explorer/<ns>/latestHyphens vs underscores (recurring source of confusion):
food_footprints.py, crop_yields.py.short_name= argument keeps hyphens: food-footprints, crop-yields.paths.create_explorer(short_name="<short-with-hyphens>") — pass the hyphenated slug.The two ends of the spectrum, plus everything in between:
| Full-YAML | Programmatic / table-driven | |
|---|---|---|
| Where views live | Hand-listed in <short>.config.yml under views: | Auto-expanded by paths.create_explorer(tb=tb, ...) from columns whose m.dimensions is set |
| Where chart text (FAUST) lives | Per-view view.config.{title, subtitle, note} in YAML | presentation.{title_public, grapher_config} on each indicator's garden metadata; common defaults via definitions.common.presentation.grapher_config |
Where chart-level config lives (hasMapTab, tab, yAxis, chartTypes) | Per-view view.config | Indicator's presentation.grapher_config — single source of truth, same as for any standalone chart on that indicator |
| Map color scale | view.indicators.y[i].display.{colorScaleScheme, colorScaleNumericBins} (semicolon-string form, explorer-flavored override) | presentation.grapher_config.map.colorScale.{baseColorScheme, binningStrategy, customNumericValues} (canonical grapher form, inherited at chart render time) |
| Python step content | Trivial: paths.create_explorer(config=config, short_name=...).save(...) | Loops columns to set m.dimensions, optionally post-processes (sort_choices, group_views, per-view display tweaks) |
It's a spectrum, not a switch. Mix freely: use table-driven for the bulk of views, hand-list a handful of bespoke ones; or stay full-YAML but still push title/subtitle for the single-indicator views into garden metadata to remove duplication.
Strong fit for table-driven:
create_explorer(tb=tb, indicator_names=..., dimensions=...) matches this shape directly — model: migration/latest/migration_flows.py.Stick with full-YAML when:
c.group_views(...))."""<one-line description of what this explorer surfaces>."""
from etl.helpers import PathFinder
paths = PathFinder(__file__)
def run() -> None:
config = paths.load_config()
c = paths.create_explorer(
config=config,
short_name="<short-with-hyphens>", # explorer slug
)
c.save(tolerate_extra_indicators=True)tolerate_extra_indicators=True is the common case: the upstream grapher dataset usually carries more indicators than the explorer references, and without this flag c.save() errors on the unused ones.
"""<one-line description>."""
from etl.helpers import PathFinder
paths = PathFinder(__file__)
# Map column → dimension tuple. "na" is the conventional empty slot for conditional dimensions
# (e.g. cost_metric is meaningful only when type=cost; affordability views set cost_metric="na").
COLUMN_DIMENSIONS: dict[str, dict[str, str]] = {
"<col_a>": {"dim1": "value_a1", "dim2": "value_a2"},
"<col_b>": {"dim1": "value_b1", "dim2": "value_b2"},
# ...
}
def run() -> None:
config = paths.load_config()
ds = paths.load_dataset("<grapher_dataset>")
tb = ds.read("<table>", load_data=False) # metadata only — faster, we don't need values
for column, dims in COLUMN_DIMENSIONS.items():
tb[column].m.dimensions = dims
tb[column].m.original_short_name = "<unifying_indicator_name>"
c = paths.create_explorer(
config=config,
tb=tb,
indicator_names=["<unifying_indicator_name>"],
dimensions={
"dim1": ["value_a1", "value_b1", ...], # explicit choice order
"dim2": ["value_a2", "value_b2", ...],
},
# common_view_config={...}, # only if not in indicator metadata
short_name="<short-with-hyphens>",
)
# Optional post-processing — see "Post-processing" below.
# c.sort_choices({"dim1": lambda x: sorted(x)})
# c.group_views([...])
c.save(tolerate_extra_indicators=True)Key APIs (see etl/viz/chart/core/expand.py and etl/viz/chart/core/create.py):
tb[col].m.dimensions: dict[str, str] — required per column. Each entry says "this column represents the (dim1=value, dim2=value) cell." Columns without m.dimensions are ignored by the expander.tb[col].m.original_short_name: str — the unifying indicator name. With indicator_names=[that_name] and a single name, the expander treats all N columns as one logical indicator with N dimension combinations and drops the auto-added "indicator" pseudo-dimension.dimensions= accepts:None → all dimensions found, arbitrary order.list[str] → restricts and orders dimensions, all values shown.dict[str, list[str] | "*"] → restricts and orders both dimensions and choices. Use "*" for "all values, arbitrary order."common_view_config= is applied uniformly to every auto-expanded view. Use it for fields that are truly shared and don't live at indicator level. Prefer indicator-level presentation.grapher_config for anything that should also flow to standalone charts.Always block style. Mappings and lists in explorer config YAML must use block style — one key per line, list items on their own line under
-. Never use flow style ({ key: value, ... }or[a, b, c]) even for tiny per-viewdimensions:blocks. PR review on a 45-view file is unreadable when half the views collapse to a single flow line. The only exception is markdown links inside a quoted-scalarsubtitle:/note:(those[text](url)brackets are content, not YAML structure).
config:
# Explorer settings rows — keys map verbatim from the legacy TSV settings section.
explorerTitle: ...
explorerSubtitle: ...
isPublished: true
hasMapTab: false
hideAlertBanner: true
hideAnnotationFieldsInTitle: true
entityType: country # or "food", "region", etc.
thumbnail: https://assets.ourworldindata.org/uploads/...
wpBlockId: "12345"
subNavId: explorers
subNavCurrentId: <slug>
selection:
- <default selected entity>
- <another>
pickerColumnSlugs: [] # an empty list is OK; non-empty must be block-style
yAxisMin: 0
# ...
definitions:
# Shared config applied to all views. Use this list (with optional `dimensions:`
# filter per entry) — NOT YAML anchors and `<<:` merge keys. The framework merges
# entries at expansion time; per-view `config:` blocks override anything here.
common_views:
- config:
type: DiscreteBar
hasMapTab: false
# Dimension-filtered overrides apply only to matching views:
# - dimensions:
# metric: share
# config:
# note: "Share values sum to 100%"
dimensions:
# one entry per dropdown / radio / checkbox the user toggles
- slug: <snake_case> # e.g. "metric"
name: <human label> # e.g. "Metric"
presentation:
type: dropdown # or radio / checkbox
choices:
- slug: <choice_snake>
name: "<as shown in widget>"
- slug: <another>
name: "..."
views:
# one entry per (dim1=x, dim2=y, …) tuple
- dimensions:
<dim_slug>: <choice_slug>
# ...
indicators:
y:
- catalogPath: <table>#<short> # short form — see "catalogPath — short forms accepted" below
display: # per-view, per-indicator overrides
colorScaleNumericBins: 0;1;2
colorScaleScheme: PuBu
config:
# Only per-view overrides here. Common stuff lives in definitions.common_views.
# No `<<:` merge keys, no `&anchor`s.
title: ...
subtitle: ...
type: <chart type> # LineChart, DiscreteBar, "LineChart DiscreteBar", StackedArea, …
hasMapTab: false
minTime: 1990
yAxisMin: 0For table-driven explorers, views: should still be present but is typically views: [] — the explorer JSON schema requires the key, and create_explorer(tb=tb, ...) populates the views at runtime.
catalogPath — short forms acceptedThe Indicator.is_a_valid_path check (etl/viz/chart/model/view.py:62) accepts three forms; pick the shortest one that still unambiguously resolves:
| Form | Example | When to use |
|---|---|---|
table#indicator | global_carbon_budget#emissions_total | Default. Resolved against the explorer's DAG dependencies via tables_by_name — fine as long as no two dependencies expose a table with the same name. |
dataset/table#indicator | global_carbon_budget/global_carbon_budget#emissions_total | When two upstream datasets happen to expose tables with the same short_name. |
grapher/<ns>/<v>/<dataset>/<table>#<indicator> | grapher/gcp/2025-11-13/global_carbon_budget/global_carbon_budget#emissions_total | Only when you need to pin a specific dataset version separate from the one in the DAG — almost never the right form to write by hand. |
Short forms are expanded at c.save() time by Indicator.expand_path(tables_by_name). If the table name doesn't exist in any dependency it raises Table name '<x>' not found in dependency tables; if multiple dependencies expose the same table name, it raises and asks you to disambiguate with the medium form.
Default to table#indicator when authoring YAML. The full path is verbose, drifts when upstream versions bump, and is only needed for genuinely ambiguous cases.
config: settings — the most common keys| Key | Type | Notes |
|---|---|---|
explorerTitle | string | Page title above the explorer. |
explorerSubtitle | string | One-liner under the title. |
isPublished | bool | true to publish; false keeps it draft. |
hasMapTab | bool | Whether any view shows the map tab by default. |
entityType | string | country (default), or food, region, species, etc. — controls picker labels. |
selection | list[str] | Default selected entities. |
pickerColumnSlugs | list[str] | Picker columns shown alongside the entity name. |
subNavId | string | Almost always explorers. |
subNavCurrentId | string | The slug — appears as the active nav item. |
wpBlockId | string | WordPress block ID for embedding (legacy). Stringify even when numeric. |
thumbnail | string | URL of preview image. |
hideAlertBanner | bool | Suppress the OWID-wide banner. |
hideAnnotationFieldsInTitle | bool | Drop time/entity from auto-titles. |
yAxisMin | number/string | Default Y-axis floor. |
yScaleToggle | bool | Allow user to toggle linear/log. |
originUrl | string | Path back to the topic page (e.g. /environmental-impacts-of-food). |
dropdown: shown as <select>. Use for >4 choices or when the choices have long labels.radio: shown as a row of pills. Use for ≤4 mutually-exclusive choices.checkbox: shown as a single toggle. Two choices only — usually "off" (slug like combined/absolute/no) and "on" (slug like the field name). Pair with presentation.choice_slug_true: <on_slug> so the framework knows which slug means "checked."- slug: by_stage
name: By stage of supply chain
presentation:
type: checkbox
choice_slug_true: stages
choices:
- slug: combined
name: ""
- slug: stages
name: By stage of supply chainWhen a dimension is only meaningful for some rows (e.g. cost_metric matters only when type=cost, not when type=affordability), include the dimension everywhere with an "na" slot:
<dim>: "na".na choice with name: "" so the widget renders as empty when applicable.agriculture/latest/food_prices.{py,config.yml}.- slug: cost_metric
name: Cost metric
presentation:
type: radio
choices:
- slug: na
name: ""
- slug: dollars_per_day
name: $ per dayFor single-indicator views, the rendered chart inherits the indicator's stored grapher_config from MySQL at render time (both standalone-chart and explorer-view paths). Push:
| Per-view config | → indicator garden metadata |
|---|---|
title | presentation.grapher_config.title |
subtitle | presentation.grapher_config.subtitle |
note | presentation.grapher_config.note |
map.colorScale | presentation.grapher_config.map.colorScale.{baseColorScheme, binningStrategy, customNumericValues} |
hasMapTab, tab, yAxis, chartTypes, hideRelativeToggle, selectedFacetStrategy, … | presentation.grapher_config.<field> |
For the chart-heading flow specifically, the priority is grapher_config.title > title_public > display.name > title (see docs/architecture/metadata/faqs.md). When migrating a chart-wrapping explorer, the chart's bespoke heading text belongs in grapher_config.title — title_public is the human-readable replacement for a dimensional indicator.title, not the chart's heading.
Cross-cutting baselines (e.g. hasMapTab: true for every indicator in a dataset) go under definitions.common.presentation.grapher_config in the garden .meta.yml. The catalog merge is recursive on presentation and grapher_config (lib/catalog/owid/catalog/core/yaml_metadata.py:_merge_variable_metadata), so each indicator inherits the common defaults plus its own overrides without manual <<: *anchor repetition.
DRY for repeated text fragments via dynamic-yaml interpolation:
definitions:
prefix: "Long shared phrase about diet X."
suffix: "Common closing sentence about methodology."
# In the indicator block:
subtitle: "{definitions.prefix} {definitions.suffix}"This composes N unique full strings from a handful of building blocks. Verified via dynamic_yaml_to_dict (lib/catalog/owid/catalog/core/utils.py).
After paths.create_explorer() returns the explorer c, you can mutate it before c.save():
c.sort_choices({dim_slug: lambda x: sorted(x)}) — control the order of dimension dropdowns. Useful when slugs sort poorly alphabetically.
c.group_views(groups=[...]) — bundle multiple existing views into a new multi-indicator view. Each entry in groups:
dimension: the dimension whose choices are being collapsed.choices: the choice slugs to combine (omit for all).choice_new_slug: name for the new collapsed choice (e.g. combined, total, breakdown).view_config: chart-level config for the new view (chartTypes, title, subtitle, selectedFacetStrategy, …). Title can be a template like "Population aged {age}" evaluated against params.view_metadata: data-page metadata (description_key, etc.) for the new view.replace=True to drop the originals; default keeps both.overwrite_dimension_choice=True if choice_new_slug collides with an existing choice and you want grouped views to win.sex={female, male} views — group_views adds sex=combined showing both timeseries on one chart. Same pattern works for age brackets, region groups, conflict types, or any dimension where users may want a single multi-line chart.c.edit_views([...]) — apply chart-level config to many views at once, optionally scoped by dimension. Each entry is {"dimensions": <filter>, "config": {...}, "metadata": {...}}; the framework merges entries by specificity (more dimensions in the filter = wins on conflicts). Use this in preference to set_global_config whenever you have per-slice overrides:
c.edit_views([
# No filter → applies to every view (defaults).
{"config": {"type": "LineChart DiscreteBar", "hasMapTab": True}},
# Scoped override — only this exact (gas, accounting, fuel, count) cell gets a Slope tab.
{
"dimensions": {"gas": "co2", "accounting": "territorial", "fuel": "all_fossil", "count": "per_capita"},
"config": {"type": "LineChart SlopeChart DiscreteBar"},
},
]) Callable values inside config (e.g. "title": lambda v: ...) are evaluated against each matching view, so dimension-aware text templates still work. set_global_config is just a one-entry shortcut for edit_views; reach for edit_views once you have more than one slice to address.
Manual loop over c.views — for things edit_views can't reach: per-indicator display blocks (numDecimalPlaces, colorScaleScheme, colorScaleNumericBins, display.color, …). These live on view.indicators.y[i].display, not on view.config. Match views via the view.matches(**kwargs) helper, which accepts a single value or a list (list = OR semantics):
for view in c.views:
if view.matches(gas=["methane", "all_ghg"], count="per_capita"):
decimals = 1
elif view.matches(gas="warming_impact", fuel=["land_use", "fossil_plus_land_use"]):
decimals = 3
else:
continue
for indicator in view.indicators.y or []:
indicator.display = {**(indicator.display or {}), "numDecimalPlaces": decimals} Pattern: migration_flows.py's add_display_settings(c). Avoid this loop when the setting can live on the indicator's garden metadata instead.
choice_renames={dim: {slug: display_name, ...}} (passed directly to create_explorer) — map slug → display name when you need to derive the display label programmatically. Model: chart/minerals/latest/minerals.py.
Sidecar <short>.dims.yaml — when the column → dimensions map exceeds ~50 entries, lift it out of the Python step into a sidecar YAML loaded at module-import time. Keeps <short>.py focused on logic and turns dim-tagging changes into a 1-line YAML edit. Model: etl/steps/viz/explorer/emissions/latest/co2.{py,dims.yaml}:
from pathlib import Path
import yaml
COLUMN_DIMENSIONS = yaml.safe_load((Path(__file__).parent / "co2.dims.yaml").read_text())In dag/<ns>.yml:
viz://explorer/<ns>/latest/<short>:
- data://grapher/<ns1>/<v1>/<dataset1>
- data://grapher/<ns2>/<v2>/<dataset2>
# ... one line per unique upstream grapher datasetPlace near related explorer entries (or alongside the upstream grapher steps) for discoverability.
Hand off to the user:
.venv/bin/etlr viz://explorer/<ns>/latest/<short> --grapher — runs the step and upserts the explorer to the staging DB.http://staging-site-<branch>/admin/explorers/preview/<slug> and spot-check:owid-grapher/explorers/<slug>.explorer.tsv. Cosmetic differences (column ordering, whitespace) are acceptable; structural differences (missing views, swapped dimension orderings) are not. The Wizard's apps/wizard/app_pages/explorer_diff/ page does this comparison interactively for staging vs production.make check.Hand the user the exact etlr command — don't run it yourself.
build_views does not propagate per-view display from indicator metadata — see TODO at etl/viz/chart/core/expand.py:313. Each auto-expanded view gets indicators.y[0] with only catalogPath, no display. If you need different color scales per view, either (a) put them in the indicator's presentation.grapher_config.map.colorScale so they apply at chart render time, or (b) post-process c.views in Python.type: LineChart is not a valid grapher_config field when authoring via indicator metadata — use chartTypes: ["LineChart"] (the schema is an array). Per-view config.type in the explorer YAML still accepts strings.tb.read(..., load_data=False) is essential when you only need column metadata to set dimensions; loading data unnecessarily slows the step.views: key in the explorer config even when empty. Pass views: [].grapher_config before the explorer view can inherit it. Re-run etlr --grapher data://grapher/... after editing garden metadata; the explorer step alone won't refresh the indicator's stored config.&common_view, <<: *common_view) — don't use them. They can't filter by dimension and add per-view noise. Use definitions.common_views instead.short_name — file is food_footprints.py but short_name="food-footprints". Mismatch produces a published explorer whose URL doesn't match the legacy slug.tolerate_extra_indicators — usually want True for explorers since you're cherry-picking indicators from larger upstream datasets.na pattern (3 choices: na, off, on) works for radio and dropdown but not checkbox — the framework rejects it with Dimension choices for 'checkbox' must have exactly two choices. If you genuinely need a third "doesn't apply" state, either use a radio with three choices, or drop the na slot and let the framework hide/disable the toggle when the current dimension state has no matching <true> view.no/yes/on/off/true/false slugs. A view.dimensions: { relative_to_world: no } reads back as {"relative_to_world": False}, which won't match the string slug "no" declared in dimensions[*].choices. Use semantic slugs (total/relative, absolute/share) or quote the strings — but choosing different slug names is the more durable fix.edit_views doesn't reach indicator-level display. Fields like numDecimalPlaces, colorScaleScheme, colorScaleNumericBins, and per-indicator color live on view.indicators.y[i].display, not on view.config. edit_views only writes view-level config/metadata. For these, fall back to a c.views loop (see Step 6) or push them into the indicator's garden presentation.grapher_config so they apply at chart render time.Full-YAML (each view hand-listed):
etl/steps/viz/explorer/agriculture/latest/crop_yields.{py,config.yml} — large indicator-based explorer with many dimensions.etl/steps/viz/explorer/agriculture/latest/food_prices.{py,config.yml} — small grapher-chart-based migration (12 chart IDs unwrapped to 12 single-indicator views) with conditional dimensions ("na" pattern).etl/steps/viz/explorer/agriculture/latest/fertilizers.{py,config.yml} — checkbox dimension with choice_slug_true; multi-namespace dependencies.etl/steps/viz/explorer/food/latest/food_footprints.{py,config.yml} — hybrid (16 grapher-chart views + 29 CSV-backed views) showing dimension-filtered common_views for differing sourceDesc per view-type.etl/steps/viz/explorer/war/latest/countries_in_conflict_data.{py,config.yml} — uses na-named choices to model conditional dimensions.etl/steps/viz/explorer/emissions/latest/ipcc_scenarios.{py,config.yml} — moderate-size YAML-driven.Table-driven (views auto-expanded from a dimensional table):
etl/steps/viz/explorer/migration/latest/migration_flows.{py,config.yml} — passes tb=tb, indicator_names=[...], dimensions=[...] to create_explorer; YAML carries only the static config and dimension presentation. Includes add_display_settings(c) post-processing.etl/steps/viz/explorer/emissions/latest/co2.{py,dims.yaml,config.yml} — sidecar .dims.yaml for the 54-entry column→dimensions map; uses c.edit_views([...]) with both an unscoped default and a 4-dim-filtered override; uses a c.views loop with view.matches(...) for per-indicator numDecimalPlaces overrides.etl/steps/viz/explorer/emissions/latest/air_pollution.{py,config.yml} — table-driven with c.group_views(...) to add facet views, c.drop_views(...) to prune cross-products.etl/steps/viz/chart/minerals/latest/minerals.py — same APIs in the multidim channel; useful read for choice_renames.Once an explorer is on paths.create_explorer(), it's a candidate for the Track-B port to MDIM (viz://chart/...) once feature parity is reached. See umbrella issue #6014.
© owid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/create-explorer of owid/etl.
Open the folder on GitHubat commit 69ab20e
Create Explorer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Create Explorer this skillowid/etl | 158 | — | ~7.2k | Automated safety check: Pass | MIT | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 598 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Glue 09 10 Migrationaws-samples/aws-glue-samples | 1.5k | — | ~2.4k | Automated safety check: Pass | MIT-0 | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Dbt Databricks PR Readydatabricks/dbt-databricks | 379 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Mz Dbt ReleaseMaterializeInc/materialize | 6.4k | — | ~1.2k | Automated safety check: Pass | Custom licence |
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Upgrade an AWS Glue ETL job from Glue version 0.9 or 1.0 to Glue 4.0.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
databricks/dbt-databricks
A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.
MaterializeInc/materialize
Cut a dbt-materialize PyPI release: bump the version in version.py and setup.py, date the Unreleased CHANGELOG entry, and open the release PR with a Ship: <url body.
liam-machine/erd-studio
Friendly, step-by-step setup for ERD Studio in an existing dbt project, for people who may be new to dbt or data modelling.
owid/etl
Find every OWID surface that references a chart, indicator, MDIM, or explorer — articles (links vs embeds), explorers, narrative charts, data insights, static viz, key-chart slots, MDIM views.
owid/etl
Add a scatter view (with GDP per capita on x) to existing OWID charts via the admin API, mirroring the admin UI's "Add scatter type" defaults, then retire the old standalone "X vs.
owid/etl
Add new survey question codes (e.g. An agent skill from owid/etl.
owid/etl
Build or refresh an OWID static visualization end to end — resolve what data it needs from an old static viz image, an indicator, or a grapher chart; check both the ETL catalog and the producer's…
owid/etl
Propose redirects from (soon-to-sunset) grapher charts to the matching views of published MDIMs.
owid/etl
Take (soon-to-sunset) OWID explorers to redirected MDIMs, end to end.
Categories
Author or modify an Our World in Data explorer (multi-dimensional dashboard with dropdown selectors, published from ETL via viz://explorer/<ns/latest/<short). Create Explorer is an agent skill from owid/etl. Author or modify an Our World in Data explorer (multi-dimensional dashboard with dropdown selectors, published from ETL via viz://explorer/<ns/latest/<short).
Create Explorer fits situations like: the user wants to build a new explorer; add/remove views; dimensions on an existing one; change the explorers chart text.
Run `npx skills add owid/etl --skill create-explorer -a claude-code`. Or copy the skill folder (.claude/skills/create-explorer in owid/etl) into .claude/skills/create-explorer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add owid/etl --skill create-explorer -a codex`. Or copy the skill folder (.claude/skills/create-explorer in owid/etl) into .agents/skills/create-explorer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add owid/etl --skill create-explorer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-explorer, .gemini/skills/create-explorer, .github/skills/create-explorer and .opencode/skills/create-explorer in your project.
Going by SKILL.md and its folder, Create Explorer needs the command-line tools its instructions call (make). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: assets.ourworldindata.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Create Explorer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Create Explorer: Crawl4AI Web Scraping (smallnest/goclaw, 598 stars), Glue 09 10 Migration (aws-samples/aws-glue-samples, 1.5k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Dbt Databricks PR Ready (databricks/dbt-databricks, 379 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
owid (a GitHub organization) maintains it in owid/etl, which has 158 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 7, 2026.
Source: owid/etl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.