Migrating Dagster To Airflow
astronomer/agents
Guide for migrating Dagster projects to Apache Airflow 3 on Astro.
A skill your agent uses when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating…
$ npx skills add kid-sid/claude-spellbook --skill data-pipelines -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install kid-sid/claude-spellbook data-pipelines --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-pipelines .claude/skills/data-pipelines && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-pipelines" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/data-pipelines into .claude/skills/data-pipelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipelines", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/kid-sid/claude-spellbook/tree/main/skills/data-pipelinesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add kid-sid/claude-spellbook --skill data-pipelines -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install kid-sid/claude-spellbook data-pipelines --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/data-pipelines .agents/skills/data-pipelines && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-pipelines" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/data-pipelines into .agents/skills/data-pipelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipelines", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kid-sid/claude-spellbook --skill data-pipelines -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install kid-sid/claude-spellbook data-pipelines --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/data-pipelines .cursor/skills/data-pipelines && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-pipelines" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/data-pipelines into .cursor/skills/data-pipelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipelines", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/kid-sid/claude-spellbook.git --path skills/data-pipelines--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add kid-sid/claude-spellbook --skill data-pipelines -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install kid-sid/claude-spellbook data-pipelines --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/data-pipelines .gemini/skills/data-pipelines && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-pipelines" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/data-pipelines into .gemini/skills/data-pipelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipelines", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install kid-sid/claude-spellbook data-pipelinesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add kid-sid/claude-spellbook --skill data-pipelines -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/data-pipelines .github/skills/data-pipelines && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-pipelines" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/data-pipelines into .github/skills/data-pipelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipelines", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add kid-sid/claude-spellbook --skill data-pipelines -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install kid-sid/claude-spellbook data-pipelines --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/data-pipelines .opencode/skills/data-pipelines && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-pipelines" agent skill from https://github.com/kid-sid/claude-spellbook/tree/main/skills/data-pipelines into .opencode/skills/data-pipelines/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipelines", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-pipelinesA skill your agent uses when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating…
Data Pipelines is an agent skill from kid-sid/claude-spellbook. Use when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating data quality, or orchestrating multi-step data workflows.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL. It works with dbt and Apache Airflow. The repository describes itself as: A curated collection of skills, prompts, and workflows that extend Claude's capabilities — your personal grimoire for AI-powered development. The licence is MIT.
Read from SKILL.md and the folder at commit a7c2ac9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dbtairflowFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Pipelines loads about 4.3k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 656 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from kid-sid/claude-spellbook at commit a7c2ac9, republished under its MIT licence (© kid-sid). 656 words, ~4,295 tokens.
.claude/skills/data-pipelines/SKILL.md (or your agent's skills folder).Orchestration, transformation, and validation patterns for production data pipelines.
| Approach | Transform where | Use when |
|---|---|---|
| ETL | Before loading (in pipeline code) | Target warehouse has limited compute; PII must be masked before storage |
| ELT | After loading (in warehouse SQL) | Modern warehouse (BigQuery, Snowflake, Redshift); raw data must be preserved |
| Streaming | Continuously (Kafka + Flink/Spark) | Sub-minute latency required; event sourcing |
Default for modern stacks: ELT — land raw data, transform with dbt, version-control SQL.
from datetime import datetime, timedelta
from airflow.decorators import dag, task
from airflow.operators.python import PythonOperator
from airflow.providers.postgres.hooks.postgres import PostgresHook
@dag(
schedule="0 6 * * *", # 6 AM daily
start_date=datetime(2026, 1, 1),
catchup=False, # don't backfill missed runs on deploy
max_active_runs=1, # prevent overlapping runs
default_args={
"retries": 3,
"retry_delay": timedelta(minutes=5),
"retry_exponential_backoff": True,
"email_on_failure": True,
},
tags=["finance", "daily"],
)
def daily_revenue_pipeline():
@task
def extract_orders(execution_date=None) -> list[dict]:
hook = PostgresHook(postgres_conn_id="source_db")
# Use execution_date for idempotent extraction
rows = hook.get_records(
"SELECT * FROM orders WHERE date = %s",
parameters=[execution_date.date()],
)
return [dict(r) for r in rows]
@task
def transform(orders: list[dict]) -> list[dict]:
return [
{**o, "revenue_usd": o["amount"] * o["fx_rate"]}
for o in orders
if o["status"] == "completed"
]
@task
def load(records: list[dict], execution_date=None):
hook = PostgresHook(postgres_conn_id="warehouse")
# Idempotent: delete-then-insert for the partition date
hook.run("DELETE FROM daily_revenue WHERE date = %s", parameters=[execution_date.date()])
hook.insert_rows("daily_revenue", [[r["date"], r["revenue_usd"]] for r in records])
orders = extract_orders()
transformed = transform(orders)
load(transformed)
dag = daily_revenue_pipeline()from airflow.operators.bash import BashOperator
from airflow.operators.python import BranchPythonOperator
from airflow.sensors.filesystem import FileSensor
from airflow.sensors.sql import SqlSensor
from airflow.providers.http.sensors.http import HttpSensor
# Wait for a file to appear (S3, GCS, local)
wait_for_export = FileSensor(
task_id="wait_for_export",
filepath="/data/exports/{{ ds }}/orders.csv",
poke_interval=60, # check every 60s
timeout=3600, # fail after 1 hour
mode="reschedule", # release worker slot while waiting
)
# Wait for upstream table to be populated
wait_for_source = SqlSensor(
task_id="wait_for_orders",
conn_id="source_db",
sql="SELECT COUNT(*) FROM orders WHERE date = '{{ ds }}' HAVING COUNT(*) > 0",
poke_interval=120,
mode="reschedule",
)
# Branch: skip load on weekends
def should_load(**context):
if context["execution_date"].weekday() >= 5:
return "skip_load"
return "load"
branch = BranchPythonOperator(task_id="check_day", python_callable=should_load)# Push value
@task
def extract() -> dict:
return {"row_count": 1042, "checksum": "abc123"} # return value auto-pushes XCom
# Pull value
@task
def validate(stats: dict): # passed as argument from task dependency
assert stats["row_count"] > 0, "Empty extract"
# Manual XCom pull (classic operators)
def load(**context):
stats = context["task_instance"].xcom_pull(task_ids="extract")
print(stats["row_count"])XCom limits: XComs are stored in the Airflow metadata DB — not suited for large data. Pass row counts, checksums, and file paths through XComs; never entire datasets.
@task
def get_regions() -> list[str]:
return ["us-east", "eu-west", "ap-south"]
@task
def process_region(region: str):
extract_and_load(region)
# Creates one task instance per region — parallelized automatically
process_region.expand(region=get_regions())from airflow.hooks.base import BaseHook
from airflow.models import Variable
# Never hardcode credentials — use Connections
conn = BaseHook.get_connection("my_postgres")
dsn = f"postgresql://{conn.login}:{conn.password}@{conn.host}/{conn.schema}"
# Runtime config — use Variables (or better: Airflow Params)
batch_size = int(Variable.get("etl_batch_size", default_var=1000))dbt_project/
├── models/
│ ├── staging/ # stg_* — raw → typed, renamed, deduplicated
│ │ └── stg_orders.sql
│ ├── intermediate/ # int_* — business logic joins
│ │ └── int_order_items.sql
│ └── marts/ # final — wide tables for BI/downstream
│ └── fct_revenue.sql
├── tests/ # custom SQL tests
├── macros/ # Jinja macros
├── seeds/ # static CSV reference data
└── dbt_project.yml-- staging/stg_orders.sql
-- Materialization: view (cheap, always fresh)
{{ config(materialized='view') }}
SELECT
order_id::VARCHAR AS order_id,
user_id::VARCHAR AS user_id,
created_at::TIMESTAMP AS created_at,
amount_cents / 100.0 AS amount_usd,
status
FROM {{ source('raw', 'orders') }}
WHERE status != 'test'-- marts/fct_revenue.sql
-- Materialization: table (fast reads, rebuilt on each run)
{{ config(materialized='table') }}
SELECT
DATE_TRUNC('day', o.created_at) AS date,
p.name AS product_name,
SUM(oi.quantity) AS units_sold,
SUM(oi.quantity * oi.unit_price) AS revenue_usd
FROM {{ ref('stg_orders') }} o -- ref() creates dependency
JOIN {{ ref('int_order_items') }} oi ON o.order_id = oi.order_id
JOIN {{ ref('stg_products') }} p ON oi.product_id = p.product_id
WHERE o.status = 'completed'
GROUP BY 1, 2-- Only process new/updated rows — essential for large tables
{{ config(
materialized='incremental',
unique_key='order_id',
incremental_strategy='merge', -- or 'delete+insert', 'insert_overwrite'
on_schema_change='append_new_columns',
) }}
SELECT
order_id,
user_id,
amount_usd,
created_at,
updated_at
FROM {{ source('raw', 'orders') }}
{% if is_incremental() %}
-- Only load rows newer than the last run
WHERE updated_at > (SELECT MAX(updated_at) FROM {{ this }})
{% endif %}# models/staging/sources.yml
version: 2
sources:
- name: raw
database: analytics
schema: raw_data
freshness:
warn_after: {count: 6, period: hour}
error_after: {count: 24, period: hour}
loaded_at_field: _loaded_at # column that holds ingestion timestamp
tables:
- name: orders
description: Raw orders from the transactional database
- name: products# Check source freshness in CI
dbt source freshness# models/staging/stg_orders.yml
version: 2
models:
- name: stg_orders
columns:
- name: order_id
tests:
- not_null
- unique
- name: status
tests:
- accepted_values:
values: ["pending", "completed", "cancelled", "refunded"]
- name: user_id
tests:
- not_null
- relationships:
to: ref('stg_users')
field: user_id
- name: amount_usd
tests:
- not_null
- dbt_utils.accepted_range:
min_value: 0
max_value: 100000-- tests/assert_revenue_non_negative.sql — custom SQL test (fails if rows returned)
SELECT date, revenue_usd
FROM {{ ref('fct_revenue') }}
WHERE revenue_usd < 0-- macros/cents_to_dollars.sql
{% macro cents_to_dollars(column_name) %}
({{ column_name }} / 100.0)::NUMERIC(10, 2)
{% endmacro %}
-- Usage in a model
SELECT {{ cents_to_dollars('amount_cents') }} AS amount_usd-- macros/generate_surrogate_key.sql (or use dbt_utils)
{% macro surrogate_key(fields) %}
MD5(CONCAT_WS('|', {% for f in fields %}COALESCE(CAST({{ f }} AS VARCHAR), ''){% if not loop.last %}, {% endif %}{% endfor %}))
{% endmacro %}dbt run # run all models
dbt run --select staging # run a directory
dbt run --select stg_orders+ # run model and all downstream
dbt run --select +fct_revenue # run model and all upstream
dbt test # run all tests
dbt test --select stg_orders # test one model
dbt build # run + test in dependency order
dbt source freshness # check source data freshness
dbt docs generate && dbt docs serve # generate + serve lineage docs
dbt compile # render SQL without runningA pipeline run is idempotent if running it twice produces the same result as running it once.
# GOOD: delete-then-insert for a known partition
def load_partition(date: str, records: list[dict]):
with engine.begin() as conn:
conn.execute(
text("DELETE FROM daily_stats WHERE date = :date"),
{"date": date}
)
conn.execute(insert(DailyStats), records)
# GOOD: UPSERT (merge) on unique key
def upsert_orders(records: list[dict]):
stmt = pg_insert(orders_table).values(records)
stmt = stmt.on_conflict_do_update(
index_elements=["order_id"],
set_={"status": stmt.excluded.status, "updated_at": stmt.excluded.updated_at}
)
with engine.begin() as conn:
conn.execute(stmt)
# BAD: append-only — reruns duplicate data
def load_orders(records):
engine.execute(insert(orders_table).values(records)) # duplicates on rerunAirflow idempotency: Use {{ ds }} (execution date, not run date) in all queries. Two runs for the same ds must produce the same output.
| Strategy | How | Use When |
|---|---|---|
| Full refresh | Truncate + reload entire table | Small tables (<1M rows), no CDC |
| Incremental by timestamp | WHERE updated_at > last_run_max | Source has reliable updated_at |
| Incremental by partition | Process one date partition per run | Append-only event data |
| CDC (change data capture) | Debezium → Kafka → warehouse | High-volume, low-latency, soft deletes |
| Snapshot | dbt snapshot (strategy: timestamp) | Track slowly-changing dimensions |
# Watermark-based incremental (Python)
def get_watermark(conn, table: str) -> datetime:
row = conn.execute(
text("SELECT COALESCE(MAX(updated_at), '1970-01-01') FROM :table", bindparams=[bindparam("table")])
).fetchone()
return row[0]
def extract_incremental(source_conn, watermark: datetime) -> list[dict]:
return source_conn.execute(
text("SELECT * FROM orders WHERE updated_at > :wm ORDER BY updated_at"),
{"wm": watermark},
).fetchall()# Generic tests: not_null, unique, accepted_values, relationships
# Package tests: dbt_utils, dbt_expectations (Great Expectations style)
- name: amount_usd
tests:
- dbt_expectations.expect_column_values_to_be_between:
min_value: 0
max_value: 50000
row_condition: "status = 'completed'"import great_expectations as gx
context = gx.get_context()
suite = context.add_expectation_suite("orders_suite")
validator = context.get_validator(
batch_request=batch_request,
expectation_suite_name="orders_suite",
)
validator.expect_column_values_to_not_be_null("order_id")
validator.expect_column_values_to_be_unique("order_id")
validator.expect_column_values_to_be_between("amount_usd", min_value=0)
validator.expect_column_pair_values_A_to_be_greater_than_B(
"completed_at", "created_at"
)
results = validator.validate()
if not results.success:
raise ValueError(f"Data quality check failed: {results}")@task
def reconcile(source_count: int, target_count: int, tolerance: float = 0.001):
delta = abs(source_count - target_count) / max(source_count, 1)
if delta > tolerance:
raise ValueError(
f"Row count mismatch: source={source_count}, target={target_count}, "
f"delta={delta:.2%} > {tolerance:.2%} tolerance"
)# Airflow: SLA miss callback
def sla_miss_callback(dag, task_list, blocking_task_list, slas, blocking_tis):
send_slack_alert(f"SLA missed for DAG {dag.dag_id}: {task_list}")
@dag(sla_miss_callback=sla_miss_callback)
def my_dag():
...
# Airflow: task-level SLA (fail if task exceeds duration)
load = PythonOperator(
task_id="load",
python_callable=load_fn,
sla=timedelta(minutes=30), # alert if this task takes >30 min
)# Emit pipeline metrics to Prometheus/StatsD
from airflow.stats import Stats
Stats.incr("pipeline.rows_processed", count=row_count, tags={"dag": dag_id})
Stats.timing("pipeline.duration_ms", value=duration_ms, tags={"dag": dag_id})catchup=True on a new DAG — Airflow will try to backfill all missed runs since start_date; set catchup=False on new DAGs and trigger backfills manually with airflow dags backfillupdated_at index on source tables — incremental loads do WHERE updated_at > watermark; without an index this is a full-table scan on every run; ensure the source has an index on the watermark columnmode="poke" on long-waiting sensors — poke mode holds a worker slot while waiting; use mode="reschedule" so the slot is released between checksunique_key + merge strategy once the table exceeds 10M rowsdbt test or row-count reconciliation as a gate before final loadscatchup=False and max_active_runs=1 unless backfill is intended{{ ds }}) used in queries, not wall-clock timemode="reschedule" not mode="poke"ref() used for all cross-model dependencies — never hardcoded table namesunique_key and handle late-arriving datadbt source freshness)not_null, unique, accepted_values, relationships on key columnsSee also:
database-design(index design, query optimization, migration patterns) See also:observability(structured logging, metrics, SLO alerting for pipeline health)
© kid-sid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/data-pipelines of kid-sid/claude-spellbook.
Open the folder on GitHubat commit a7c2ac9
Data Pipelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Pipelines this skillkid-sid/claude-spellbook | 189 | — | ~4.3k | Automated safety check: Pass | MIT | |
| Migrating Dagster To Airflowastronomer/agents | 450 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Deploying Airflowastronomer/agents | 450 | 1 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Senior Data Engineerbenchflow-ai/skillsbench | 1.8k | — | ~5.9k | Automated safety check: Pass | MIT | |
| Senior Data Engineeralirezarezvani/claude-skills | 28k | 3 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Data Engineerdavila7/claude-code-templates | 32k | 7 repos | ~2.8k | Automated safety check: Pass | MIT |
astronomer/agents
Guide for migrating Dagster projects to Apache Airflow 3 on Astro.
astronomer/agents
Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
alirezarezvani/claude-skills
Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.
davila7/claude-code-templates
Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.
davila7/claude-code-templates
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.
kid-sid/claude-spellbook
A skill your agent uses when building or reviewing UI components for keyboard and screen reader compatibility, adding ARIA to custom widgets, auditing a page for WCAG AA conformance, or preparing…
kid-sid/claude-spellbook
A skill your agent uses when building, wiring, or debugging an Agentex agent — choosing agent type, configuring acp.py and manifest.yaml, using adk.messages or adk.state, or resolving…
kid-sid/claude-spellbook
A skill your agent uses when building production LLM applications — designing RAG pipelines, choosing vector databases, implementing agent orchestration, optimizing cost, or adding AI safety…
kid-sid/claude-spellbook
A skill your agent uses when building or refactoring Angular applications — choosing between signals, RxJS, and NgRx for state, configuring routing with guards and lazy loading, optimizing change…
kid-sid/claude-spellbook
A skill your agent uses when designing new REST endpoints, reviewing an existing API contract, adding pagination or filtering, planning a versioning strategy, or building a public or partner-facing…
kid-sid/claude-spellbook
A skill your agent uses when implementing login flows, issuing or validating JWTs, setting up OAuth2/OIDC with a provider, designing role-based or attribute-based access control, securing API…
Works with
Categories
A skill your agent uses when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating…. Data Pipelines is an agent skill from kid-sid/claude-spellbook. Use when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating data quality, or orchestrating multi-step data workflows.
Data Pipelines fits situations like: debugging data pipelines with Airflow; writing dbt models; designing incremental loads; implementing idempotent ETL/ELT jobs.
Run `npx skills add kid-sid/claude-spellbook --skill data-pipelines -a claude-code`. Or copy the skill folder (skills/data-pipelines in kid-sid/claude-spellbook) into .claude/skills/data-pipelines in your project. Claude Code loads it when a task matches its description.
Run `npx skills add kid-sid/claude-spellbook --skill data-pipelines -a codex`. Or copy the skill folder (skills/data-pipelines in kid-sid/claude-spellbook) into .agents/skills/data-pipelines in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kid-sid/claude-spellbook --skill data-pipelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-pipelines, .gemini/skills/data-pipelines, .github/skills/data-pipelines and .opencode/skills/data-pipelines in your project.
Going by SKILL.md and its folder, Data Pipelines needs the command-line tools its instructions call (dbt and airflow). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Pipelines is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Pipelines: Migrating Dagster To Airflow (astronomer/agents, 450 stars), Deploying Airflow (astronomer/agents, 450 stars), Senior Data Engineer (benchflow-ai/skillsbench, 1.8k stars) and Senior Data Engineer (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
kid-sid (a GitHub user) maintains it in kid-sid/claude-spellbook, which has 189 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on August 5, 2026.
Source: kid-sid/claude-spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.