Agent skill

Data Pipelines

by kid-sid in kid-sid/claude-spellbook

A skill your agent uses when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating…

MITAuto-check passedData & Analytics

Install Data Pipelines

skills CLI
$ npx skills add kid-sid/claude-spellbook --skill data-pipelines -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kid-sid/claude-spellbook data-pipelines --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kid-sid/claude-spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-pipelines .claude/skills/data-pipelines && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-pipelines
GitHub stars
189
Token cost
~4.3k tokens
SKILL.md length
656 words
Files
1
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating…

  • Debugging data pipelines with Airflow
  • SKILL.md covers When to Activate, ETL vs ELT Decision, Airflow and dbt, plus 6 more sections
  • Calls dbt and airflow
  • Writing dbt models

What it does

Data Pipelines is an agent skill from kid-sid/claude-spellbook. Use when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating data quality, or orchestrating multi-step data workflows.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL. It works with dbt and Apache Airflow. The repository describes itself as: A curated collection of skills, prompts, and workflows that extend Claude's capabilities — your personal grimoire for AI-powered development. The licence is MIT.

When your agent uses it

  • Debugging data pipelines with Airflow
  • Writing dbt models
  • Designing incremental loads
  • Implementing idempotent ETL/ELT jobs

Example prompts

  • “/data-pipelines”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit a7c2ac9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • dbt
    • airflow

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Pipelines loads about 4.3k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 656 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kid-sid/claude-spellbook at commit a7c2ac9, republished under its MIT licence (© kid-sid). 656 words, ~4,295 tokens.

Download SKILL.mdSave it as .claude/skills/data-pipelines/SKILL.md (or your agent's skills folder).
name
data-pipelines
description
Use when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating data quality, or orchestrating multi-step data workflows.

Data Pipelines

Orchestration, transformation, and validation patterns for production data pipelines.

When to Activate

  • Writing Airflow DAGs, operators, sensors, or XComs
  • Building dbt models, sources, tests, or macros
  • Designing incremental vs full-load strategies
  • Implementing idempotent pipeline runs
  • Validating data quality with dbt tests or Great Expectations
  • Orchestrating multi-step ELT/ETL workflows
  • Debugging failed runs, backfills, or data freshness issues

ETL vs ELT Decision

ApproachTransform whereUse when
ETLBefore loading (in pipeline code)Target warehouse has limited compute; PII must be masked before storage
ELTAfter loading (in warehouse SQL)Modern warehouse (BigQuery, Snowflake, Redshift); raw data must be preserved
StreamingContinuously (Kafka + Flink/Spark)Sub-minute latency required; event sourcing

Default for modern stacks: ELT — land raw data, transform with dbt, version-control SQL.

Airflow

DAG Structure
python
from datetime import datetime, timedelta
from airflow.decorators import dag, task
from airflow.operators.python import PythonOperator
from airflow.providers.postgres.hooks.postgres import PostgresHook

@dag(
    schedule="0 6 * * *",          # 6 AM daily
    start_date=datetime(2026, 1, 1),
    catchup=False,                  # don't backfill missed runs on deploy
    max_active_runs=1,              # prevent overlapping runs
    default_args={
        "retries": 3,
        "retry_delay": timedelta(minutes=5),
        "retry_exponential_backoff": True,
        "email_on_failure": True,
    },
    tags=["finance", "daily"],
)
def daily_revenue_pipeline():

    @task
    def extract_orders(execution_date=None) -> list[dict]:
        hook = PostgresHook(postgres_conn_id="source_db")
        # Use execution_date for idempotent extraction
        rows = hook.get_records(
            "SELECT * FROM orders WHERE date = %s",
            parameters=[execution_date.date()],
        )
        return [dict(r) for r in rows]

    @task
    def transform(orders: list[dict]) -> list[dict]:
        return [
            {**o, "revenue_usd": o["amount"] * o["fx_rate"]}
            for o in orders
            if o["status"] == "completed"
        ]

    @task
    def load(records: list[dict], execution_date=None):
        hook = PostgresHook(postgres_conn_id="warehouse")
        # Idempotent: delete-then-insert for the partition date
        hook.run("DELETE FROM daily_revenue WHERE date = %s", parameters=[execution_date.date()])
        hook.insert_rows("daily_revenue", [[r["date"], r["revenue_usd"]] for r in records])

    orders = extract_orders()
    transformed = transform(orders)
    load(transformed)

dag = daily_revenue_pipeline()
Operators & Sensors
python
from airflow.operators.bash import BashOperator
from airflow.operators.python import BranchPythonOperator
from airflow.sensors.filesystem import FileSensor
from airflow.sensors.sql import SqlSensor
from airflow.providers.http.sensors.http import HttpSensor

# Wait for a file to appear (S3, GCS, local)
wait_for_export = FileSensor(
    task_id="wait_for_export",
    filepath="/data/exports/{{ ds }}/orders.csv",
    poke_interval=60,    # check every 60s
    timeout=3600,        # fail after 1 hour
    mode="reschedule",   # release worker slot while waiting
)

# Wait for upstream table to be populated
wait_for_source = SqlSensor(
    task_id="wait_for_orders",
    conn_id="source_db",
    sql="SELECT COUNT(*) FROM orders WHERE date = '{{ ds }}' HAVING COUNT(*) > 0",
    poke_interval=120,
    mode="reschedule",
)

# Branch: skip load on weekends
def should_load(**context):
    if context["execution_date"].weekday() >= 5:
        return "skip_load"
    return "load"

branch = BranchPythonOperator(task_id="check_day", python_callable=should_load)
XComs — Task Communication
python
# Push value
@task
def extract() -> dict:
    return {"row_count": 1042, "checksum": "abc123"}  # return value auto-pushes XCom

# Pull value
@task
def validate(stats: dict):   # passed as argument from task dependency
    assert stats["row_count"] > 0, "Empty extract"

# Manual XCom pull (classic operators)
def load(**context):
    stats = context["task_instance"].xcom_pull(task_ids="extract")
    print(stats["row_count"])

XCom limits: XComs are stored in the Airflow metadata DB — not suited for large data. Pass row counts, checksums, and file paths through XComs; never entire datasets.

Dynamic Task Mapping
python
@task
def get_regions() -> list[str]:
    return ["us-east", "eu-west", "ap-south"]

@task
def process_region(region: str):
    extract_and_load(region)

# Creates one task instance per region — parallelized automatically
process_region.expand(region=get_regions())
Connections & Variables
python
from airflow.hooks.base import BaseHook
from airflow.models import Variable

# Never hardcode credentials — use Connections
conn = BaseHook.get_connection("my_postgres")
dsn = f"postgresql://{conn.login}:{conn.password}@{conn.host}/{conn.schema}"

# Runtime config — use Variables (or better: Airflow Params)
batch_size = int(Variable.get("etl_batch_size", default_var=1000))

dbt

Project Structure
dbt_project/
├── models/
│   ├── staging/          # stg_* — raw → typed, renamed, deduplicated
│   │   └── stg_orders.sql
│   ├── intermediate/     # int_* — business logic joins
│   │   └── int_order_items.sql
│   └── marts/            # final — wide tables for BI/downstream
│       └── fct_revenue.sql
├── tests/                # custom SQL tests
├── macros/               # Jinja macros
├── seeds/                # static CSV reference data
└── dbt_project.yml
Model Types & Materializations
sql
-- staging/stg_orders.sql
-- Materialization: view (cheap, always fresh)
{{ config(materialized='view') }}

SELECT
    order_id::VARCHAR      AS order_id,
    user_id::VARCHAR       AS user_id,
    created_at::TIMESTAMP  AS created_at,
    amount_cents / 100.0   AS amount_usd,
    status
FROM {{ source('raw', 'orders') }}
WHERE status != 'test'
sql
-- marts/fct_revenue.sql
-- Materialization: table (fast reads, rebuilt on each run)
{{ config(materialized='table') }}

SELECT
    DATE_TRUNC('day', o.created_at) AS date,
    p.name                          AS product_name,
    SUM(oi.quantity)                AS units_sold,
    SUM(oi.quantity * oi.unit_price) AS revenue_usd
FROM {{ ref('stg_orders') }}      o    -- ref() creates dependency
JOIN {{ ref('int_order_items') }} oi ON o.order_id = oi.order_id
JOIN {{ ref('stg_products') }}    p  ON oi.product_id = p.product_id
WHERE o.status = 'completed'
GROUP BY 1, 2
Incremental Models
sql
-- Only process new/updated rows — essential for large tables
{{ config(
    materialized='incremental',
    unique_key='order_id',
    incremental_strategy='merge',    -- or 'delete+insert', 'insert_overwrite'
    on_schema_change='append_new_columns',
) }}

SELECT
    order_id,
    user_id,
    amount_usd,
    created_at,
    updated_at
FROM {{ source('raw', 'orders') }}

{% if is_incremental() %}
    -- Only load rows newer than the last run
    WHERE updated_at > (SELECT MAX(updated_at) FROM {{ this }})
{% endif %}
Sources & Freshness
yaml
# models/staging/sources.yml
version: 2

sources:
  - name: raw
    database: analytics
    schema: raw_data
    freshness:
      warn_after: {count: 6, period: hour}
      error_after: {count: 24, period: hour}
    loaded_at_field: _loaded_at       # column that holds ingestion timestamp
    tables:
      - name: orders
        description: Raw orders from the transactional database
      - name: products
bash
# Check source freshness in CI
dbt source freshness
dbt Tests
yaml
# models/staging/stg_orders.yml
version: 2

models:
  - name: stg_orders
    columns:
      - name: order_id
        tests:
          - not_null
          - unique
      - name: status
        tests:
          - accepted_values:
              values: ["pending", "completed", "cancelled", "refunded"]
      - name: user_id
        tests:
          - not_null
          - relationships:
              to: ref('stg_users')
              field: user_id
      - name: amount_usd
        tests:
          - not_null
          - dbt_utils.accepted_range:
              min_value: 0
              max_value: 100000
sql
-- tests/assert_revenue_non_negative.sql — custom SQL test (fails if rows returned)
SELECT date, revenue_usd
FROM {{ ref('fct_revenue') }}
WHERE revenue_usd < 0
Macros
sql
-- macros/cents_to_dollars.sql
{% macro cents_to_dollars(column_name) %}
    ({{ column_name }} / 100.0)::NUMERIC(10, 2)
{% endmacro %}

-- Usage in a model
SELECT {{ cents_to_dollars('amount_cents') }} AS amount_usd
sql
-- macros/generate_surrogate_key.sql (or use dbt_utils)
{% macro surrogate_key(fields) %}
    MD5(CONCAT_WS('|', {% for f in fields %}COALESCE(CAST({{ f }} AS VARCHAR), ''){% if not loop.last %}, {% endif %}{% endfor %}))
{% endmacro %}
dbt Commands
bash
dbt run                              # run all models
dbt run --select staging             # run a directory
dbt run --select stg_orders+         # run model and all downstream
dbt run --select +fct_revenue        # run model and all upstream
dbt test                             # run all tests
dbt test --select stg_orders         # test one model
dbt build                            # run + test in dependency order
dbt source freshness                 # check source data freshness
dbt docs generate && dbt docs serve  # generate + serve lineage docs
dbt compile                          # render SQL without running

Idempotency Patterns

A pipeline run is idempotent if running it twice produces the same result as running it once.

python
# GOOD: delete-then-insert for a known partition
def load_partition(date: str, records: list[dict]):
    with engine.begin() as conn:
        conn.execute(
            text("DELETE FROM daily_stats WHERE date = :date"),
            {"date": date}
        )
        conn.execute(insert(DailyStats), records)

# GOOD: UPSERT (merge) on unique key
def upsert_orders(records: list[dict]):
    stmt = pg_insert(orders_table).values(records)
    stmt = stmt.on_conflict_do_update(
        index_elements=["order_id"],
        set_={"status": stmt.excluded.status, "updated_at": stmt.excluded.updated_at}
    )
    with engine.begin() as conn:
        conn.execute(stmt)

# BAD: append-only — reruns duplicate data
def load_orders(records):
    engine.execute(insert(orders_table).values(records))  # duplicates on rerun

Airflow idempotency: Use {{ ds }} (execution date, not run date) in all queries. Two runs for the same ds must produce the same output.


Incremental Load Strategies

StrategyHowUse When
Full refreshTruncate + reload entire tableSmall tables (<1M rows), no CDC
Incremental by timestampWHERE updated_at > last_run_maxSource has reliable updated_at
Incremental by partitionProcess one date partition per runAppend-only event data
CDC (change data capture)Debezium → Kafka → warehouseHigh-volume, low-latency, soft deletes
Snapshotdbt snapshot (strategy: timestamp)Track slowly-changing dimensions
python
# Watermark-based incremental (Python)
def get_watermark(conn, table: str) -> datetime:
    row = conn.execute(
        text("SELECT COALESCE(MAX(updated_at), '1970-01-01') FROM :table", bindparams=[bindparam("table")])
    ).fetchone()
    return row[0]

def extract_incremental(source_conn, watermark: datetime) -> list[dict]:
    return source_conn.execute(
        text("SELECT * FROM orders WHERE updated_at > :wm ORDER BY updated_at"),
        {"wm": watermark},
    ).fetchall()

Data Validation

dbt-native (preferred)
yaml
# Generic tests: not_null, unique, accepted_values, relationships
# Package tests: dbt_utils, dbt_expectations (Great Expectations style)
- name: amount_usd
  tests:
    - dbt_expectations.expect_column_values_to_be_between:
        min_value: 0
        max_value: 50000
        row_condition: "status = 'completed'"
Python validation (Great Expectations)
python
import great_expectations as gx

context = gx.get_context()
suite = context.add_expectation_suite("orders_suite")

validator = context.get_validator(
    batch_request=batch_request,
    expectation_suite_name="orders_suite",
)
validator.expect_column_values_to_not_be_null("order_id")
validator.expect_column_values_to_be_unique("order_id")
validator.expect_column_values_to_be_between("amount_usd", min_value=0)
validator.expect_column_pair_values_A_to_be_greater_than_B(
    "completed_at", "created_at"
)

results = validator.validate()
if not results.success:
    raise ValueError(f"Data quality check failed: {results}")
Row-count reconciliation
python
@task
def reconcile(source_count: int, target_count: int, tolerance: float = 0.001):
    delta = abs(source_count - target_count) / max(source_count, 1)
    if delta > tolerance:
        raise ValueError(
            f"Row count mismatch: source={source_count}, target={target_count}, "
            f"delta={delta:.2%} > {tolerance:.2%} tolerance"
        )

Monitoring & Alerting

python
# Airflow: SLA miss callback
def sla_miss_callback(dag, task_list, blocking_task_list, slas, blocking_tis):
    send_slack_alert(f"SLA missed for DAG {dag.dag_id}: {task_list}")

@dag(sla_miss_callback=sla_miss_callback)
def my_dag():
    ...

# Airflow: task-level SLA (fail if task exceeds duration)
load = PythonOperator(
    task_id="load",
    python_callable=load_fn,
    sla=timedelta(minutes=30),   # alert if this task takes >30 min
)
python
# Emit pipeline metrics to Prometheus/StatsD
from airflow.stats import Stats

Stats.incr("pipeline.rows_processed", count=row_count, tags={"dag": dag_id})
Stats.timing("pipeline.duration_ms", value=duration_ms, tags={"dag": dag_id})

Red Flags

  • catchup=True on a new DAG — Airflow will try to backfill all missed runs since start_date; set catchup=False on new DAGs and trigger backfills manually with airflow dags backfill
  • Passing datasets through XComs — XComs are stored in the Airflow metadata DB (SQLite or Postgres); passing DataFrames or large lists corrupts the DB and kills performance; pass file paths, row counts, or checksums only
  • Non-idempotent pipeline — if a run fails halfway and must be retried, appending duplicates corrupts the target; always upsert or delete-then-insert on a partition key
  • No updated_at index on source tables — incremental loads do WHERE updated_at > watermark; without an index this is a full-table scan on every run; ensure the source has an index on the watermark column
  • Hard-coded credentials in DAG code — DAGs are stored in version control and Airflow logs; always use Airflow Connections or environment variables, never string literals
  • mode="poke" on long-waiting sensors — poke mode holds a worker slot while waiting; use mode="reschedule" so the slot is released between checks
  • Unbounded full-refresh on large tables — a full refresh of a 500M-row table is slow and expensive; use incremental models with unique_key + merge strategy once the table exceeds 10M rows
  • No data quality tests before downstream loads — failing silently and loading bad data is worse than failing loudly; add dbt test or row-count reconciliation as a gate before final loads
Show full SKILL.md (140 more words)Show less

Checklist

  • DAG has catchup=False and max_active_runs=1 unless backfill is intended
  • All tasks are idempotent — reruns produce the same result
  • Execution date ({{ ds }}) used in queries, not wall-clock time
  • XComs carry only metadata (counts, paths, checksums) — not datasets
  • Airflow Connections used for all credentials — no hardcoded secrets
  • Sensors use mode="reschedule" not mode="poke"
  • dbt staging models rename, cast, and deduplicate raw source data
  • ref() used for all cross-model dependencies — never hardcoded table names
  • Incremental models have unique_key and handle late-arriving data
  • Source freshness checks configured and run in CI (dbt source freshness)
  • dbt tests cover: not_null, unique, accepted_values, relationships on key columns
  • Row-count reconciliation between source and target after each load
  • SLA alerts configured for critical DAGs
  • Backfill procedure documented and tested

See also: database-design (index design, query optimization, migration patterns) See also: observability (structured logging, metrics, SLO alerting for pipeline health)

© kid-sid, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/data-pipelines of kid-sid/claude-spellbook.

Open the folder on GitHubat commit a7c2ac9

Compare with similar skills

Data Pipelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Pipelines compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Pipelines this skillkid-sid/claude-spellbook189—~4.3kAutomated safety check: PassMIT
Migrating Dagster To Airflowastronomer/agents450—~3.8kAutomated safety check: PassApache-2.0
Deploying Airflowastronomer/agents4501 repos~2.8kAutomated safety check: PassApache-2.0
Senior Data Engineerbenchflow-ai/skillsbench1.8k—~5.9kAutomated safety check: PassMIT
Senior Data Engineeralirezarezvani/claude-skills28k3 repos~1.4kAutomated safety check: PassMIT
Data Engineerdavila7/claude-code-templates32k7 repos~2.8kAutomated safety check: PassMIT

Similar skills

  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    450 GitHub stars~3.8k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Deploying Airflow

    astronomer/agents

    Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~2.8k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    alirezarezvani/claude-skills

    Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    28k GitHub starsUsed in 3 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Data Engineer

    davila7/claude-code-templates

    Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

    32k GitHub starsUsed in 7 repos~2.8k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    davila7/claude-code-templates

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    32k GitHub starsUsed in 1 repo~1.4k tokens
    Data & AnalyticsAuto-check passed

More from kid-sid/claude-spellbook

All 54 skills in this repo
  • Accessibility

    kid-sid/claude-spellbook

    A skill your agent uses when building or reviewing UI components for keyboard and screen reader compatibility, adding ARIA to custom widgets, auditing a page for WCAG AA conformance, or preparing…

    189 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Agentex

    kid-sid/claude-spellbook

    A skill your agent uses when building, wiring, or debugging an Agentex agent — choosing agent type, configuring acp.py and manifest.yaml, using adk.messages or adk.state, or resolving…

    189 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • AI Engineer

    kid-sid/claude-spellbook

    A skill your agent uses when building production LLM applications — designing RAG pipelines, choosing vector databases, implementing agent orchestration, optimizing cost, or adding AI safety…

    189 GitHub stars~3.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Angular

    kid-sid/claude-spellbook

    A skill your agent uses when building or refactoring Angular applications — choosing between signals, RxJS, and NgRx for state, configuring routing with guards and lazy loading, optimizing change…

    189 GitHub stars~5k tokensUpdated 2 mo ago
    Auto-check passed
  • API Design

    kid-sid/claude-spellbook

    A skill your agent uses when designing new REST endpoints, reviewing an existing API contract, adding pagination or filtering, planning a versioning strategy, or building a public or partner-facing…

    189 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Auth

    kid-sid/claude-spellbook

    A skill your agent uses when implementing login flows, issuing or validating JWTs, setting up OAuth2/OIDC with a provider, designing role-based or attribute-based access control, securing API…

    189 GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Data Pipelines

What does Data Pipelines do?

A skill your agent uses when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating…. Data Pipelines is an agent skill from kid-sid/claude-spellbook. Use when building or debugging data pipelines with Airflow or Prefect, writing dbt models or tests, designing incremental loads, implementing idempotent ETL/ELT jobs, validating data quality, or orchestrating multi-step data workflows.

When should I use Data Pipelines?

Data Pipelines fits situations like: debugging data pipelines with Airflow; writing dbt models; designing incremental loads; implementing idempotent ETL/ELT jobs.

How do I install Data Pipelines in Claude Code?

Run `npx skills add kid-sid/claude-spellbook --skill data-pipelines -a claude-code`. Or copy the skill folder (skills/data-pipelines in kid-sid/claude-spellbook) into .claude/skills/data-pipelines in your project. Claude Code loads it when a task matches its description.

How do I install Data Pipelines in Codex?

Run `npx skills add kid-sid/claude-spellbook --skill data-pipelines -a codex`. Or copy the skill folder (skills/data-pipelines in kid-sid/claude-spellbook) into .agents/skills/data-pipelines in your project. Codex loads it when a task matches its description.

Can I use Data Pipelines in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kid-sid/claude-spellbook --skill data-pipelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-pipelines, .gemini/skills/data-pipelines, .github/skills/data-pipelines and .opencode/skills/data-pipelines in your project.

What does Data Pipelines need to run?

Going by SKILL.md and its folder, Data Pipelines needs the command-line tools its instructions call (dbt and airflow). Our summary lists: Python 3.

Does Data Pipelines access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Pipelines safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Pipelines use?

Data Pipelines is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Pipelines use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Pipelines?

Skills that share tags, products or a category with Data Pipelines: Migrating Dagster To Airflow (astronomer/agents, 450 stars), Deploying Airflow (astronomer/agents, 450 stars), Senior Data Engineer (benchflow-ai/skillsbench, 1.8k stars) and Senior Data Engineer (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Pipelines?

kid-sid (a GitHub user) maintains it in kid-sid/claude-spellbook, which has 189 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on August 5, 2026.

Source: kid-sid/claude-spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.