Agent skill

Data Pipeline Engineer

by curiositech in curiositech/some_claude_skills

Expert data engineer for ETL/ELT pipelines, streaming, data warehousing.

MITAuto-check passedData & Analytics

Install Data Pipeline Engineer

skills CLI
$ npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install curiositech/some_claude_skills data-pipeline-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/data-pipeline-engineer .claude/skills/data-pipeline-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-pipeline-engineer
GitHub stars
243
Token cost
~1.5k tokens
SKILL.md length
516 words
Files
8 (incl. scripts, references)
Skills in repo
109
Repo updated
First seen
Licence
MIT

At a glance

Expert data engineer for ETL/ELT pipelines, streaming, data warehousing.

  • Works in 10 steps: Full Table Refreshes → Tight Coupling to Source Schemas → Monolithic DAGs → …
  • Tasks that involve Data pipelines and ETL
  • SKILL.md covers Quick Start, Core Capabilities, Architecture Patterns and Reference Examples, plus 4 more sections
  • Runs Python and Shell scripts from its folder

What it does

Data Pipeline Engineer is an agent skill from curiositech/some_claude_skills. Expert data engineer for ETL/ELT pipelines, streaming, data warehousing. Activate on: data pipeline, ETL, ELT, data warehouse, Spark, Kafka, Airflow, dbt, data modeling, star schema, streaming data, batch processing, data quality. NOT for: API design (use api-architect), ML training (use ML skills), dashboards (use design skills).

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `.claude-plugin/plugin.json`, `CHANGELOG.md` and `references/airflow-dag.py`).

It sits in Data & Analytics, covering Data pipelines and ETL and Data warehousing. It works with dbt, Apache Airflow and Apache Kafka. The repository describes itself as: Claude skills that make my life easier. The licence is MIT.

When your agent uses it

  • Tasks that involve Data pipelines and ETL
  • Tasks that involve Data warehousing

Example prompts

  • “/data-pipeline-engineer”

Requirements

  • Python 3
  • A Bash shell
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(dbt:*,spark-submit:*,airflow:*,python:*)

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Full Table Refreshes
  2. Tight Coupling to Source Schemas
  3. Monolithic DAGs
  4. No Data Quality Gates
  5. Processing Before Archiving
  6. Hardcoded Dates in Queries
  7. Missing Watermarks in Streaming
  8. No Retry/Backoff Strategy
  9. Undocumented Data Lineage
  10. Testing Only in Production

What it can do on your machine

Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(dbt:*
    • spark-submit:*
    • airflow:*
    • python:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.getdbt.com
    • airflow.apache.org
    • docs.greatexpectations.io
    • docs.delta.io
    • kafka.apache.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Pipeline Engineer loads about 1.5k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 516 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its MIT licence (© curiositech). 516 words, ~1,473 tokens.

Download SKILL.mdSave it as .claude/skills/data-pipeline-engineer/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
data-pipeline-engineer
description
Expert data engineer for ETL/ELT pipelines, streaming, data warehousing. Activate on: data pipeline, ETL, ELT, data warehouse, Spark, Kafka, Airflow, dbt, data modeling, star schema, streaming data, batch processing, data quality. NOT for: API design (use api-architect), ML training (use ML skills), dashboards (use design skills).
allowed-tools
Read, Write, Edit, Bash(dbt:*,spark-submit:*,airflow:*,python:*)
metadata.category
Data & Analytics
metadata.tags
etl, spark, kafka, airflow, data-warehouse

Data Pipeline Engineer

Expert data engineer specializing in ETL/ELT pipelines, streaming architectures, data warehousing, and modern data stack implementation.

Quick Start

  1. Identify sources - data formats, volumes, freshness requirements
  2. Choose architecture - Medallion (Bronze/Silver/Gold), Lambda, or Kappa
  3. Design layers - staging → intermediate → marts (dbt pattern)
  4. Add quality gates - Great Expectations or dbt tests at each layer
  5. Orchestrate - Airflow DAGs with sensors and retries
  6. Monitor - lineage, freshness, anomaly detection

Core Capabilities

CapabilityTechnologiesKey Patterns
Batch ProcessingSpark, dbt, DatabricksIncremental, partitioning, Delta/Iceberg
Stream ProcessingKafka, Flink, Spark StreamingWatermarks, exactly-once, windowing
OrchestrationAirflow, Dagster, PrefectDAG design, sensors, task groups
Data Modelingdbt, SQLKimball, Data Vault, SCD
Data QualityGreat Expectations, dbt testsValidation suites, freshness

Architecture Patterns

BRONZE (Raw)     → Exact source copy, schema-on-read, partitioned by ingestion
      ↓ Cleaning, Deduplication
SILVER (Cleansed) → Validated, standardized, business logic applied
      ↓ Aggregation, Enrichment
GOLD (Business)   → Dimensional models, aggregates, ready for BI/ML
Lambda vs Kappa
  • Lambda: Batch + Stream layers → merged serving layer (complex but complete)
  • Kappa: Stream-only with replay → simpler but requires robust streaming

Reference Examples

Full implementation examples in ./references/:

FileDescription
dbt-project-structure.mdComplete dbt layout with staging, intermediate, marts
airflow-dag.pyProduction DAG with sensors, task groups, quality checks
spark-streaming.pyKafka-to-Delta processor with windowing
great-expectations-suite.jsonComprehensive data quality expectation suite

Anti-Patterns (10 Critical Mistakes)

1. Full Table Refreshes

Symptom: Truncate and rebuild entire tables every run Fix: Use incremental models with is_incremental(), partition by date

2. Tight Coupling to Source Schemas

Symptom: Pipeline breaks when upstream adds/removes columns Fix: Explicit source contracts, select only needed columns in staging

3. Monolithic DAGs

Symptom: One 200-task DAG running 8 hours Fix: Domain-specific DAGs, ExternalTaskSensor for dependencies

4. No Data Quality Gates

Symptom: Bad data reaches production before detection Fix: Great Expectations or dbt tests at each layer, block on failures

5. Processing Before Archiving

Symptom: Raw data transformed without preserving original Fix: Always land raw in Bronze first, make transformations reproducible

6. Hardcoded Dates in Queries

Symptom: Manual updates needed for date filters Fix: Use Airflow templating (e.g., ds variable) or dynamic date functions

Show full SKILL.md (200 more words)Show less
7. Missing Watermarks in Streaming

Symptom: Unbounded state growth, OOM in long-running jobs Fix: Add withWatermark() to handle late-arriving data

8. No Retry/Backoff Strategy

Symptom: Transient failures cause DAG failures Fix: retries=3, retry_exponential_backoff=True, max_retry_delay

9. Undocumented Data Lineage

Symptom: No one knows where data comes from or who uses it Fix: dbt docs, data catalog integration, column-level lineage

10. Testing Only in Production

Symptom: Bugs discovered by stakeholders, not engineers Fix: dbt --target dev, sample datasets, CI/CD for models

Quality Checklist

Pipeline Design:

  • Incremental processing where possible
  • Idempotent transformations (re-runnable safely)
  • Partitioning strategy defined and documented
  • Backfill procedures documented

Data Quality:

  • Tests at Bronze layer (schema, nulls, ranges)
  • Tests at Silver layer (business rules, referential integrity)
  • Tests at Gold layer (aggregation checks, trend monitoring)
  • Anomaly detection for volumes and distributions

Orchestration:

  • Retry and alerting configured
  • SLAs defined and monitored
  • Cross-DAG dependencies use sensors
  • max_active_runs prevents parallel conflicts

Operations:

  • Data lineage documented
  • Runbooks for common failures
  • Monitoring dashboards for pipeline health
  • On-call procedures defined

Validation Script

Run ./scripts/validate-pipeline.sh to check:

  • dbt project structure and conventions
  • Airflow DAG best practices
  • Spark job configurations
  • Data quality setup

External Resources

© curiositech, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in .claude/skills/data-pipeline-engineer of curiositech/some_claude_skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • CHANGELOG.md
  • references/airflow-dag.py
  • references/dbt-project-structure.md
  • references/great-expectations-suite.json
  • references/spark-streaming.py
  • scripts/validate-pipeline.sh

Open the folder on GitHubat commit 6713fc7

Compare with similar skills

Data Pipeline Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Pipeline Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Pipeline Engineer this skillcuriositech/some_claude_skills243—~1.5kAutomated safety check: PassMIT
Senior Data Engineerbenchflow-ai/skillsbench1.8k—~5.9kAutomated safety check: PassMIT
Data Engineerdavila7/claude-code-templates32k8 repos~2.8kAutomated safety check: PassMIT
Senior Data Engineeralirezarezvani/claude-skills28k3 repos~1.4kAutomated safety check: PassMIT
Senior Data Engineerdavila7/claude-code-templates32k1 repos~1.4kAutomated safety check: PassMIT
Airflow State Storeastronomer/agents451—~6.1kAutomated safety check: PassApache-2.0

Similar skills

  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Data Engineer

    davila7/claude-code-templates

    Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

    32k GitHub starsUsed in 8 repos~2.8k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    alirezarezvani/claude-skills

    Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    28k GitHub starsUsed in 3 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    davila7/claude-code-templates

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    32k GitHub starsUsed in 1 repo~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Airflow State Store

    astronomer/agents

    Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.

    451 GitHub stars~6.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Engineering Data Pipelines

    telagod/code-abyss

    Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.

    243 GitHub stars~236 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed

More from curiositech/some_claude_skills

All 109 skills in this repo
  • Crisis Detection Intervention AI

    curiositech/some_claude_skills

    Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.

    243 GitHub starsUsed in 3 repos~3.8k tokens
    Auto-check passed
  • Form Validation Architect

    curiositech/some_claude_skills

    End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.

    243 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Competitive Cartographer

    curiositech/some_claude_skills

    Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.

    243 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • GitHub Actions Pipeline Builder

    curiositech/some_claude_skills

    Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.

    243 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: notes
  • Computer Vision Pipeline

    curiositech/some_claude_skills

    Build production computer vision pipelines for object detection, tracking, and video analysis.

    243 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed
  • Design Archivist

    curiositech/some_claude_skills

    Long-running design anthropologist that builds comprehensive visual databases from 500-1000 real-world examples, extracting color palettes, typography patterns, layout systems, and interaction…

    243 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Data Pipeline Engineer

What does Data Pipeline Engineer do?

Expert data engineer for ETL/ELT pipelines, streaming, data warehousing. Data Pipeline Engineer is an agent skill from curiositech/some_claude_skills. Expert data engineer for ETL/ELT pipelines, streaming, data warehousing.

When should I use Data Pipeline Engineer?

Data Pipeline Engineer fits situations like: tasks that involve Data pipelines and ETL; tasks that involve Data warehousing.

How do I install Data Pipeline Engineer in Claude Code?

Run `npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a claude-code`. Or copy the skill folder (.claude/skills/data-pipeline-engineer in curiositech/some_claude_skills) into .claude/skills/data-pipeline-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Data Pipeline Engineer in Codex?

Run `npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a codex`. Or copy the skill folder (.claude/skills/data-pipeline-engineer in curiositech/some_claude_skills) into .agents/skills/data-pipeline-engineer in your project. Codex loads it when a task matches its description.

Can I use Data Pipeline Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-pipeline-engineer, .gemini/skills/data-pipeline-engineer, .github/skills/data-pipeline-engineer and .opencode/skills/data-pipeline-engineer in your project.

What does Data Pipeline Engineer need to run?

Going by SKILL.md and its folder, Data Pipeline Engineer needs Python and a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(dbt:*,spark-submit:*,airflow:*,python:*).

Does Data Pipeline Engineer access the network?

SKILL.md names 5 domains. As links in the text: docs.getdbt.com, airflow.apache.org, docs.greatexpectations.io, docs.delta.io and kafka.apache.org. This is read from the text; nothing was executed.

Is Data Pipeline Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Pipeline Engineer use?

Data Pipeline Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Pipeline Engineer use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Data Pipeline Engineer?

Skills that share tags, products or a category with Data Pipeline Engineer: Senior Data Engineer (benchflow-ai/skillsbench, 1.8k stars), Data Engineer (davila7/claude-code-templates, 32k stars), Senior Data Engineer (alirezarezvani/claude-skills, 28k stars) and Senior Data Engineer (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Pipeline Engineer?

curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 243 GitHub stars. The repository holds 109 skills in this directory. The repository was last updated on September 6, 2026.

Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.