Senior Data Engineer
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
Expert data engineer for ETL/ELT pipelines, streaming, data warehousing.
$ npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install curiositech/some_claude_skills data-pipeline-engineer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/data-pipeline-engineer .claude/skills/data-pipeline-engineer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-pipeline-engineer" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/.claude/skills/data-pipeline-engineer into .claude/skills/data-pipeline-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipeline-engineer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/curiositech/some_claude_skills/tree/main/.claude/skills/data-pipeline-engineerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install curiositech/some_claude_skills data-pipeline-engineer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/data-pipeline-engineer .agents/skills/data-pipeline-engineer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-pipeline-engineer" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/.claude/skills/data-pipeline-engineer into .agents/skills/data-pipeline-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipeline-engineer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install curiositech/some_claude_skills data-pipeline-engineer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/data-pipeline-engineer .cursor/skills/data-pipeline-engineer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-pipeline-engineer" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/.claude/skills/data-pipeline-engineer into .cursor/skills/data-pipeline-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipeline-engineer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/curiositech/some_claude_skills.git --path .claude/skills/data-pipeline-engineer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install curiositech/some_claude_skills data-pipeline-engineer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/data-pipeline-engineer .gemini/skills/data-pipeline-engineer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-pipeline-engineer" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/.claude/skills/data-pipeline-engineer into .gemini/skills/data-pipeline-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipeline-engineer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install curiositech/some_claude_skills data-pipeline-engineerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/data-pipeline-engineer .github/skills/data-pipeline-engineer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-pipeline-engineer" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/.claude/skills/data-pipeline-engineer into .github/skills/data-pipeline-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipeline-engineer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install curiositech/some_claude_skills data-pipeline-engineer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/data-pipeline-engineer .opencode/skills/data-pipeline-engineer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-pipeline-engineer" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/.claude/skills/data-pipeline-engineer into .opencode/skills/data-pipeline-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-pipeline-engineer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-pipeline-engineerExpert data engineer for ETL/ELT pipelines, streaming, data warehousing.
Data Pipeline Engineer is an agent skill from curiositech/some_claude_skills. Expert data engineer for ETL/ELT pipelines, streaming, data warehousing. Activate on: data pipeline, ETL, ELT, data warehouse, Spark, Kafka, Airflow, dbt, data modeling, star schema, streaming data, batch processing, data quality. NOT for: API design (use api-architect), ML training (use ML skills), dashboards (use design skills).
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `.claude-plugin/plugin.json`, `CHANGELOG.md` and `references/airflow-dag.py`).
It sits in Data & Analytics, covering Data pipelines and ETL and Data warehousing. It works with dbt, Apache Airflow and Apache Kafka. The repository describes itself as: Claude skills that make my life easier. The licence is MIT.
10 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBash(dbt:*spark-submit:*airflow:*python:*)From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python and Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.getdbt.comairflow.apache.orgdocs.greatexpectations.iodocs.delta.iokafka.apache.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Pipeline Engineer loads about 1.5k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 516 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its MIT licence (© curiositech). 516 words, ~1,473 tokens.
.claude/skills/data-pipeline-engineer/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Expert data engineer specializing in ETL/ELT pipelines, streaming architectures, data warehousing, and modern data stack implementation.
| Capability | Technologies | Key Patterns |
|---|---|---|
| Batch Processing | Spark, dbt, Databricks | Incremental, partitioning, Delta/Iceberg |
| Stream Processing | Kafka, Flink, Spark Streaming | Watermarks, exactly-once, windowing |
| Orchestration | Airflow, Dagster, Prefect | DAG design, sensors, task groups |
| Data Modeling | dbt, SQL | Kimball, Data Vault, SCD |
| Data Quality | Great Expectations, dbt tests | Validation suites, freshness |
BRONZE (Raw) → Exact source copy, schema-on-read, partitioned by ingestion
↓ Cleaning, Deduplication
SILVER (Cleansed) → Validated, standardized, business logic applied
↓ Aggregation, Enrichment
GOLD (Business) → Dimensional models, aggregates, ready for BI/MLFull implementation examples in ./references/:
| File | Description |
|---|---|
dbt-project-structure.md | Complete dbt layout with staging, intermediate, marts |
airflow-dag.py | Production DAG with sensors, task groups, quality checks |
spark-streaming.py | Kafka-to-Delta processor with windowing |
great-expectations-suite.json | Comprehensive data quality expectation suite |
Symptom: Truncate and rebuild entire tables every run
Fix: Use incremental models with is_incremental(), partition by date
Symptom: Pipeline breaks when upstream adds/removes columns Fix: Explicit source contracts, select only needed columns in staging
Symptom: One 200-task DAG running 8 hours Fix: Domain-specific DAGs, ExternalTaskSensor for dependencies
Symptom: Bad data reaches production before detection Fix: Great Expectations or dbt tests at each layer, block on failures
Symptom: Raw data transformed without preserving original Fix: Always land raw in Bronze first, make transformations reproducible
Symptom: Manual updates needed for date filters
Fix: Use Airflow templating (e.g., ds variable) or dynamic date functions
Symptom: Unbounded state growth, OOM in long-running jobs
Fix: Add withWatermark() to handle late-arriving data
Symptom: Transient failures cause DAG failures
Fix: retries=3, retry_exponential_backoff=True, max_retry_delay
Symptom: No one knows where data comes from or who uses it Fix: dbt docs, data catalog integration, column-level lineage
Symptom: Bugs discovered by stakeholders, not engineers
Fix: dbt --target dev, sample datasets, CI/CD for models
Pipeline Design:
Data Quality:
Orchestration:
Operations:
Run ./scripts/validate-pipeline.sh to check:
© curiositech, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in .claude/skills/data-pipeline-engineer of curiositech/some_claude_skills.
Open the folder on GitHubat commit 6713fc7
Data Pipeline Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Pipeline Engineer this skillcuriositech/some_claude_skills | 243 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Senior Data Engineerbenchflow-ai/skillsbench | 1.8k | — | ~5.9k | Automated safety check: Pass | MIT | |
| Data Engineerdavila7/claude-code-templates | 32k | 8 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Senior Data Engineeralirezarezvani/claude-skills | 28k | 3 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Senior Data Engineerdavila7/claude-code-templates | 32k | 1 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Airflow State Storeastronomer/agents | 451 | — | ~6.1k | Automated safety check: Pass | Apache-2.0 |
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
davila7/claude-code-templates
Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.
alirezarezvani/claude-skills
Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.
davila7/claude-code-templates
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.
astronomer/agents
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.
telagod/code-abyss
Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.
curiositech/some_claude_skills
Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.
curiositech/some_claude_skills
End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.
curiositech/some_claude_skills
Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.
curiositech/some_claude_skills
Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.
curiositech/some_claude_skills
Build production computer vision pipelines for object detection, tracking, and video analysis.
curiositech/some_claude_skills
Long-running design anthropologist that builds comprehensive visual databases from 500-1000 real-world examples, extracting color palettes, typography patterns, layout systems, and interaction…
Works with
Categories
Expert data engineer for ETL/ELT pipelines, streaming, data warehousing. Data Pipeline Engineer is an agent skill from curiositech/some_claude_skills. Expert data engineer for ETL/ELT pipelines, streaming, data warehousing.
Data Pipeline Engineer fits situations like: tasks that involve Data pipelines and ETL; tasks that involve Data warehousing.
Run `npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a claude-code`. Or copy the skill folder (.claude/skills/data-pipeline-engineer in curiositech/some_claude_skills) into .claude/skills/data-pipeline-engineer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a codex`. Or copy the skill folder (.claude/skills/data-pipeline-engineer in curiositech/some_claude_skills) into .agents/skills/data-pipeline-engineer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill data-pipeline-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-pipeline-engineer, .gemini/skills/data-pipeline-engineer, .github/skills/data-pipeline-engineer and .opencode/skills/data-pipeline-engineer in your project.
Going by SKILL.md and its folder, Data Pipeline Engineer needs Python and a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(dbt:*,spark-submit:*,airflow:*,python:*).
SKILL.md names 5 domains. As links in the text: docs.getdbt.com, airflow.apache.org, docs.greatexpectations.io, docs.delta.io and kafka.apache.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Data Pipeline Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Pipeline Engineer: Senior Data Engineer (benchflow-ai/skillsbench, 1.8k stars), Data Engineer (davila7/claude-code-templates, 32k stars), Senior Data Engineer (alirezarezvani/claude-skills, 28k stars) and Senior Data Engineer (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 243 GitHub stars. The repository holds 109 skills in this directory. The repository was last updated on September 6, 2026.
Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.