Agent skill

Senior Data Engineer

by borghei in borghei/Claude-Skills

Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.

MITAuto-check passedData & Analytics

Install Senior Data Engineer

skills CLI
$ npx skills add borghei/Claude-Skills --skill senior-data-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills senior-data-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/senior-data-engineer .claude/skills/senior-data-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
senior-data-engineer
GitHub stars
874
Token cost
~1.4k tokens
SKILL.md length
411 words
Files
9 (incl. scripts, references)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.

  • Designing data architectures
  • SKILL.md covers Core Capabilities, When to Use, Clarify First and Quick Start, plus 3 more sections
  • Runs Python scripts from its folder; calls python
  • Building pipelines

What it does

Senior Data Engineer is an agent skill from borghei/Claude-Skills. Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka. Use when designing data architectures, building pipelines, adding data-quality checks, optimizing ETL/ELT, or troubleshooting pipeline failures.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/data_modeling_patterns.md`, `references/data_pipeline_architecture.md` and `references/dataops_best_practices.md`).

It sits in Data & Analytics, covering Data pipelines and ETL. It works with Apache Airflow, dbt, Apache Kafka and Dagster. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Designing data architectures
  • Building pipelines
  • Adding data-quality checks
  • Optimizing ETL/ELT

Example prompts

  • “/senior-data-engineer”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Senior Data Engineer loads about 1.4k tokens when it runs, and up to ~29k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 411 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~29k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 411 words, ~1,373 tokens.

Download SKILL.mdSave it as .claude/skills/senior-data-engineer/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
senior-data-engineer
description
Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka. Use when designing data architectures, building pipelines, adding data-quality checks, optimizing ETL/ELT, or troubleshooting pipeline failures.
license
MIT + Commons Clause
metadata.version
1.2.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
data-engineering
metadata.updated
2026-06-17
metadata.tags
airflow, spark, data-pipelines, warehousing, etl
metadata.python-tools
pipeline_orchestrator.py, data_quality_validator.py, etl_performance_optimizer.py
metadata.tech-stack
python, sql, spark, airflow, dbt, kafka

Senior Data Engineer

Generate pipeline configurations (Airflow, Prefect, Dagster), validate data quality with profiling and anomaly detection, and optimize SQL/Spark performance with actionable recommendations.

Core Capabilities

  • Pipeline generation — Airflow/Prefect/Dagster DAG code for batch and incremental loads, with DAG validation.
  • Data quality — schema validation, profiling, anomaly detection, data contracts, and Great Expectations suite generation.
  • ETL/ELT optimization — SQL and Spark analysis, partition strategy, and query cost estimation per warehouse.
  • Architecture decisions — batch vs streaming and warehouse vs lakehouse trade-off frameworks.
  • Reliability patterns — incremental watermarks, dead letter queues, freshness checks, and schema-drift detection.

When to Use

  • Designing a data architecture or choosing batch vs streaming / warehouse vs lakehouse.
  • Building or generating Airflow/Spark/dbt pipelines.
  • Adding data-quality checks or data contracts.
  • Optimizing slow ETL/ELT queries or troubleshooting pipeline failures.

Clarify First

Before generating pipelines, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Orchestrator — Airflow / Prefect / Dagster (--type; changes the generated DAG code)
  • Source, destination & load mode — systems involved and batch vs incremental (--source/--destination/--mode; shapes the pipeline)
  • Data-quality expectations — the schema and contracts to enforce (drives the Great Expectations suite generation)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

bash
# Generate an Airflow DAG for incremental PostgreSQL -> Snowflake
python scripts/pipeline_orchestrator.py generate \
  --type airflow --source postgres --destination snowflake \
  --tables orders,customers --mode incremental --schedule "0 5 * * *"

# Validate data quality against a schema
python scripts/data_quality_validator.py validate data.csv \
  --schema schema.json --detect-anomalies --json

# Profile a dataset
python scripts/data_quality_validator.py profile data.csv --json

# Optimize a slow SQL query
python scripts/etl_performance_optimizer.py analyze-sql query.sql \
  --warehouse snowflake --json

# Estimate query cost
python scripts/etl_performance_optimizer.py estimate-cost query.sql \
  --warehouse bigquery --stats data_stats.json --json

Tools

ToolSubcommandsPurpose
pipeline_orchestrator.pygenerate, validate, templateGenerate Airflow/Prefect/Dagster pipeline code, validate DAGs
data_quality_validator.pyvalidate, profile, generate-suite, contract, schemaSchema validation, profiling, anomaly detection, Great Expectations
etl_performance_optimizer.pyanalyze-sql, analyze-spark, optimize-partition, estimate-cost, templateSQL/Spark optimization, partition strategy, cost estimation

All subcommands support --json for machine-readable output and --output for file writing.

Show full SKILL.md (149 more words)Show less

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

Integration Points

SkillIntegration
senior-data-scientistFeature engineering consumes curated mart data
senior-ml-engineerML pipelines depend on feature store tables
senior-devopsCI/CD for dbt, Airflow deployment, container orchestration
senior-architectArchitecture reviews for lakehouse vs warehouse decisions
code-reviewerPipeline code reviews for DAGs, dbt models, Spark jobs

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in engineering/senior-data-engineer of borghei/Claude-Skills.

  • SKILL.md
  • references/data_modeling_patterns.md
  • references/data_pipeline_architecture.md
  • references/dataops_best_practices.md
  • references/decisions-and-troubleshooting.md
  • references/pipeline-workflows.md
  • scripts/data_quality_validator.py
  • scripts/etl_performance_optimizer.py
  • scripts/pipeline_orchestrator.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

Senior Data Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Senior Data Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Senior Data Engineer this skillborghei/Claude-Skills874—~1.4kAutomated safety check: PassMIT
Engineering Data Pipelinestelagod/code-abyss244—~236Automated safety check: PassMIT
Migrating Dagster To Airflowastronomer/agents450—~3.8kAutomated safety check: PassApache-2.0
Senior Data Engineerbenchflow-ai/skillsbench1.8k—~5.9kAutomated safety check: PassMIT
Senior Data Engineeralirezarezvani/claude-skills28k3 repos~1.4kAutomated safety check: PassMIT
Senior Data Engineerdavila7/claude-code-templates32k1 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Engineering Data Pipelines

    telagod/code-abyss

    Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.

    244 GitHub stars~236 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    450 GitHub stars~3.8k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    alirezarezvani/claude-skills

    Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    28k GitHub starsUsed in 3 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    davila7/claude-code-templates

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    32k GitHub starsUsed in 1 repo~1.4k tokens
    Data & AnalyticsAuto-check passed
  • AI Data Engineering

    ancoleman/ai-design-components

    Data pipelines, feature stores, and embedding generation for AI/ML systems.

    526 GitHub starsUsed in 1 repo~3.5k tokens
    Data & AnalyticsAuto-check passed

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Senior Data Engineer

What does Senior Data Engineer do?

Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka. Senior Data Engineer is an agent skill from borghei/Claude-Skills. Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.

When should I use Senior Data Engineer?

Senior Data Engineer fits situations like: designing data architectures; building pipelines; adding data-quality checks; optimizing ETL/ELT.

How do I install Senior Data Engineer in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill senior-data-engineer -a claude-code`. Or copy the skill folder (engineering/senior-data-engineer in borghei/Claude-Skills) into .claude/skills/senior-data-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Senior Data Engineer in Codex?

Run `npx skills add borghei/Claude-Skills --skill senior-data-engineer -a codex`. Or copy the skill folder (engineering/senior-data-engineer in borghei/Claude-Skills) into .agents/skills/senior-data-engineer in your project. Codex loads it when a task matches its description.

Can I use Senior Data Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill senior-data-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/senior-data-engineer, .gemini/skills/senior-data-engineer, .github/skills/senior-data-engineer and .opencode/skills/senior-data-engineer in your project.

What does Senior Data Engineer need to run?

Going by SKILL.md and its folder, Senior Data Engineer needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Senior Data Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Senior Data Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Senior Data Engineer use?

Senior Data Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Senior Data Engineer use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 28k tokens, read only when the agent opens those files.

What are the alternatives to Senior Data Engineer?

Skills that share tags, products or a category with Senior Data Engineer: Engineering Data Pipelines (telagod/code-abyss, 244 stars), Migrating Dagster To Airflow (astronomer/agents, 450 stars), Senior Data Engineer (benchflow-ai/skillsbench, 1.8k stars) and Senior Data Engineer (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Senior Data Engineer?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.