Agent skill

Senior Data Engineer

by davila7 in davila7/claude-code-templates

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

MITAuto-check passedData & Analytics

Install Senior Data Engineer

skills CLI
$ npx skills add davila7/claude-code-templates --skill senior-data-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates senior-data-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/development/senior-data-engineer .claude/skills/senior-data-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
senior-data-engineer
GitHub stars
32k
Used in
1 other repo
Token cost
~1.4k tokens
SKILL.md length
438 words
Files
7 (incl. scripts, references)
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

  • Works in 3 steps: Data Pipeline Architecture → Data Modeling Patterns → Dataops Best Practices
  • Designing data architectures
  • SKILL.md covers Quick Start, Core Expertise, Tech Stack and Reference Documentation, plus 7 more sections
  • Runs Python scripts from its folder; calls python, kubectl and docker

What it does

Senior Data Engineer is an agent skill from davila7/claude-code-templates. World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/data_modeling_patterns.md`, `references/data_pipeline_architecture.md` and `references/dataops_best_practices.md`).

It sits in Data & Analytics, covering Data pipelines and ETL. It works with dbt, Python, SQL and Apache Airflow. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Designing data architectures
  • Building data pipelines
  • Optimizing data workflows
  • Implementing data governance

Example prompts

  • “/senior-data-engineer”

Requirements

  • Python 3
  • Docker

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Data Pipeline Architecture
  2. Data Modeling Patterns
  3. Dataops Best Practices

What it can do on your machine

Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • kubectl
    • docker
    • helm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl, docker and helm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Senior Data Engineer loads about 1.4k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 438 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 438 words, ~1,377 tokens.

Download SKILL.mdSave it as .claude/skills/senior-data-engineer/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
senior-data-engineer
description
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

Senior Data Engineer

World-class senior data engineer skill for production-grade AI/ML/Data systems.

Quick Start

Main Capabilities
bash
# Core Tool 1
python scripts/pipeline_orchestrator.py --input data/ --output results/

# Core Tool 2  
python scripts/data_quality_validator.py --target project/ --analyze

# Core Tool 3
python scripts/etl_performance_optimizer.py --config config.yaml --deploy

Core Expertise

This skill covers world-class capabilities in:

  • Advanced production patterns and architectures
  • Scalable system design and implementation
  • Performance optimization at scale
  • MLOps and DataOps best practices
  • Real-time processing and inference
  • Distributed computing frameworks
  • Model deployment and monitoring
  • Security and compliance
  • Cost optimization
  • Team leadership and mentoring

Tech Stack

Languages: Python, SQL, R, Scala, Go ML Frameworks: PyTorch, TensorFlow, Scikit-learn, XGBoost Data Tools: Spark, Airflow, dbt, Kafka, Databricks LLM Frameworks: LangChain, LlamaIndex, DSPy Deployment: Docker, Kubernetes, AWS/GCP/Azure Monitoring: MLflow, Weights & Biases, Prometheus Databases: PostgreSQL, BigQuery, Snowflake, Pinecone

Reference Documentation

1. Data Pipeline Architecture

Comprehensive guide available in references/data_pipeline_architecture.md covering:

  • Advanced patterns and best practices
  • Production implementation strategies
  • Performance optimization techniques
  • Scalability considerations
  • Security and compliance
  • Real-world case studies
2. Data Modeling Patterns

Complete workflow documentation in references/data_modeling_patterns.md including:

  • Step-by-step processes
  • Architecture design patterns
  • Tool integration guides
  • Performance tuning strategies
  • Troubleshooting procedures
3. Dataops Best Practices

Technical reference guide in references/dataops_best_practices.md with:

  • System design principles
  • Implementation examples
  • Configuration best practices
  • Deployment strategies
  • Monitoring and observability

Production Patterns

Pattern 1: Scalable Data Processing

Enterprise-scale data processing with distributed computing:

  • Horizontal scaling architecture
  • Fault-tolerant design
  • Real-time and batch processing
  • Data quality validation
  • Performance monitoring
Pattern 2: ML Model Deployment

Production ML system with high availability:

  • Model serving with low latency
  • A/B testing infrastructure
  • Feature store integration
  • Model monitoring and drift detection
  • Automated retraining pipelines
Pattern 3: Real-Time Inference

High-throughput inference system:

  • Batching and caching strategies
  • Load balancing
  • Auto-scaling
  • Latency optimization
  • Cost optimization

Best Practices

Show full SKILL.md (181 more words)Show less
Development
  • Test-driven development
  • Code reviews and pair programming
  • Documentation as code
  • Version control everything
  • Continuous integration
Production
  • Monitor everything critical
  • Automate deployments
  • Feature flags for releases
  • Canary deployments
  • Comprehensive logging
Team Leadership
  • Mentor junior engineers
  • Drive technical decisions
  • Establish coding standards
  • Foster learning culture
  • Cross-functional collaboration

Performance Targets

Latency:

  • P50: < 50ms
  • P95: < 100ms
  • P99: < 200ms

Throughput:

  • Requests/second: > 1000
  • Concurrent users: > 10,000

Availability:

  • Uptime: 99.9%
  • Error rate: < 0.1%

Security & Compliance

  • Authentication & authorization
  • Data encryption (at rest & in transit)
  • PII handling and anonymization
  • GDPR/CCPA compliance
  • Regular security audits
  • Vulnerability management

Common Commands

bash
# Development
python -m pytest tests/ -v --cov
python -m black src/
python -m pylint src/

# Training
python scripts/train.py --config prod.yaml
python scripts/evaluate.py --model best.pth

# Deployment
docker build -t service:v1 .
kubectl apply -f k8s/
helm upgrade service ./charts/

# Monitoring
kubectl logs -f deployment/service
python scripts/health_check.py

Resources

  • Advanced Patterns: references/data_pipeline_architecture.md
  • Implementation Guide: references/data_modeling_patterns.md
  • Technical Reference: references/dataops_best_practices.md
  • Automation Scripts: scripts/ directory

Senior-Level Responsibilities

As a world-class senior professional:

  1. Technical Leadership

    • Drive architectural decisions
    • Mentor team members
    • Establish best practices
    • Ensure code quality
  2. Strategic Thinking

    • Align with business goals
    • Evaluate trade-offs
    • Plan for scale
    • Manage technical debt
  3. Collaboration

    • Work across teams
    • Communicate effectively
    • Build consensus
    • Share knowledge
  4. Innovation

    • Stay current with research
    • Experiment with new approaches
    • Contribute to community
    • Drive continuous improvement
  5. Production Excellence

    • Ensure high availability
    • Monitor proactively
    • Optimize performance
    • Respond to incidents

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in cli-tool/components/skills/development/senior-data-engineer of davila7/claude-code-templates.

  • SKILL.md
  • references/data_modeling_patterns.md
  • references/data_pipeline_architecture.md
  • references/dataops_best_practices.md
  • scripts/data_quality_validator.py
  • scripts/etl_performance_optimizer.py
  • scripts/pipeline_orchestrator.py

Open the folder on GitHubat commit 14680ec

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Senior Data Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Senior Data Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Senior Data Engineer this skilldavila7/claude-code-templates32k1 repos~1.4kAutomated safety check: PassMIT
Senior Data Engineerbenchflow-ai/skillsbench1.8k—~5.9kAutomated safety check: PassMIT
Senior Data Engineeralirezarezvani/claude-skills28k3 repos~1.4kAutomated safety check: PassMIT
Transforming Dataancoleman/ai-design-components526—~3kAutomated safety check: PassMIT
Databricks JobsKilo-Org/kilo-marketplace1901 repos~3.1kAutomated safety check: PassCustom licence
Engineering Data Pipelinestelagod/code-abyss243—~236Automated safety check: PassMIT

Similar skills

  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    alirezarezvani/claude-skills

    Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

    28k GitHub starsUsed in 3 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    526 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Databricks Jobs

    Kilo-Org/kilo-marketplace

    Develop and deploy Lakeflow Jobs on Databricks via DABs, Python SDK, or the CLI.

    190 GitHub starsUsed in 1 repo~3.1k tokens
    Data & AnalyticsAuto-check passed
  • Engineering Data Pipelines

    telagod/code-abyss

    Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.

    243 GitHub stars~236 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    borghei/Claude-Skills

    Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Questions about Senior Data Engineer

What does Senior Data Engineer do?

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Senior Data Engineer is an agent skill from davila7/claude-code-templates. World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure.

When should I use Senior Data Engineer?

Senior Data Engineer fits situations like: designing data architectures; building data pipelines; optimizing data workflows; implementing data governance.

How do I install Senior Data Engineer in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill senior-data-engineer -a claude-code`. Or copy the skill folder (cli-tool/components/skills/development/senior-data-engineer in davila7/claude-code-templates) into .claude/skills/senior-data-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Senior Data Engineer in Codex?

Run `npx skills add davila7/claude-code-templates --skill senior-data-engineer -a codex`. Or copy the skill folder (cli-tool/components/skills/development/senior-data-engineer in davila7/claude-code-templates) into .agents/skills/senior-data-engineer in your project. Codex loads it when a task matches its description.

Can I use Senior Data Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill senior-data-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/senior-data-engineer, .gemini/skills/senior-data-engineer, .github/skills/senior-data-engineer and .opencode/skills/senior-data-engineer in your project.

What does Senior Data Engineer need to run?

Going by SKILL.md and its folder, Senior Data Engineer needs Python for the scripts in its folder and the command-line tools its instructions call (python, kubectl, docker and helm). Our summary lists: Python 3; Docker.

Does Senior Data Engineer access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Senior Data Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Senior Data Engineer use?

Senior Data Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Senior Data Engineer use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to Senior Data Engineer?

Skills that share tags, products or a category with Senior Data Engineer: Senior Data Engineer (benchflow-ai/skillsbench, 1.8k stars), Senior Data Engineer (alirezarezvani/claude-skills, 28k stars), Transforming Data (ancoleman/ai-design-components, 526 stars) and Databricks Jobs (Kilo-Org/kilo-marketplace, 190 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Senior Data Engineer?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.