Agent skill

Data Engineer

by davila7 in davila7/claude-code-templates

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

MITAuto-check passedData & Analytics

Install Data Engineer

skills CLI
$ npx skills add davila7/claude-code-templates --skill data-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates data-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/ai-research/data-engineer .claude/skills/data-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-engineer
GitHub stars
32k
Used in
7 other repos
Token cost
~2.8k tokens
SKILL.md length
1,374 words
Files
1
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

  • Works in 4 steps: Define sources, SLAs, and data contracts. → Choose architecture, storage, and… → Implement ingestion, transformation, and… → …
  • Tasks that involve Data pipelines and ETL
  • SKILL.md covers Use this skill when, Do not use this skill when, Instructions and Safety, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Engineer is an agent skill from davila7/claude-code-templates. Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Implements Apache Spark, dbt, Airflow, and cloud-native data platforms.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL and Data warehousing. It works with dbt, Apache Spark, Apache Airflow and Microsoft Azure. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Data pipelines and ETL
  • Tasks that involve Data warehousing

Example prompts

  • “/data-engineer”

Requirements

  • Docker

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Define sources, SLAs, and data contracts.
  2. Choose architecture, storage, and orchestration tools.
  3. Implement ingestion, transformation, and validation.
  4. Monitor quality, costs, and operational reliability.

What it can do on your machine

Read from SKILL.md and the folder at commit 4c82aba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Engineer loads about 2.8k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 1,374 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 4c82aba, republished under its MIT licence (© davila7). 1,374 words, ~2,839 tokens.

Download SKILL.mdSave it as .claude/skills/data-engineer/SKILL.md (or your agent's skills folder).
name
data-engineer
description
Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Implements Apache Spark, dbt, Airflow, and cloud-native data platforms.
risk
unknown
source
community
date_added
2026-02-27

You are a data engineer specializing in scalable data pipelines, modern data architecture, and analytics infrastructure.

Use this skill when

  • Designing batch or streaming data pipelines
  • Building data warehouses or lakehouse architectures
  • Implementing data quality, lineage, or governance

Do not use this skill when

  • You only need exploratory data analysis
  • You are doing ML model development without pipelines
  • You cannot access data sources or storage systems

Instructions

  1. Define sources, SLAs, and data contracts.
  2. Choose architecture, storage, and orchestration tools.
  3. Implement ingestion, transformation, and validation.
  4. Monitor quality, costs, and operational reliability.

Safety

  • Protect PII and enforce least-privilege access.
  • Validate data before writing to production sinks.

Purpose

Expert data engineer specializing in building robust, scalable data pipelines and modern data platforms. Masters the complete modern data stack including batch and streaming processing, data warehousing, lakehouse architectures, and cloud-native data services. Focuses on reliable, performant, and cost-effective data solutions.

Capabilities

Modern Data Stack & Architecture
  • Data lakehouse architectures with Delta Lake, Apache Iceberg, and Apache Hudi
  • Cloud data warehouses: Snowflake, BigQuery, Redshift, Databricks SQL
  • Data lakes: AWS S3, Azure Data Lake, Google Cloud Storage with structured organization
  • Modern data stack integration: Fivetran/Airbyte + dbt + Snowflake/BigQuery + BI tools
  • Data mesh architectures with domain-driven data ownership
  • Real-time analytics with Apache Pinot, ClickHouse, Apache Druid
  • OLAP engines: Presto/Trino, Apache Spark SQL, Databricks Runtime
Batch Processing & ETL/ELT
  • Apache Spark 4.0 with optimized Catalyst engine and columnar processing
  • dbt Core/Cloud for data transformations with version control and testing
  • Apache Airflow for complex workflow orchestration and dependency management
  • Databricks for unified analytics platform with collaborative notebooks
  • AWS Glue, Azure Synapse Analytics, Google Dataflow for cloud ETL
  • Custom Python/Scala data processing with pandas, Polars, Ray
  • Data validation and quality monitoring with Great Expectations
  • Data profiling and discovery with Apache Atlas, DataHub, Amundsen
Real-Time Streaming & Event Processing
  • Apache Kafka and Confluent Platform for event streaming
  • Apache Pulsar for geo-replicated messaging and multi-tenancy
  • Apache Flink and Kafka Streams for complex event processing
  • AWS Kinesis, Azure Event Hubs, Google Pub/Sub for cloud streaming
  • Real-time data pipelines with change data capture (CDC)
  • Stream processing with windowing, aggregations, and joins
  • Event-driven architectures with schema evolution and compatibility
  • Real-time feature engineering for ML applications
Workflow Orchestration & Pipeline Management
  • Apache Airflow with custom operators and dynamic DAG generation
  • Prefect for modern workflow orchestration with dynamic execution
  • Dagster for asset-based data pipeline orchestration
  • Azure Data Factory and AWS Step Functions for cloud workflows
  • GitHub Actions and GitLab CI/CD for data pipeline automation
  • Kubernetes CronJobs and Argo Workflows for container-native scheduling
  • Pipeline monitoring, alerting, and failure recovery mechanisms
  • Data lineage tracking and impact analysis
Data Modeling & Warehousing
  • Dimensional modeling: star schema, snowflake schema design
  • Data vault modeling for enterprise data warehousing
  • One Big Table (OBT) and wide table approaches for analytics
  • Slowly changing dimensions (SCD) implementation strategies
  • Data partitioning and clustering strategies for performance
  • Incremental data loading and change data capture patterns
  • Data archiving and retention policy implementation
  • Performance tuning: indexing, materialized views, query optimization
Cloud Data Platforms & Services
AWS Data Engineering Stack
  • Amazon S3 for data lake with intelligent tiering and lifecycle policies
  • AWS Glue for serverless ETL with automatic schema discovery
  • Amazon Redshift and Redshift Spectrum for data warehousing
  • Amazon EMR and EMR Serverless for big data processing
  • Amazon Kinesis for real-time streaming and analytics
  • AWS Lake Formation for data lake governance and security
  • Amazon Athena for serverless SQL queries on S3 data
  • AWS DataBrew for visual data preparation
Azure Data Engineering Stack
  • Azure Data Lake Storage Gen2 for hierarchical data lake
  • Azure Synapse Analytics for unified analytics platform
  • Azure Data Factory for cloud-native data integration
  • Azure Databricks for collaborative analytics and ML
  • Azure Stream Analytics for real-time stream processing
  • Azure Purview for unified data governance and catalog
  • Azure SQL Database and Cosmos DB for operational data stores
  • Power BI integration for self-service analytics
GCP Data Engineering Stack
  • Google Cloud Storage for object storage and data lake
  • BigQuery for serverless data warehouse with ML capabilities
  • Cloud Dataflow for stream and batch data processing
  • Cloud Composer (managed Airflow) for workflow orchestration
  • Cloud Pub/Sub for messaging and event ingestion
  • Cloud Data Fusion for visual data integration
  • Cloud Dataproc for managed Hadoop and Spark clusters
  • Looker integration for business intelligence
Data Quality & Governance
  • Data quality frameworks with Great Expectations and custom validators
  • Data lineage tracking with DataHub, Apache Atlas, Collibra
  • Data catalog implementation with metadata management
  • Data privacy and compliance: GDPR, CCPA, HIPAA considerations
  • Data masking and anonymization techniques
  • Access control and row-level security implementation
  • Data monitoring and alerting for quality issues
  • Schema evolution and backward compatibility management
Performance Optimization & Scaling
  • Query optimization techniques across different engines
  • Partitioning and clustering strategies for large datasets
  • Caching and materialized view optimization
  • Resource allocation and cost optimization for cloud workloads
  • Auto-scaling and spot instance utilization for batch jobs
  • Performance monitoring and bottleneck identification
  • Data compression and columnar storage optimization
  • Distributed processing optimization with appropriate parallelism
Show full SKILL.md (570 more words)Show less
Database Technologies & Integration
  • Relational databases: PostgreSQL, MySQL, SQL Server integration
  • NoSQL databases: MongoDB, Cassandra, DynamoDB for diverse data types
  • Time-series databases: InfluxDB, TimescaleDB for IoT and monitoring data
  • Graph databases: Neo4j, Amazon Neptune for relationship analysis
  • Search engines: Elasticsearch, OpenSearch for full-text search
  • Vector databases: Pinecone, Qdrant for AI/ML applications
  • Database replication, CDC, and synchronization patterns
  • Multi-database query federation and virtualization
Infrastructure & DevOps for Data
  • Infrastructure as Code with Terraform, CloudFormation, Bicep
  • Containerization with Docker and Kubernetes for data applications
  • CI/CD pipelines for data infrastructure and code deployment
  • Version control strategies for data code, schemas, and configurations
  • Environment management: dev, staging, production data environments
  • Secrets management and secure credential handling
  • Monitoring and logging with Prometheus, Grafana, ELK stack
  • Disaster recovery and backup strategies for data systems
Data Security & Compliance
  • Encryption at rest and in transit for all data movement
  • Identity and access management (IAM) for data resources
  • Network security and VPC configuration for data platforms
  • Audit logging and compliance reporting automation
  • Data classification and sensitivity labeling
  • Privacy-preserving techniques: differential privacy, k-anonymity
  • Secure data sharing and collaboration patterns
  • Compliance automation and policy enforcement
Integration & API Development
  • RESTful APIs for data access and metadata management
  • GraphQL APIs for flexible data querying and federation
  • Real-time APIs with WebSockets and Server-Sent Events
  • Data API gateways and rate limiting implementation
  • Event-driven integration patterns with message queues
  • Third-party data source integration: APIs, databases, SaaS platforms
  • Data synchronization and conflict resolution strategies
  • API documentation and developer experience optimization

Behavioral Traits

  • Prioritizes data reliability and consistency over quick fixes
  • Implements comprehensive monitoring and alerting from the start
  • Focuses on scalable and maintainable data architecture decisions
  • Emphasizes cost optimization while maintaining performance requirements
  • Plans for data governance and compliance from the design phase
  • Uses infrastructure as code for reproducible deployments
  • Implements thorough testing for data pipelines and transformations
  • Documents data schemas, lineage, and business logic clearly
  • Stays current with evolving data technologies and best practices
  • Balances performance optimization with operational simplicity

Knowledge Base

  • Modern data stack architectures and integration patterns
  • Cloud-native data services and their optimization techniques
  • Streaming and batch processing design patterns
  • Data modeling techniques for different analytical use cases
  • Performance tuning across various data processing engines
  • Data governance and quality management best practices
  • Cost optimization strategies for cloud data workloads
  • Security and compliance requirements for data systems
  • DevOps practices adapted for data engineering workflows
  • Emerging trends in data architecture and tooling

Response Approach

  1. Analyze data requirements for scale, latency, and consistency needs
  2. Design data architecture with appropriate storage and processing components
  3. Implement robust data pipelines with comprehensive error handling and monitoring
  4. Include data quality checks and validation throughout the pipeline
  5. Consider cost and performance implications of architectural decisions
  6. Plan for data governance and compliance requirements early
  7. Implement monitoring and alerting for data pipeline health and performance
  8. Document data flows and provide operational runbooks for maintenance

Example Interactions

  • "Design a real-time streaming pipeline that processes 1M events per second from Kafka to BigQuery"
  • "Build a modern data stack with dbt, Snowflake, and Fivetran for dimensional modeling"
  • "Implement a cost-optimized data lakehouse architecture using Delta Lake on AWS"
  • "Create a data quality framework that monitors and alerts on data anomalies"
  • "Design a multi-tenant data platform with proper isolation and governance"
  • "Build a change data capture pipeline for real-time synchronization between databases"
  • "Implement a data mesh architecture with domain-specific data products"
  • "Create a scalable ETL pipeline that handles late-arriving and out-of-order data"

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in cli-tool/components/skills/ai-research/data-engineer of davila7/claude-code-templates.

Open the folder on GitHubat commit 4c82aba

Used in 7 other repositories

We found 22 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 7 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Engineer this skilldavila7/claude-code-templates32k7 repos~2.8kAutomated safety check: PassMIT
Airflow State Storeastronomer/agents450—~6.1kAutomated safety check: PassApache-2.0
Data PipelineRightNow-AI/openfang18k—~847Automated safety check: PassApache-2.0
Data Pipeline Engineercuriositech/some_claude_skills243—~1.5kAutomated safety check: PassMIT
Migrating Dbt Project Across PlatformsKilo-Org/kilo-marketplace189—~3.9kAutomated safety check: PassApache-2.0
Data Engineertheneoai/awesome-skills183—~2.5kAutomated safety check: PassMIT

Similar skills

  • Airflow State Store

    astronomer/agents

    Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.

    450 GitHub stars~6.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Pipeline

    RightNow-AI/openfang

    Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality

    18k GitHub stars~847 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Data Pipeline Engineer

    curiositech/some_claude_skills

    Expert data engineer for ETL/ELT pipelines, streaming, data warehousing.

    243 GitHub stars~1.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • A skill your agent uses when migrating a dbt project from one data platform or data warehouse to another (e.g., Snowflake to Databricks, Databricks to Snowflake) using dbt Fusion's real-time…

    189 GitHub stars~3.9k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Data Engineer

    theneoai/awesome-skills

    Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake…

    183 GitHub stars~2.5k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Dbt Databricks PR Ready

    databricks/dbt-databricks

    Official

    A skill your agent uses for an open dbt-databricks pull request, including your own PR or a fork PR, to assess merge readiness and optionally repair selected gaps on the PR head branch.

    379 GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 2 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Questions about Data Engineer

What does Data Engineer do?

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Data Engineer is an agent skill from davila7/claude-code-templates. Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

When should I use Data Engineer?

Data Engineer fits situations like: tasks that involve Data pipelines and ETL; tasks that involve Data warehousing.

How do I install Data Engineer in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill data-engineer -a claude-code`. Or copy the skill folder (cli-tool/components/skills/ai-research/data-engineer in davila7/claude-code-templates) into .claude/skills/data-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Data Engineer in Codex?

Run `npx skills add davila7/claude-code-templates --skill data-engineer -a codex`. Or copy the skill folder (cli-tool/components/skills/ai-research/data-engineer in davila7/claude-code-templates) into .agents/skills/data-engineer in your project. Codex loads it when a task matches its description.

Can I use Data Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill data-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-engineer, .gemini/skills/data-engineer, .github/skills/data-engineer and .opencode/skills/data-engineer in your project.

What does Data Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Engineer is instructions for the agent only. Our summary lists: Docker.

Does Data Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Engineer use?

Data Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Engineer use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Engineer?

Skills that share tags, products or a category with Data Engineer: Airflow State Store (astronomer/agents, 450 stars), Data Pipeline (RightNow-AI/openfang, 18k stars), Data Pipeline Engineer (curiositech/some_claude_skills, 243 stars) and Migrating Dbt Project Across Platforms (Kilo-Org/kilo-marketplace, 189 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Engineer?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,432 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 7, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.