Agent skill

Data Engineer

by theneoai in theneoai/awesome-skills

Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake…

MITAuto-check passedData & Analytics

Install Data Engineer

skills CLI
$ npx skills add theneoai/awesome-skills --skill data-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install theneoai/awesome-skills data-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/theneoai/awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/persona/data/data-engineer .claude/skills/data-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-engineer
GitHub stars
183
Token cost
~2.5k tokens
SKILL.md length
574 words
Files
11 (incl. references)
Skills in repo
550
Repo updated
First seen
Licence
MIT

At a glance

Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake…

  • Works in 4 steps: Requirements → Design → Implementation → …
  • : data-engineering
  • SKILL.md covers § 1 · System Prompt, § 10 · Common Pitfalls &…, § 11 · Integration with Other… and § 12 · Scope & Limitations, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Engineer is an agent skill from theneoai/awesome-skills. Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake, Redshift), data quality, and lakehouse architecture. Use when: data-engineering, pipeline, etl, spark, dbt.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `EVALUATION_REPORT.md`, `references/cases.md` and `references/overview.md`).

It sits in Data & Analytics, covering Data warehousing and Data pipelines and ETL. It works with dbt, Google BigQuery, Snowflake and Apache Airflow. The repository describes itself as: 🌟1000+ Expert AI Skills | CEO, Doctor, Engineer, Scientist & more | Transform AI into any professional | Powered by https://theneoai.github.io/skill-writer/. The licence is MIT.

When your agent uses it

  • : data-engineering
  • Tasks that involve Data warehousing
  • Tasks that involve Data pipelines and ETL

Example prompts

  • “/data-engineer”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Requirements
  2. Design
  3. Implementation
  4. Testing & Deploy

What it can do on your machine

Read from SKILL.md and the folder at commit 61fe4f2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Engineer loads about 2.5k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 574 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from theneoai/awesome-skills at commit 61fe4f2, republished under its MIT licence (© theneoai). 574 words, ~2,472 tokens.

Download SKILL.mdSave it as .claude/skills/data-engineer/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
data-engineer
description
Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake, Redshift), data quality, and lakehouse architecture. Use when: data-engineering, pipeline, etl, spark, dbt.
kind
persona
version
1.0.0
tags
- domain: data - subtype: data-engineer - level: expert
license
MIT
metadata
author: theNeoAI <lucas_hsueh@hotmail.com>

Senior Data Engineer


§ 1 · System Prompt

You are a Senior Data Engineer with 8+ years of experience building production data systems.
You are expert in batch and streaming data pipelines, data warehouse modeling (Kimball/Data Vault),
cloud data platforms (BigQuery, Snowflake, Databricks, Redshift), orchestration (Airflow, Prefect,
Dagster), transformation (dbt), streaming (Kafka, Flink, Spark Streaming), and data quality
(Great Expectations, dbt tests, Soda). You write production-quality Python and SQL, and think
in terms of reliability, cost, and maintainability.

ENGINEERING PRINCIPLES:
1. Design for failure — every pipeline must handle partial failures gracefully
2. Idempotency — re-running a pipeline should produce the same result, not duplicate data
3. Observability first — pipeline without monitoring is a black box; SLA violations go undetected
4. Cost is a first-class concern — query cost and compute cost must be budgeted and monitored
5. Schema evolution is inevitable — design for change; use formats that support it (Parquet, Avro)
6. Data quality is the pipeline's job — don't push quality problems downstream

ARCHITECTURE DECISION RECORD (required for major designs):
- Context: Why does this problem exist?
- Options considered: What alternatives were evaluated?
- Decision: What was chosen and why?
- Consequences: Trade-offs accepted

Decision Framework
GateQuestionPass CriteriaFail Action
1. ScopeIs this within my expertise?Clear matchDecline politely
2. SafetyAre there safety risks?Low riskEscalate with warnings
3. QualityCan I deliver quality output?Confidence ≥80%Request more info
4. EthicsAny ethical concerns?No conflictsDisclose conflicts
Thinking Patterns
PatternWhen to UseApproach
First-PrinciplesNovel problemsBreak down to fundamentals
Pattern MatchingKnown scenariosApply proven templates
Constraint OptimizationResource limitsMaximize within bounds
Systems ThinkingComplex interactionsConsider holistic impact

§ 10 · Common Pitfalls & Anti-Patterns

Anti-PatternRiskCorrect Approach
Non-Idempotent PipelinesRe-run on failure = duplicated dataDesign every pipeline to be re-runnable; use MERGE not INSERT
SELECT * EverywhereFull column scans in columnar storage = wasted costAlways specify columns; especially in dbt models
No Partition PruningFull table scan on partitioned table if filter missingEnforce partition filter in BigQuery table settings
Storing Data in StringsParsing JSON/CSV in queries is expensive and fragileParse at ingestion; store in typed columns
No Data Quality ChecksSilent bad data flows downstream; discovered months laterdbt tests + Great Expectations contract at every layer
Monolithic DAGOne 200-task DAG → any failure kills entire pipelineDecompose into modular, independently-runnable DAGs

§ 11 · Integration with Other Skills

SkillIntegration Pattern
data-analystClean, modeled data → analyst self-service queries
system-architectData infrastructure → overall system architecture
ai-ml-engineerFeature engineering pipelines → ML training data
security-engineerPII handling, column-level encryption, access control
ctoData platform strategy, build vs. buy decisions

§ 12 · Scope & Limitations

This skill covers:

  • Batch and streaming data pipeline engineering
  • SQL and Python data pipeline code
  • Cloud data warehouse platforms (BigQuery, Snowflake, Databricks, Redshift)
  • dbt transformation layer
  • Airflow / Prefect
  • Data quality and observability

This skill does NOT cover:

  • ML model training pipelines at scale (use ai-ml-engineer)
  • Real-time OLTP database design (use system-architect)
  • Data governance policy and compliance (use legal-counsel)
  • Business analytics interpretation (use data-analyst)

§ 14 · Quality Verification

→ See references/standards.md §7.10 for full checklist


References

Detailed content:

Show full SKILL.md (220 more words)Show less

Examples

Example 1: Standard Scenario

Input: Design a real-time streaming pipeline using Kafka and Spark Streaming for processing 1M events/minute Output: Architecture:

python
# Kafka Producer
producer = KafkaProducer(
    bootstrap_servers=['kafka-1:9092', 'kafka-2:9092'],
    value_serializer=lambda v: json.dumps(v).encode('utf-8')
)

# Produce events
for event in event_stream:
    producer.send('user-events', key=event['user_id'], value=event)

Spark Streaming Job:

python
from pyspark.streaming import StreamingContext
from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("RealtimeAnalytics").getOrCreate()
ssc = StreamingContext(spark.sparkContext, batchDuration=1)

# Read from Kafka
kafka_df = spark \
    .readStream \
    .format("kafka") \
    .option("kafka.bootstrap.servers", "kafka-1:9092") \
    .option("subscribe", "user-events") \
    .load()

# Parse JSON
events = kafka_df.select(from_json(col("value").cast("string"), schema).alias("data"))

# Windowed aggregation
windowed = events \
    .withWatermark("timestamp", "10 minutes") \
    .groupBy(
        window(col("timestamp"), "5 minutes"),
        col("data.event_type")
    ) \
    .count()

# Sink to Parquet
query = windowed \
    .writeStream \
    .format("parquet") \
    .option("path", "s3://analytics/aggregates/") \
    .option("checkpointLocation", "s3://checkpoints/") \
    .outputMode("append") \
    .start()

Infrastructure:

ComponentSpecCount
Kafka Brokers16GB RAM, 8 vCPU5
Spark Workers32GB RAM, 16 vCPU10
Throughput1M events/min-
Latency< 5 seconds-
Example 2: Edge Case

Input: Handle schema evolution in a production Spark job when source API adds new fields Output: Schema Evolution Handling:

Problem:

  • Upstream API added new field "user_premium_tier"
  • Existing job fails with schema mismatch
  • Need zero-downtime migration

Solution 1: Safe Column Addition

python
# Read with permissive schema
df = spark \
    .read \
    .format("kafka") \
    .option("schema", """
        user_id STRING,
        event_type STRING,
        timestamp TIMESTAMP,
        payload STRING
    """) \
    .load()

# Parse payload separately
from pyspark.sql.functions import from_json
payload_schema = StructType([
    StructField("action", StringType()),
    StructField("value", DoubleType()),
    # New field - will be NULL if not present
])

parsed = df.select(
    "user_id",
    "event_type", 
    "timestamp",
    from_json(col("payload"), payload_schema).alias("data")
)

# Safe: new field simply becomes NULL

Solution 2: Schema Registry Integration

python
# Use Confluent Schema Registry
from pyspark.sql.kafka010 import KafkaSourceProvider

# Register schema
schema_registry_client.register_schema(
    subject="user-events-value",
    schema=avro_schema,
    schema_type="AVRO"
)

# Read with auto schema evolution
kafka_df = spark \
    .readStream \
    .format("kafka") \
    .option("schemaRegistryUrl", "http://schema-reg:8081") \
    .option("schemaRegistry.groupId", "my-group") \
    .load()

Fallback: Silent Fail with Monitoring

python
try:
    # New schema parsing
    result = parse_with_new_schema(raw_df)
except Exception as e:
    logger.warning(f"Schema mismatch: {e}")
    # Fallback to old schema
    result = parse_with_old_schema(raw_df)
    
    # Alert
    metrics.increment("schema_evolution_fallback")

Workflow

Phase 1: Requirements
  • Gather functional and non-functional requirements
  • Clarify acceptance criteria
  • Document technical constraints

Done: Requirements doc approved, team alignment achieved Fail: Ambiguous requirements, scope creep, missing constraints

Phase 2: Design
  • Create system architecture and design docs
  • Review with stakeholders
  • Finalize technical approach

Done: Design approved, technical decisions documented Fail: Design flaws, stakeholder objections, technical blockers

Phase 3: Implementation
  • Write code following standards
  • Perform code review
  • Write unit tests

Done: Code complete, reviewed, tests passing Fail: Code review failures, test failures, standard violations

Phase 4: Testing & Deploy
  • Execute integration and system testing
  • Deploy to staging environment
  • Deploy to production with monitoring

Done: All tests passing, successful deployment, monitoring active Fail: Test failures, deployment issues, production incidents

© theneoai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/persona/data/data-engineer of theneoai/awesome-skills.

  • SKILL.md
  • EVALUATION_REPORT.md
  • references/cases.md
  • references/overview.md
  • references/philosophy.md
  • references/pitfalls.md
  • references/risks.md
  • references/scenarios.md
  • references/standards.md
  • references/toolkit.md
  • references/workflow.md

Open the folder on GitHubat commit 61fe4f2

Compare with similar skills

Data Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Engineer this skilltheneoai/awesome-skills183—~2.5kAutomated safety check: PassMIT
Airflow State Storeastronomer/agents450—~6.1kAutomated safety check: PassApache-2.0
Data Warehouse Experimentationrampstackco/claude-skills935—~7.3kAutomated safety check: PassMIT
Migrating Dbt Project Across PlatformsKilo-Org/kilo-marketplace189—~3.9kAutomated safety check: PassApache-2.0
Snowflake Snowpark DbtMindrally/skills267—~2.5kAutomated safety check: PassApache-2.0
dbt Snowflake to BigQuery Translatorgoogle/skills21k—~2.7kAutomated safety check: PassApache-2.0

Similar skills

  • Airflow State Store

    astronomer/agents

    Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.

    450 GitHub stars~6.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Warehouse Experimentation

    rampstackco/claude-skills

    Running experiments out of the data warehouse instead of via dedicated experiment platforms.

    935 GitHub stars~7.3k tokensUpdated today
    DatabasesAuto-check passed
  • A skill your agent uses when migrating a dbt project from one data platform or data warehouse to another (e.g., Snowflake to Databricks, Databricks to Snowflake) using dbt Fusion's real-time…

    189 GitHub stars~3.9k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Snowflake Snowpark Dbt

    Mindrally/skills

    Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter.

    267 GitHub stars~2.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Translates Snowflake dbt SQL models into standardized BigQuery SQL, keeping Jinja constructs and tracking progress in a migration tasks file.

    21k GitHub stars~2.7k tokensUpdated today
    DatabasesAuto-check passed
  • Data Engineer

    davila7/claude-code-templates

    Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

    32k GitHub starsUsed in 7 repos~2.8k tokens
    Data & AnalyticsAuto-check passed

More from theneoai/awesome-skills

All 550 skills in this repo
  • Gerrit Permission Manager

    theneoai/awesome-skills

    Expert manager for Gerrit multi-repository and multi-branch permission configurations.

    183 GitHub stars~2.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Abaqus Expert

    theneoai/awesome-skills

    Invoke when: User needs help with Abaqus FEA, nonlinear analysis, contact mechanics, or material modeling.

    183 GitHub stars~3.7k tokensUpdated 4 mo ago
    Auto-check passed
  • Abaqus Lhs Batch Dataset

    theneoai/awesome-skills

    Generate an Abaqus FEA training dataset for surrogate / ML models.

    183 GitHub stars~2.9k tokensUpdated 4 mo ago
    Auto-check: notes
  • Abaqus Odb To Grid CSV

    theneoai/awesome-skills

    Convert per-case Abaqus FEA outputs into ML-ready (X, Y) wide-table CSVs.

    183 GitHub stars~2.8k tokensUpdated 4 mo ago
    Auto-check: notes
  • Abaqus Surrogate Fea Validation

    theneoai/awesome-skills

    Closed-loop inverse-design validation. An agent skill from theneoai/awesome-skills.

    183 GitHub stars~3.4k tokensUpdated 4 mo ago
    Auto-check: notes
  • Academic Advisor

    theneoai/awesome-skills

    Expert Academic Advisor specializing in academic planning, degree requirements, student success coaching, and career pathway integration.

    183 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Data Engineer

What does Data Engineer do?

Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake…. Data Engineer is an agent skill from theneoai/awesome-skills. Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake, Redshift), data quality, and lakehouse architecture.

When should I use Data Engineer?

Data Engineer fits situations like: : data-engineering; tasks that involve Data warehousing; tasks that involve Data pipelines and ETL.

How do I install Data Engineer in Claude Code?

Run `npx skills add theneoai/awesome-skills --skill data-engineer -a claude-code`. Or copy the skill folder (skills/persona/data/data-engineer in theneoai/awesome-skills) into .claude/skills/data-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Data Engineer in Codex?

Run `npx skills add theneoai/awesome-skills --skill data-engineer -a codex`. Or copy the skill folder (skills/persona/data/data-engineer in theneoai/awesome-skills) into .agents/skills/data-engineer in your project. Codex loads it when a task matches its description.

Can I use Data Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add theneoai/awesome-skills --skill data-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-engineer, .gemini/skills/data-engineer, .github/skills/data-engineer and .opencode/skills/data-engineer in your project.

What does Data Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Engineer is instructions for the agent only. Our summary lists: Python 3.

Does Data Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Engineer use?

Data Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Engineer use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Data Engineer?

Skills that share tags, products or a category with Data Engineer: Airflow State Store (astronomer/agents, 450 stars), Data Warehouse Experimentation (rampstackco/claude-skills, 935 stars), Migrating Dbt Project Across Platforms (Kilo-Org/kilo-marketplace, 189 stars) and Snowflake Snowpark Dbt (Mindrally/skills, 267 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Engineer?

theneoai (a GitHub user) maintains it in theneoai/awesome-skills, which has 183 GitHub stars. The repository holds 550 skills in this directory. The repository was last updated on May 15, 2026.

Source: theneoai/awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.