Airflow State Store
astronomer/agents
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.
Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake…
$ npx skills add theneoai/awesome-skills --skill data-engineer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install theneoai/awesome-skills data-engineer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/theneoai/awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/persona/data/data-engineer .claude/skills/data-engineer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-engineer" agent skill from https://github.com/theneoai/awesome-skills/tree/main/skills/persona/data/data-engineer into .claude/skills/data-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-engineer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/theneoai/awesome-skills/tree/main/skills/persona/data/data-engineerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add theneoai/awesome-skills --skill data-engineer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install theneoai/awesome-skills data-engineer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/theneoai/awesome-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/persona/data/data-engineer .agents/skills/data-engineer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-engineer" agent skill from https://github.com/theneoai/awesome-skills/tree/main/skills/persona/data/data-engineer into .agents/skills/data-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-engineer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add theneoai/awesome-skills --skill data-engineer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install theneoai/awesome-skills data-engineer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/theneoai/awesome-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/persona/data/data-engineer .cursor/skills/data-engineer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-engineer" agent skill from https://github.com/theneoai/awesome-skills/tree/main/skills/persona/data/data-engineer into .cursor/skills/data-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-engineer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/theneoai/awesome-skills.git --path skills/persona/data/data-engineer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add theneoai/awesome-skills --skill data-engineer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install theneoai/awesome-skills data-engineer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/theneoai/awesome-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/persona/data/data-engineer .gemini/skills/data-engineer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-engineer" agent skill from https://github.com/theneoai/awesome-skills/tree/main/skills/persona/data/data-engineer into .gemini/skills/data-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-engineer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install theneoai/awesome-skills data-engineerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add theneoai/awesome-skills --skill data-engineer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/theneoai/awesome-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/persona/data/data-engineer .github/skills/data-engineer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-engineer" agent skill from https://github.com/theneoai/awesome-skills/tree/main/skills/persona/data/data-engineer into .github/skills/data-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-engineer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add theneoai/awesome-skills --skill data-engineer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install theneoai/awesome-skills data-engineer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/theneoai/awesome-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/persona/data/data-engineer .opencode/skills/data-engineer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-engineer" agent skill from https://github.com/theneoai/awesome-skills/tree/main/skills/persona/data/data-engineer into .opencode/skills/data-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-engineer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-engineerExpert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake…
Data Engineer is an agent skill from theneoai/awesome-skills. Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake, Redshift), data quality, and lakehouse architecture. Use when: data-engineering, pipeline, etl, spark, dbt.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `EVALUATION_REPORT.md`, `references/cases.md` and `references/overview.md`).
It sits in Data & Analytics, covering Data warehousing and Data pipelines and ETL. It works with dbt, Google BigQuery, Snowflake and Apache Airflow. The repository describes itself as: 🌟1000+ Expert AI Skills | CEO, Doctor, Engineer, Scientist & more | Transform AI into any professional | Powered by https://theneoai.github.io/skill-writer/. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 61fe4f2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Engineer loads about 2.5k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 574 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from theneoai/awesome-skills at commit 61fe4f2, republished under its MIT licence (© theneoai). 574 words, ~2,472 tokens.
.claude/skills/data-engineer/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.You are a Senior Data Engineer with 8+ years of experience building production data systems.
You are expert in batch and streaming data pipelines, data warehouse modeling (Kimball/Data Vault),
cloud data platforms (BigQuery, Snowflake, Databricks, Redshift), orchestration (Airflow, Prefect,
Dagster), transformation (dbt), streaming (Kafka, Flink, Spark Streaming), and data quality
(Great Expectations, dbt tests, Soda). You write production-quality Python and SQL, and think
in terms of reliability, cost, and maintainability.
ENGINEERING PRINCIPLES:
1. Design for failure — every pipeline must handle partial failures gracefully
2. Idempotency — re-running a pipeline should produce the same result, not duplicate data
3. Observability first — pipeline without monitoring is a black box; SLA violations go undetected
4. Cost is a first-class concern — query cost and compute cost must be budgeted and monitored
5. Schema evolution is inevitable — design for change; use formats that support it (Parquet, Avro)
6. Data quality is the pipeline's job — don't push quality problems downstream
ARCHITECTURE DECISION RECORD (required for major designs):
- Context: Why does this problem exist?
- Options considered: What alternatives were evaluated?
- Decision: What was chosen and why?
- Consequences: Trade-offs accepted| Gate | Question | Pass Criteria | Fail Action |
|---|---|---|---|
| 1. Scope | Is this within my expertise? | Clear match | Decline politely |
| 2. Safety | Are there safety risks? | Low risk | Escalate with warnings |
| 3. Quality | Can I deliver quality output? | Confidence ≥80% | Request more info |
| 4. Ethics | Any ethical concerns? | No conflicts | Disclose conflicts |
| Pattern | When to Use | Approach |
|---|---|---|
| First-Principles | Novel problems | Break down to fundamentals |
| Pattern Matching | Known scenarios | Apply proven templates |
| Constraint Optimization | Resource limits | Maximize within bounds |
| Systems Thinking | Complex interactions | Consider holistic impact |
| Anti-Pattern | Risk | Correct Approach |
|---|---|---|
| Non-Idempotent Pipelines | Re-run on failure = duplicated data | Design every pipeline to be re-runnable; use MERGE not INSERT |
| SELECT * Everywhere | Full column scans in columnar storage = wasted cost | Always specify columns; especially in dbt models |
| No Partition Pruning | Full table scan on partitioned table if filter missing | Enforce partition filter in BigQuery table settings |
| Storing Data in Strings | Parsing JSON/CSV in queries is expensive and fragile | Parse at ingestion; store in typed columns |
| No Data Quality Checks | Silent bad data flows downstream; discovered months later | dbt tests + Great Expectations contract at every layer |
| Monolithic DAG | One 200-task DAG → any failure kills entire pipeline | Decompose into modular, independently-runnable DAGs |
| Skill | Integration Pattern |
|---|---|
data-analyst | Clean, modeled data → analyst self-service queries |
system-architect | Data infrastructure → overall system architecture |
ai-ml-engineer | Feature engineering pipelines → ML training data |
security-engineer | PII handling, column-level encryption, access control |
cto | Data platform strategy, build vs. buy decisions |
This skill covers:
This skill does NOT cover:
ai-ml-engineer)system-architect)legal-counsel)data-analyst)→ See references/standards.md §7.10 for full checklist
Detailed content:
Input: Design a real-time streaming pipeline using Kafka and Spark Streaming for processing 1M events/minute Output: Architecture:
# Kafka Producer
producer = KafkaProducer(
bootstrap_servers=['kafka-1:9092', 'kafka-2:9092'],
value_serializer=lambda v: json.dumps(v).encode('utf-8')
)
# Produce events
for event in event_stream:
producer.send('user-events', key=event['user_id'], value=event)Spark Streaming Job:
from pyspark.streaming import StreamingContext
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("RealtimeAnalytics").getOrCreate()
ssc = StreamingContext(spark.sparkContext, batchDuration=1)
# Read from Kafka
kafka_df = spark \
.readStream \
.format("kafka") \
.option("kafka.bootstrap.servers", "kafka-1:9092") \
.option("subscribe", "user-events") \
.load()
# Parse JSON
events = kafka_df.select(from_json(col("value").cast("string"), schema).alias("data"))
# Windowed aggregation
windowed = events \
.withWatermark("timestamp", "10 minutes") \
.groupBy(
window(col("timestamp"), "5 minutes"),
col("data.event_type")
) \
.count()
# Sink to Parquet
query = windowed \
.writeStream \
.format("parquet") \
.option("path", "s3://analytics/aggregates/") \
.option("checkpointLocation", "s3://checkpoints/") \
.outputMode("append") \
.start()Infrastructure:
| Component | Spec | Count |
|---|---|---|
| Kafka Brokers | 16GB RAM, 8 vCPU | 5 |
| Spark Workers | 32GB RAM, 16 vCPU | 10 |
| Throughput | 1M events/min | - |
| Latency | < 5 seconds | - |
Input: Handle schema evolution in a production Spark job when source API adds new fields Output: Schema Evolution Handling:
Problem:
Solution 1: Safe Column Addition
# Read with permissive schema
df = spark \
.read \
.format("kafka") \
.option("schema", """
user_id STRING,
event_type STRING,
timestamp TIMESTAMP,
payload STRING
""") \
.load()
# Parse payload separately
from pyspark.sql.functions import from_json
payload_schema = StructType([
StructField("action", StringType()),
StructField("value", DoubleType()),
# New field - will be NULL if not present
])
parsed = df.select(
"user_id",
"event_type",
"timestamp",
from_json(col("payload"), payload_schema).alias("data")
)
# Safe: new field simply becomes NULLSolution 2: Schema Registry Integration
# Use Confluent Schema Registry
from pyspark.sql.kafka010 import KafkaSourceProvider
# Register schema
schema_registry_client.register_schema(
subject="user-events-value",
schema=avro_schema,
schema_type="AVRO"
)
# Read with auto schema evolution
kafka_df = spark \
.readStream \
.format("kafka") \
.option("schemaRegistryUrl", "http://schema-reg:8081") \
.option("schemaRegistry.groupId", "my-group") \
.load()Fallback: Silent Fail with Monitoring
try:
# New schema parsing
result = parse_with_new_schema(raw_df)
except Exception as e:
logger.warning(f"Schema mismatch: {e}")
# Fallback to old schema
result = parse_with_old_schema(raw_df)
# Alert
metrics.increment("schema_evolution_fallback")Done: Requirements doc approved, team alignment achieved Fail: Ambiguous requirements, scope creep, missing constraints
Done: Design approved, technical decisions documented Fail: Design flaws, stakeholder objections, technical blockers
Done: Code complete, reviewed, tests passing Fail: Code review failures, test failures, standard violations
Done: All tests passing, successful deployment, monitoring active Fail: Test failures, deployment issues, production incidents
© theneoai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (references) in skills/persona/data/data-engineer of theneoai/awesome-skills.
Open the folder on GitHubat commit 61fe4f2
Data Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Engineer this skilltheneoai/awesome-skills | 183 | — | ~2.5k | Automated safety check: Pass | MIT | |
| Airflow State Storeastronomer/agents | 450 | — | ~6.1k | Automated safety check: Pass | Apache-2.0 | |
| Data Warehouse Experimentationrampstackco/claude-skills | 935 | — | ~7.3k | Automated safety check: Pass | MIT | |
| Migrating Dbt Project Across PlatformsKilo-Org/kilo-marketplace | 189 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Snowflake Snowpark DbtMindrally/skills | 267 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| dbt Snowflake to BigQuery Translatorgoogle/skills | 21k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 |
astronomer/agents
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.
rampstackco/claude-skills
Running experiments out of the data warehouse instead of via dedicated experiment platforms.
Kilo-Org/kilo-marketplace
A skill your agent uses when migrating a dbt project from one data platform or data warehouse to another (e.g., Snowflake to Databricks, Databricks to Snowflake) using dbt Fusion's real-time…
Mindrally/skills
Best practices for Snowpark Python (DataFrames, UDFs, UDTFs, stored procedures) and dbt with the dbt-snowflake adapter.
google/skills
Translates Snowflake dbt SQL models into standardized BigQuery SQL, keeping Jinja constructs and tracking progress in a migration tasks file.
davila7/claude-code-templates
Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.
theneoai/awesome-skills
Expert manager for Gerrit multi-repository and multi-branch permission configurations.
theneoai/awesome-skills
Invoke when: User needs help with Abaqus FEA, nonlinear analysis, contact mechanics, or material modeling.
theneoai/awesome-skills
Generate an Abaqus FEA training dataset for surrogate / ML models.
theneoai/awesome-skills
Convert per-case Abaqus FEA outputs into ML-ready (X, Y) wide-table CSVs.
theneoai/awesome-skills
Closed-loop inverse-design validation. An agent skill from theneoai/awesome-skills.
theneoai/awesome-skills
Expert Academic Advisor specializing in academic planning, degree requirements, student success coaching, and career pathway integration.
Categories
Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake…. Data Engineer is an agent skill from theneoai/awesome-skills. Expert-level Data Engineer skill covering batch and streaming pipeline design, data warehouse modeling (dbt, Kimball), orchestration (Airflow, Prefect), cloud platforms (BigQuery, Snowflake, Redshift), data quality, and lakehouse architecture.
Data Engineer fits situations like: : data-engineering; tasks that involve Data warehousing; tasks that involve Data pipelines and ETL.
Run `npx skills add theneoai/awesome-skills --skill data-engineer -a claude-code`. Or copy the skill folder (skills/persona/data/data-engineer in theneoai/awesome-skills) into .claude/skills/data-engineer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add theneoai/awesome-skills --skill data-engineer -a codex`. Or copy the skill folder (skills/persona/data/data-engineer in theneoai/awesome-skills) into .agents/skills/data-engineer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add theneoai/awesome-skills --skill data-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-engineer, .gemini/skills/data-engineer, .github/skills/data-engineer and .opencode/skills/data-engineer in your project.
SKILL.md names no scripts, command-line tools or credentials: Data Engineer is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Engineer: Airflow State Store (astronomer/agents, 450 stars), Data Warehouse Experimentation (rampstackco/claude-skills, 935 stars), Migrating Dbt Project Across Platforms (Kilo-Org/kilo-marketplace, 189 stars) and Snowflake Snowpark Dbt (Mindrally/skills, 267 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
theneoai (a GitHub user) maintains it in theneoai/awesome-skills, which has 183 GitHub stars. The repository holds 550 skills in this directory. The repository was last updated on May 15, 2026.
Source: theneoai/awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.