Migrating Dagster To Airflow
astronomer/agents
Guide for migrating Dagster projects to Apache Airflow 3 on Astro.
Data pipelines, feature stores, and embedding generation for AI/ML systems.
$ npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ancoleman/ai-design-components ai-data-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-data-engineering .claude/skills/ai-data-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-data-engineering" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/ai-data-engineering into .claude/skills/ai-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-data-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ancoleman/ai-design-components/tree/main/skills/ai-data-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ancoleman/ai-design-components ai-data-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ai-data-engineering .agents/skills/ai-data-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-data-engineering" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/ai-data-engineering into .agents/skills/ai-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-data-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ancoleman/ai-design-components ai-data-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ai-data-engineering .cursor/skills/ai-data-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-data-engineering" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/ai-data-engineering into .cursor/skills/ai-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-data-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ancoleman/ai-design-components.git --path skills/ai-data-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ancoleman/ai-design-components ai-data-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ai-data-engineering .gemini/skills/ai-data-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-data-engineering" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/ai-data-engineering into .gemini/skills/ai-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-data-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ancoleman/ai-design-components ai-data-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ai-data-engineering .github/skills/ai-data-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-data-engineering" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/ai-data-engineering into .github/skills/ai-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-data-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ancoleman/ai-design-components ai-data-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ai-data-engineering .opencode/skills/ai-data-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-data-engineering" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/ai-data-engineering into .opencode/skills/ai-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-data-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-data-engineeringData pipelines, feature stores, and embedding generation for AI/ML systems.
AI Data Engineering is an agent skill from ancoleman/ai-design-components. Data pipelines, feature stores, and embedding generation for AI/ML systems. Use when building RAG pipelines, ML feature serving, or data transformations. Covers feature stores (Feast, Tecton), embedding pipelines, chunking strategies, orchestration (Dagster, Prefect, Airflow), dbt transformations, data versioning (LakeFS), and experiment tracking (MLflow, W&B).
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 33 other files, including scripts and reference files (for example `examples/dagster-pipelines/embedding_pipeline.py`, `examples/feast-features/README.md` and `examples/feast-features/setup_features.py`).
It sits in Data & Analytics, covering Data pipelines and ETL, Embeddings and MLOps. It works with MLflow, dbt, Apache Airflow and Dagster. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.
Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pippythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AI Data Engineering loads about 3.5k tokens when it runs, and up to ~32k if it reads all its reference files. Until then it costs about 96 tokens; SKILL.md has 798 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 798 words, ~3,501 tokens.
.claude/skills/ai-data-engineering/SKILL.md (or your agent's skills folder). This skill also uses 28 other files; get the full folder from GitHub.Build data infrastructure for AI/ML systems including RAG pipelines, feature stores, and embedding generation. Provides architecture patterns, orchestration workflows, and evaluation metrics for production AI applications.
Use this skill when:
Skip this skill if:
RAG pipelines have 5 distinct stages. Understanding this architecture is critical for production implementations.
┌─────────────────────────────────────────────────────────────┐
│ RAG Pipeline (5 Stages) │
├─────────────────────────────────────────────────────────────┤
│ │
│ 1. INGESTION → Load documents (PDF, DOCX, Markdown) │
│ 2. INDEXING → Chunk (512 tokens) + Embed + Store │
│ 3. RETRIEVAL → Query embedding + Vector search + Filters │
│ 4. GENERATION → Context injection + LLM streaming │
│ 5. EVALUATION → RAGAS metrics (faithfulness, relevancy) │
│ │
└─────────────────────────────────────────────────────────────┘For complete RAG architecture with implementation patterns, see:
references/rag-architecture.md - Detailed 5-stage breakdownexamples/langchain-rag/basic_rag.py - Working implementationChunking is the most critical decision for RAG quality. Poor chunking breaks retrieval.
Default Recommendation:
Why these values:
Alternative strategies for special cases:
# Code-aware chunking (preserves functions/classes)
from langchain.text_splitter import RecursiveCharacterTextSplitter
code_splitter = RecursiveCharacterTextSplitter.from_language(
language="python",
chunk_size=512,
chunk_overlap=50
)
# Semantic chunking (splits on meaning, not tokens)
from langchain.text_splitter import SemanticChunker
semantic_splitter = SemanticChunker(
embeddings=embeddings,
breakpoint_threshold_type="percentile" # Split at semantic boundaries
)See: references/chunking-strategies.md for complete decision framework
Embedding quality directly impacts retrieval accuracy. Voyage AI is currently best-in-class.
Primary Recommendation: Voyage AI voyage-3
Cost-Effective Alternative: OpenAI text-embedding-3-small
Implementation:
from langchain_voyageai import VoyageAIEmbeddings
from langchain_openai import OpenAIEmbeddings
# Production (best quality)
embeddings = VoyageAIEmbeddings(
model="voyage-3",
voyage_api_key="your-api-key"
)
# Development (cost-effective)
embeddings = OpenAIEmbeddings(
model="text-embedding-3-small",
openai_api_key="your-api-key"
)See: references/embedding-strategies.md for complete provider comparison
Traditional metrics (BLEU, ROUGE) don't measure RAG quality. RAGAS provides LLM-as-judge evaluation.
4 Core Metrics:
| Metric | Measures | Good Score |
|---|---|---|
| Faithfulness | Factual consistency with retrieved context | > 0.8 |
| Answer Relevancy | Does answer address the user's question? | > 0.7 |
| Context Precision | Are retrieved chunks actually relevant? | > 0.6 |
| Context Recall | Were all necessary chunks retrieved? | > 0.7 |
Quick evaluation script:
# Run RAGAS evaluation (TOKEN-FREE script execution)
python scripts/evaluate_rag.py --dataset eval_data.json --output results.jsonManual implementation:
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy
dataset = {
"question": ["What is the capital of France?"],
"answer": ["Paris is the capital of France."],
"contexts": [["France's capital is Paris."]],
"ground_truth": ["Paris"]
}
result = evaluate(dataset, metrics=[faithfulness, answer_relevancy])
print(f"Faithfulness: {result['faithfulness']}")
print(f"Answer Relevancy: {result['answer_relevancy']}")See: references/evaluation-metrics.md for complete RAGAS implementation guide
Feature stores solve the "training-serving skew" problem by providing consistent feature computation.
Primary Recommendation: Feast - Open source, works with any backend (PostgreSQL, Redis, DynamoDB, S3, BigQuery, Snowflake)
Basic usage:
from feast import FeatureStore
store = FeatureStore(repo_path="feature_repo/")
# Online serving (low-latency)
features = store.get_online_features(
features=["user_features:total_orders"],
entity_rows=[{"user_id": 1001}]
).to_dict()See: references/feature-stores.md for complete Feast setup and alternatives (Tecton, Hopsworks)
LangChain is the primary framework for LLM orchestration with the largest ecosystem (24,215+ API reference snippets).
Context7 Library ID: /websites/langchain_oss_python_langchain (Trust: High, Snippets: 435)
Basic RAG Chain:
from langchain_core.prompts import ChatPromptTemplate
from langchain_qdrant import QdrantVectorStore
from langchain_voyageai import VoyageAIEmbeddings
# Setup retriever
vectorstore = QdrantVectorStore(
client=qdrant_client,
embedding=VoyageAIEmbeddings(model="voyage-3")
)
retriever = vectorstore.as_retriever(search_type="mmr", search_kwargs={"k": 5})
# Build chain
prompt = ChatPromptTemplate.from_template(
"Answer based on context:\n{context}\n\nQuestion: {question}"
)
chain = {"context": retriever, "question": lambda x: x} | prompt | ChatOpenAI() | StrOutputParser()
# Stream response
for chunk in chain.stream("What is the capital of France?"):
print(chunk, end="", flush=True)See: references/langchain-patterns.md - Complete LangChain 0.3+ patterns with streaming and hybrid search
Modern AI pipelines require workflow orchestration beyond cron jobs.
Primary Recommendation: Dagster (for ML/AI pipelines) - Asset-centric design, best lineage tracking, perfect for RAG
Example: Embedding Pipeline
from dagster import asset
from langchain_voyageai import VoyageAIEmbeddings
@asset
def raw_documents():
"""Load documents from S3."""
return documents
@asset
def chunked_documents(raw_documents):
"""Split into 512-token chunks with 50-token overlap."""
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=50)
return splitter.split_documents(raw_documents)
@asset
def embedded_documents(chunked_documents):
"""Generate embeddings with Voyage AI."""
embeddings = VoyageAIEmbeddings(model="voyage-3")
return embeddings.embed_documents([doc.page_content for doc in chunked_documents])See: references/orchestration-tools.md for complete Dagster patterns and alternatives (Prefect, Airflow 3.0, dbt)
The ai-chat skill consumes RAG pipeline outputs for streaming responses.
Backend API (FastAPI):
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
@app.post("/api/rag/stream")
async def stream_rag(query: str):
async def generate():
chain = RetrievalQA.from_chain_type(llm=OpenAI(streaming=True), retriever=vectorstore.as_retriever())
async for chunk in chain.astream(query):
yield chunk
return StreamingResponse(generate(), media_type="text/plain")See: references/rag-architecture.md for complete frontend integration patterns
The search-filter skill uses semantic search backends for vector similarity.
Backend (Qdrant + Voyage AI):
from qdrant_client import QdrantClient
from langchain_voyageai import VoyageAIEmbeddings
@app.post("/api/search/semantic")
async def semantic_search(query: str, filters: dict):
query_vector = VoyageAIEmbeddings(model="voyage-3").embed_query(query)
results = QdrantClient().search(
collection_name="documents",
query_vector=query_vector,
query_filter=filters,
limit=10
)
return {"results": results}Primary Recommendation: LakeFS (acquired DVC team November 2025)
Git-like operations on data lakes: branch, commit, merge, time travel. Works with S3/Azure/GCS.
import lakefs
branch = lakefs.Branch("main").create("experiment-voyage-3")
branch.commit("Updated embeddings to voyage-3")
branch.merge_into("main")See: references/data-versioning.md for complete LakeFS setup
1. Set up vector database:
# Run Qdrant setup script (TOKEN-FREE execution)
python scripts/setup_qdrant.py --collection docs --dimension 10242. Chunk and embed documents:
# Chunk documents (TOKEN-FREE execution)
python scripts/chunk_documents.py \
--input data/documents/ \
--chunk-size 512 \
--overlap 50 \
--output data/chunks/3. Implement RAG pipeline:
See examples/langchain-rag/basic_rag.py for complete working example.
4. Evaluate with RAGAS:
# Run evaluation (TOKEN-FREE execution)
python scripts/evaluate_rag.py \
--dataset data/eval_qa.json \
--output results/ragas_metrics.json5. Deploy with orchestration:
See examples/dagster-pipelines/embedding_pipeline.py for production deployment.
Required Python packages:
# Core RAG
pip install langchain langchain-core langchain-openai langchain-voyageai langchain-qdrant
# Vector database
pip install qdrant-client
# Evaluation
pip install ragas datasets
# Feature stores
pip install feast
# Orchestration
pip install dagster dagster-webserver
# Data versioning
pip install lakefs-clientOptional for alternatives:
# LlamaIndex (alternative to LangChain)
pip install llama-index
# dbt (SQL transformations)
pip install dbt-core dbt-postgres
# Prefect (alternative orchestration)
pip install prefectCommon Issues:
1. Poor retrieval quality - Check chunk size (try 512 tokens), increase overlap (50-100), try hybrid search, re-rank with Cohere
2. Slow embedding generation - Batch documents (100-1000), use async APIs, cache with Redis, use smaller model for dev
3. High LLM costs - Reduce retrieved chunks (k=3), use cheaper re-ranking models, cache frequent queries
See: references/rag-architecture.md for complete troubleshooting guide
Chunking: Default to 512 tokens with 50-token overlap. Use semantic chunking for complex documents. Preserve code structure for source code.
Embeddings: Use Voyage AI voyage-3 for production, OpenAI text-embedding-3-small for development. Never mix embedding models (re-embed everything if changing).
Evaluation: Run RAGAS metrics on every pipeline change. Maintain test dataset of 50+ question-answer pairs. Track metrics over time.
Orchestration: Use Dagster for ML/AI pipelines, dbt for SQL transformations only. Version control all pipeline code.
Frontend Integration: Always stream LLM responses. Implement retry logic. Show citations/sources to users. Handle empty results gracefully.
Reference Documentation:
references/rag-architecture.md - Complete RAG pipeline guidereferences/chunking-strategies.md - Decision framework for chunkingreferences/embedding-strategies.md - Embedding model comparisonreferences/langchain-patterns.md - LangChain 0.3+ patternsreferences/feature-stores.md - Feast setup and alternativesreferences/evaluation-metrics.md - RAGAS implementation guideWorking Examples:
examples/langchain-rag/basic_rag.py - Simple RAG chainexamples/langchain-rag/streaming_rag.py - Streaming responsesexamples/langchain-rag/hybrid_search.py - Vector + BM25examples/llamaindex-agents/query_engine.py - LlamaIndex alternativeexamples/feast-features/ - Complete feature store setupexamples/dagster-pipelines/embedding_pipeline.py - Production pipelineExecutable Scripts (TOKEN-FREE):
scripts/evaluate_rag.py - RAGAS evaluation runnerscripts/chunk_documents.py - Document chunking utilityscripts/benchmark_retrieval.py - Retrieval quality benchmarkscripts/setup_qdrant.py - Qdrant collection setup© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 28 other files (scripts, references) in skills/ai-data-engineering of ancoleman/ai-design-components.
Open the folder on GitHubat commit 76551b7
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ancoleman/ai-design-components, which our catalogue first saw on October 7, 2026.
AI Data Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AI Data Engineering this skillancoleman/ai-design-components | 526 | 1 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Migrating Dagster To Airflowastronomer/agents | 451 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Engineering Data Pipelinestelagod/code-abyss | 243 | — | ~236 | Automated safety check: Pass | MIT | |
| Senior Data Engineerborghei/Claude-Skills | 881 | — | ~1.4k | Automated safety check: Pass | MIT | |
| ML Pipeline ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.9k | Automated safety check: Pass | MIT | |
| ML Pipeline Workflowwshobson/agents | 40k | 13 repos | ~1.8k | Automated safety check: Pass | MIT |
astronomer/agents
Guide for migrating Dagster projects to Apache Airflow 3 on Astro.
telagod/code-abyss
Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.
borghei/Claude-Skills
Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.
Jeffallan/claude-skills
Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.
wshobson/agents
Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.
sickn33/agentic-awesome-skills
Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.
ancoleman/ai-design-components
Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.
ancoleman/ai-design-components
Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.
ancoleman/ai-design-components
Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.
ancoleman/ai-design-components
Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.
ancoleman/ai-design-components
Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.
ancoleman/ai-design-components
Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.
Works with
Categories
Data pipelines, feature stores, and embedding generation for AI/ML systems. AI Data Engineering is an agent skill from ancoleman/ai-design-components. Data pipelines, feature stores, and embedding generation for AI/ML systems.
AI Data Engineering fits situations like: building RAG pipelines; ML feature serving; data transformations.
Run `npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a claude-code`. Or copy the skill folder (skills/ai-data-engineering in ancoleman/ai-design-components) into .claude/skills/ai-data-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a codex`. Or copy the skill folder (skills/ai-data-engineering in ancoleman/ai-design-components) into .agents/skills/ai-data-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-data-engineering, .gemini/skills/ai-data-engineering, .github/skills/ai-data-engineering and .opencode/skills/ai-data-engineering in your project.
Going by SKILL.md and its folder, AI Data Engineering needs Python for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
AI Data Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 29k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with AI Data Engineering: Migrating Dagster To Airflow (astronomer/agents, 451 stars), Engineering Data Pipelines (telagod/code-abyss, 243 stars), Senior Data Engineer (borghei/Claude-Skills, 881 stars) and ML Pipeline Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.
Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.