Agent skill

AI Data Engineering

by ancoleman in ancoleman/ai-design-components

Data pipelines, feature stores, and embedding generation for AI/ML systems.

MITAuto-check passedData & Analytics

Install AI Data Engineering

skills CLI
$ npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ancoleman/ai-design-components ai-data-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-data-engineering .claude/skills/ai-data-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-data-engineering
GitHub stars
526
Used in
1 other repo
Token cost
~3.5k tokens
SKILL.md length
798 words
Files
29 (incl. scripts, references)
Skills in repo
75
Repo updated
First seen
Licence
MIT

At a glance

Data pipelines, feature stores, and embedding generation for AI/ML systems.

  • Building RAG pipelines
  • SKILL.md covers Purpose, When to Use, RAG Pipeline Architecture and Chunking Strategies, plus 12 more sections
  • Runs Python scripts from its folder; calls pip and python
  • ML feature serving

What it does

AI Data Engineering is an agent skill from ancoleman/ai-design-components. Data pipelines, feature stores, and embedding generation for AI/ML systems. Use when building RAG pipelines, ML feature serving, or data transformations. Covers feature stores (Feast, Tecton), embedding pipelines, chunking strategies, orchestration (Dagster, Prefect, Airflow), dbt transformations, data versioning (LakeFS), and experiment tracking (MLflow, W&B).

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 33 other files, including scripts and reference files (for example `examples/dagster-pipelines/embedding_pipeline.py`, `examples/feast-features/README.md` and `examples/feast-features/setup_features.py`).

It sits in Data & Analytics, covering Data pipelines and ETL, Embeddings and MLOps. It works with MLflow, dbt, Apache Airflow and Dagster. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.

When your agent uses it

  • Building RAG pipelines
  • ML feature serving
  • Data transformations

Example prompts

  • “/ai-data-engineering”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Data Engineering loads about 3.5k tokens when it runs, and up to ~32k if it reads all its reference files. Until then it costs about 96 tokens; SKILL.md has 798 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~32k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 798 words, ~3,501 tokens.

Download SKILL.mdSave it as .claude/skills/ai-data-engineering/SKILL.md (or your agent's skills folder). This skill also uses 28 other files; get the full folder from GitHub.
name
ai-data-engineering
description
Data pipelines, feature stores, and embedding generation for AI/ML systems. Use when building RAG pipelines, ML feature serving, or data transformations. Covers feature stores (Feast, Tecton), embedding pipelines, chunking strategies, orchestration (Dagster, Prefect, Airflow), dbt transformations, data versioning (LakeFS), and experiment tracking (MLflow, W&B).

AI Data Engineering

Purpose

Build data infrastructure for AI/ML systems including RAG pipelines, feature stores, and embedding generation. Provides architecture patterns, orchestration workflows, and evaluation metrics for production AI applications.

When to Use

Use this skill when:

  • Building RAG (Retrieval-Augmented Generation) pipelines
  • Implementing semantic search or vector databases
  • Setting up ML feature stores for real-time serving
  • Creating embedding generation pipelines
  • Evaluating RAG quality with RAGAS metrics
  • Orchestrating data workflows for AI systems
  • Integrating with frontend skills (ai-chat, search-filter)

Skip this skill if:

  • Building traditional CRUD applications (use databases-relational)
  • Simple key-value storage (use databases-nosql)
  • No AI/ML components in the application

RAG Pipeline Architecture

RAG pipelines have 5 distinct stages. Understanding this architecture is critical for production implementations.

┌─────────────────────────────────────────────────────────────┐
│                    RAG Pipeline (5 Stages)                   │
├─────────────────────────────────────────────────────────────┤
│                                                              │
│  1. INGESTION → Load documents (PDF, DOCX, Markdown)        │
│  2. INDEXING → Chunk (512 tokens) + Embed + Store           │
│  3. RETRIEVAL → Query embedding + Vector search + Filters   │
│  4. GENERATION → Context injection + LLM streaming          │
│  5. EVALUATION → RAGAS metrics (faithfulness, relevancy)    │
│                                                              │
└─────────────────────────────────────────────────────────────┘

For complete RAG architecture with implementation patterns, see:

  • references/rag-architecture.md - Detailed 5-stage breakdown
  • examples/langchain-rag/basic_rag.py - Working implementation

Chunking Strategies

Chunking is the most critical decision for RAG quality. Poor chunking breaks retrieval.

Default Recommendation:

  • Size: 512 tokens
  • Overlap: 50-100 tokens
  • Method: Fixed token-based

Why these values:

  • Too small (<256 tokens): Loses context, requires many retrievals
  • Too large (>1024 tokens): Includes irrelevant content, hits token limits
  • Overlap prevents information loss at chunk boundaries

Alternative strategies for special cases:

python
# Code-aware chunking (preserves functions/classes)
from langchain.text_splitter import RecursiveCharacterTextSplitter

code_splitter = RecursiveCharacterTextSplitter.from_language(
    language="python",
    chunk_size=512,
    chunk_overlap=50
)

# Semantic chunking (splits on meaning, not tokens)
from langchain.text_splitter import SemanticChunker

semantic_splitter = SemanticChunker(
    embeddings=embeddings,
    breakpoint_threshold_type="percentile"  # Split at semantic boundaries
)

See: references/chunking-strategies.md for complete decision framework

Embedding Generation

Embedding quality directly impacts retrieval accuracy. Voyage AI is currently best-in-class.

Primary Recommendation: Voyage AI voyage-3

  • Dimensions: 1024
  • MTEB Score: 69.0 (highest as of Dec 2025)
  • Cost: $$$ but 9.74% better than OpenAI
  • Use for: Production systems requiring best retrieval quality

Cost-Effective Alternative: OpenAI text-embedding-3-small

  • Dimensions: 1536
  • MTEB Score: 62.3
  • Cost: $ (5x cheaper than voyage-3)
  • Use for: Development, prototyping, cost-sensitive applications

Implementation:

python
from langchain_voyageai import VoyageAIEmbeddings
from langchain_openai import OpenAIEmbeddings

# Production (best quality)
embeddings = VoyageAIEmbeddings(
    model="voyage-3",
    voyage_api_key="your-api-key"
)

# Development (cost-effective)
embeddings = OpenAIEmbeddings(
    model="text-embedding-3-small",
    openai_api_key="your-api-key"
)

See: references/embedding-strategies.md for complete provider comparison

RAGAS Evaluation Metrics

Traditional metrics (BLEU, ROUGE) don't measure RAG quality. RAGAS provides LLM-as-judge evaluation.

4 Core Metrics:

MetricMeasuresGood Score
FaithfulnessFactual consistency with retrieved context> 0.8
Answer RelevancyDoes answer address the user's question?> 0.7
Context PrecisionAre retrieved chunks actually relevant?> 0.6
Context RecallWere all necessary chunks retrieved?> 0.7

Quick evaluation script:

bash
# Run RAGAS evaluation (TOKEN-FREE script execution)
python scripts/evaluate_rag.py --dataset eval_data.json --output results.json

Manual implementation:

python
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy

dataset = {
    "question": ["What is the capital of France?"],
    "answer": ["Paris is the capital of France."],
    "contexts": [["France's capital is Paris."]],
    "ground_truth": ["Paris"]
}

result = evaluate(dataset, metrics=[faithfulness, answer_relevancy])
print(f"Faithfulness: {result['faithfulness']}")
print(f"Answer Relevancy: {result['answer_relevancy']}")

See: references/evaluation-metrics.md for complete RAGAS implementation guide

Feature Stores

Feature stores solve the "training-serving skew" problem by providing consistent feature computation.

Primary Recommendation: Feast - Open source, works with any backend (PostgreSQL, Redis, DynamoDB, S3, BigQuery, Snowflake)

Basic usage:

python
from feast import FeatureStore
store = FeatureStore(repo_path="feature_repo/")

# Online serving (low-latency)
features = store.get_online_features(
    features=["user_features:total_orders"],
    entity_rows=[{"user_id": 1001}]
).to_dict()

See: references/feature-stores.md for complete Feast setup and alternatives (Tecton, Hopsworks)

LangChain Orchestration

LangChain is the primary framework for LLM orchestration with the largest ecosystem (24,215+ API reference snippets).

Context7 Library ID: /websites/langchain_oss_python_langchain (Trust: High, Snippets: 435)

Basic RAG Chain:

python
from langchain_core.prompts import ChatPromptTemplate
from langchain_qdrant import QdrantVectorStore
from langchain_voyageai import VoyageAIEmbeddings

# Setup retriever
vectorstore = QdrantVectorStore(
    client=qdrant_client,
    embedding=VoyageAIEmbeddings(model="voyage-3")
)
retriever = vectorstore.as_retriever(search_type="mmr", search_kwargs={"k": 5})

# Build chain
prompt = ChatPromptTemplate.from_template(
    "Answer based on context:\n{context}\n\nQuestion: {question}"
)
chain = {"context": retriever, "question": lambda x: x} | prompt | ChatOpenAI() | StrOutputParser()

# Stream response
for chunk in chain.stream("What is the capital of France?"):
    print(chunk, end="", flush=True)

See: references/langchain-patterns.md - Complete LangChain 0.3+ patterns with streaming and hybrid search

Orchestration Tools

Modern AI pipelines require workflow orchestration beyond cron jobs.

Primary Recommendation: Dagster (for ML/AI pipelines) - Asset-centric design, best lineage tracking, perfect for RAG

Example: Embedding Pipeline

python
from dagster import asset
from langchain_voyageai import VoyageAIEmbeddings

@asset
def raw_documents():
    """Load documents from S3."""
    return documents

@asset
def chunked_documents(raw_documents):
    """Split into 512-token chunks with 50-token overlap."""
    from langchain.text_splitter import RecursiveCharacterTextSplitter
    splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=50)
    return splitter.split_documents(raw_documents)

@asset
def embedded_documents(chunked_documents):
    """Generate embeddings with Voyage AI."""
    embeddings = VoyageAIEmbeddings(model="voyage-3")
    return embeddings.embed_documents([doc.page_content for doc in chunked_documents])

See: references/orchestration-tools.md for complete Dagster patterns and alternatives (Prefect, Airflow 3.0, dbt)

Integration with Frontend Skills

ai-chat Skill → RAG Backend

The ai-chat skill consumes RAG pipeline outputs for streaming responses.

Backend API (FastAPI):

python
from fastapi import FastAPI
from fastapi.responses import StreamingResponse

@app.post("/api/rag/stream")
async def stream_rag(query: str):
    async def generate():
        chain = RetrievalQA.from_chain_type(llm=OpenAI(streaming=True), retriever=vectorstore.as_retriever())
        async for chunk in chain.astream(query):
            yield chunk
    return StreamingResponse(generate(), media_type="text/plain")

See: references/rag-architecture.md for complete frontend integration patterns

Show full SKILL.md (317 more words)Show less

The search-filter skill uses semantic search backends for vector similarity.

Backend (Qdrant + Voyage AI):

python
from qdrant_client import QdrantClient
from langchain_voyageai import VoyageAIEmbeddings

@app.post("/api/search/semantic")
async def semantic_search(query: str, filters: dict):
    query_vector = VoyageAIEmbeddings(model="voyage-3").embed_query(query)
    results = QdrantClient().search(
        collection_name="documents",
        query_vector=query_vector,
        query_filter=filters,
        limit=10
    )
    return {"results": results}

Data Versioning

Primary Recommendation: LakeFS (acquired DVC team November 2025)

Git-like operations on data lakes: branch, commit, merge, time travel. Works with S3/Azure/GCS.

python
import lakefs

branch = lakefs.Branch("main").create("experiment-voyage-3")
branch.commit("Updated embeddings to voyage-3")
branch.merge_into("main")

See: references/data-versioning.md for complete LakeFS setup

Quick Start Workflow

1. Set up vector database:

bash
# Run Qdrant setup script (TOKEN-FREE execution)
python scripts/setup_qdrant.py --collection docs --dimension 1024

2. Chunk and embed documents:

bash
# Chunk documents (TOKEN-FREE execution)
python scripts/chunk_documents.py \
  --input data/documents/ \
  --chunk-size 512 \
  --overlap 50 \
  --output data/chunks/

3. Implement RAG pipeline:

See examples/langchain-rag/basic_rag.py for complete working example.

4. Evaluate with RAGAS:

bash
# Run evaluation (TOKEN-FREE execution)
python scripts/evaluate_rag.py \
  --dataset data/eval_qa.json \
  --output results/ragas_metrics.json

5. Deploy with orchestration:

See examples/dagster-pipelines/embedding_pipeline.py for production deployment.

Dependencies

Required Python packages:

bash
# Core RAG
pip install langchain langchain-core langchain-openai langchain-voyageai langchain-qdrant

# Vector database
pip install qdrant-client

# Evaluation
pip install ragas datasets

# Feature stores
pip install feast

# Orchestration
pip install dagster dagster-webserver

# Data versioning
pip install lakefs-client

Optional for alternatives:

bash
# LlamaIndex (alternative to LangChain)
pip install llama-index

# dbt (SQL transformations)
pip install dbt-core dbt-postgres

# Prefect (alternative orchestration)
pip install prefect

Troubleshooting

Common Issues:

1. Poor retrieval quality - Check chunk size (try 512 tokens), increase overlap (50-100), try hybrid search, re-rank with Cohere

2. Slow embedding generation - Batch documents (100-1000), use async APIs, cache with Redis, use smaller model for dev

3. High LLM costs - Reduce retrieved chunks (k=3), use cheaper re-ranking models, cache frequent queries

See: references/rag-architecture.md for complete troubleshooting guide

Best Practices

Chunking: Default to 512 tokens with 50-token overlap. Use semantic chunking for complex documents. Preserve code structure for source code.

Embeddings: Use Voyage AI voyage-3 for production, OpenAI text-embedding-3-small for development. Never mix embedding models (re-embed everything if changing).

Evaluation: Run RAGAS metrics on every pipeline change. Maintain test dataset of 50+ question-answer pairs. Track metrics over time.

Orchestration: Use Dagster for ML/AI pipelines, dbt for SQL transformations only. Version control all pipeline code.

Frontend Integration: Always stream LLM responses. Implement retry logic. Show citations/sources to users. Handle empty results gracefully.

Additional Resources

Reference Documentation:

  • references/rag-architecture.md - Complete RAG pipeline guide
  • references/chunking-strategies.md - Decision framework for chunking
  • references/embedding-strategies.md - Embedding model comparison
  • references/langchain-patterns.md - LangChain 0.3+ patterns
  • references/feature-stores.md - Feast setup and alternatives
  • references/evaluation-metrics.md - RAGAS implementation guide

Working Examples:

  • examples/langchain-rag/basic_rag.py - Simple RAG chain
  • examples/langchain-rag/streaming_rag.py - Streaming responses
  • examples/langchain-rag/hybrid_search.py - Vector + BM25
  • examples/llamaindex-agents/query_engine.py - LlamaIndex alternative
  • examples/feast-features/ - Complete feature store setup
  • examples/dagster-pipelines/embedding_pipeline.py - Production pipeline

Executable Scripts (TOKEN-FREE):

  • scripts/evaluate_rag.py - RAGAS evaluation runner
  • scripts/chunk_documents.py - Document chunking utility
  • scripts/benchmark_retrieval.py - Retrieval quality benchmark
  • scripts/setup_qdrant.py - Qdrant collection setup

© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 28 other files (scripts, references) in skills/ai-data-engineering of ancoleman/ai-design-components.

  • SKILL.md
  • examples/dagster-pipelines/embedding_pipeline.py
  • examples/feast-features/README.md
  • examples/feast-features/requirements.txt
  • examples/feast-features/setup_features.py
  • examples/langchain-rag/README.md
  • examples/langchain-rag/basic_rag.py
  • examples/langchain-rag/hybrid_search.py
  • examples/langchain-rag/main.py
  • examples/langchain-rag/requirements.txt
  • examples/langchain-rag/streaming_rag.py
  • examples/llamaindex-agents/README.md
  • examples/llamaindex-agents/query_engine.py
  • examples/llamaindex-agents/requirements.txt
  • outputs.yaml
  • references
  • … and 13 more

Open the folder on GitHubat commit 76551b7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ancoleman/ai-design-components, which our catalogue first saw on October 7, 2026.

Compare with similar skills

AI Data Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Data Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Data Engineering this skillancoleman/ai-design-components5261 repos~3.5kAutomated safety check: PassMIT
Migrating Dagster To Airflowastronomer/agents451—~3.8kAutomated safety check: PassApache-2.0
Engineering Data Pipelinestelagod/code-abyss243—~236Automated safety check: PassMIT
Senior Data Engineerborghei/Claude-Skills881—~1.4kAutomated safety check: PassMIT
ML Pipeline ExpertJeffallan/claude-skills12k1 repos~1.9kAutomated safety check: PassMIT
ML Pipeline Workflowwshobson/agents40k13 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    451 GitHub stars~3.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Engineering Data Pipelines

    telagod/code-abyss

    Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.

    243 GitHub stars~236 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    borghei/Claude-Skills

    Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • ML Pipeline Workflow

    wshobson/agents

    Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.

    40k GitHub starsUsed in 13 repos~1.8k tokens
    DevOps & CloudAuto-check passed
  • AI Pipeline Orchestration

    sickn33/agentic-awesome-skills

    Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed

More from ancoleman/ai-design-components

All 75 skills in this repo
  • Building AI Chat

    ancoleman/ai-design-components

    Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.

    526 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Building Forms

    ancoleman/ai-design-components

    Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.

    526 GitHub stars~3.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Building Tables

    ancoleman/ai-design-components

    Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.

    526 GitHub stars~1.8k tokensUpdated 10 mo ago
    Auto-check passed
  • Creating Dashboards

    ancoleman/ai-design-components

    Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.

    526 GitHub stars~3.5k tokensUpdated 10 mo ago
    Auto-check passed
  • Designing Layouts

    ancoleman/ai-design-components

    Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.

    526 GitHub stars~1.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Displaying Timelines

    ancoleman/ai-design-components

    Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.

    526 GitHub stars~2.7k tokensUpdated 10 mo ago
    Auto-check passed

Questions about AI Data Engineering

What does AI Data Engineering do?

Data pipelines, feature stores, and embedding generation for AI/ML systems. AI Data Engineering is an agent skill from ancoleman/ai-design-components. Data pipelines, feature stores, and embedding generation for AI/ML systems.

When should I use AI Data Engineering?

AI Data Engineering fits situations like: building RAG pipelines; ML feature serving; data transformations.

How do I install AI Data Engineering in Claude Code?

Run `npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a claude-code`. Or copy the skill folder (skills/ai-data-engineering in ancoleman/ai-design-components) into .claude/skills/ai-data-engineering in your project. Claude Code loads it when a task matches its description.

How do I install AI Data Engineering in Codex?

Run `npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a codex`. Or copy the skill folder (skills/ai-data-engineering in ancoleman/ai-design-components) into .agents/skills/ai-data-engineering in your project. Codex loads it when a task matches its description.

Can I use AI Data Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill ai-data-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-data-engineering, .gemini/skills/ai-data-engineering, .github/skills/ai-data-engineering and .opencode/skills/ai-data-engineering in your project.

What does AI Data Engineering need to run?

Going by SKILL.md and its folder, AI Data Engineering needs Python for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3.

Does AI Data Engineering access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is AI Data Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does AI Data Engineering use?

AI Data Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Data Engineering use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 29k tokens, read only when the agent opens those files.

What are the alternatives to AI Data Engineering?

Skills that share tags, products or a category with AI Data Engineering: Migrating Dagster To Airflow (astronomer/agents, 451 stars), Engineering Data Pipelines (telagod/code-abyss, 243 stars), Senior Data Engineer (borghei/Claude-Skills, 881 stars) and ML Pipeline Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Data Engineering?

ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.

Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.