AI Pipeline Orchestration
sickn33/agentic-awesome-skills
Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.
Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install BagelHole/DevOps-Security-Agent-Skills ai-pipeline-orchestration --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/devops/ai/ai-pipeline-orchestration .claude/skills/ai-pipeline-orchestration && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-pipeline-orchestration" agent skill from https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-pipeline-orchestration into .claude/skills/ai-pipeline-orchestration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-pipeline-orchestration", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-pipeline-orchestrationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install BagelHole/DevOps-Security-Agent-Skills ai-pipeline-orchestration --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/devops/ai/ai-pipeline-orchestration .agents/skills/ai-pipeline-orchestration && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-pipeline-orchestration" agent skill from https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-pipeline-orchestration into .agents/skills/ai-pipeline-orchestration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-pipeline-orchestration", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install BagelHole/DevOps-Security-Agent-Skills ai-pipeline-orchestration --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/devops/ai/ai-pipeline-orchestration .cursor/skills/ai-pipeline-orchestration && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-pipeline-orchestration" agent skill from https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-pipeline-orchestration into .cursor/skills/ai-pipeline-orchestration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-pipeline-orchestration", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/BagelHole/DevOps-Security-Agent-Skills.git --path devops/ai/ai-pipeline-orchestration--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install BagelHole/DevOps-Security-Agent-Skills ai-pipeline-orchestration --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/devops/ai/ai-pipeline-orchestration .gemini/skills/ai-pipeline-orchestration && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-pipeline-orchestration" agent skill from https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-pipeline-orchestration into .gemini/skills/ai-pipeline-orchestration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-pipeline-orchestration", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install BagelHole/DevOps-Security-Agent-Skills ai-pipeline-orchestrationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/devops/ai/ai-pipeline-orchestration .github/skills/ai-pipeline-orchestration && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-pipeline-orchestration" agent skill from https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-pipeline-orchestration into .github/skills/ai-pipeline-orchestration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-pipeline-orchestration", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install BagelHole/DevOps-Security-Agent-Skills ai-pipeline-orchestration --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/devops/ai/ai-pipeline-orchestration .opencode/skills/ai-pipeline-orchestration && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-pipeline-orchestration" agent skill from https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-pipeline-orchestration into .opencode/skills/ai-pipeline-orchestration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-pipeline-orchestration", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-pipeline-orchestrationOrchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.
AI Pipeline Orchestration is an agent skill from BagelHole/DevOps-Security-Agent-Skills. Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster. Build reliable, observable, and retriable workflows for production AI systems.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL and Fine-tuning. It works with Apache Airflow and Dagster. The repository describes itself as: Agent-ready DevOps, security, infrastructure, and compliance knowledge base with 80+ skills across Kubernetes, Terraform, AWS/Azure/GCP, AI platform operations, container… The licence is MIT.
Read from SKILL.md and the folder at commit 0365f57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
prefectpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AI Pipeline Orchestration loads about 2.3k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 207 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from BagelHole/DevOps-Security-Agent-Skills at commit 0365f57, republished under its MIT licence (© BagelHole). 207 words, ~2,296 tokens.
.claude/skills/ai-pipeline-orchestration/SKILL.md (or your agent's skills folder).Build reliable, observable AI workflows — from document ingestion to batch inference to model training pipelines.
Use this skill when:
| Tool | Best For | Complexity | GPU Jobs |
|---|---|---|---|
| Prefect | Modern Python-first; easy to adopt | Low | Good |
| Airflow | Complex DAGs; large teams; existing usage | High | Good |
| Dagster | Asset-centric; strong data lineage | Medium | Excellent |
| Temporal | Long-running workflows; reliability-first | Medium | Good |
pip install prefect prefect-kubernetes
# Start Prefect server (or use Prefect Cloud)
prefect server start
# In another terminal
prefect worker start --pool default-agent-poolfrom prefect import flow, task, get_run_logger
from prefect.tasks import task_input_hash
from datetime import timedelta
import hashlib
@task(cache_key_fn=task_input_hash, cache_expiration=timedelta(hours=24))
def fetch_documents(source_url: str) -> list[dict]:
"""Fetch documents from source; cached to avoid re-fetching."""
logger = get_run_logger()
logger.info(f"Fetching from {source_url}")
# ... fetch logic
return documents
@task(retries=3, retry_delay_seconds=30)
def chunk_and_embed(documents: list[dict]) -> list[dict]:
"""Chunk documents and generate embeddings with retry on failure."""
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-large-en-v1.5")
chunks = []
for doc in documents:
doc_chunks = chunk_text(doc["content"])
embeddings = model.encode(doc_chunks, batch_size=64)
for chunk, emb in zip(doc_chunks, embeddings):
chunks.append({"text": chunk, "embedding": emb.tolist(),
"source": doc["url"], "doc_hash": doc["hash"]})
return chunks
@task(retries=2)
def upsert_to_vector_store(chunks: list[dict]) -> int:
"""Upsert embeddings to Qdrant, skip unchanged documents."""
from qdrant_client import QdrantClient
client = QdrantClient("http://qdrant:6333")
client.upsert(collection_name="knowledge-base", points=[...])
return len(chunks)
@flow(name="rag-ingestion", log_prints=True)
def rag_ingestion_pipeline(sources: list[str]):
"""Full RAG ingestion flow — runs daily."""
logger = get_run_logger()
total = 0
for source in sources:
docs = fetch_documents(source)
chunks = chunk_and_embed(docs)
count = upsert_to_vector_store(chunks)
total += count
logger.info(f"Ingested {count} chunks from {source}")
logger.info(f"Pipeline complete: {total} total chunks indexed")
if __name__ == "__main__":
rag_ingestion_pipeline.serve(
name="daily-rag-ingestion",
cron="0 2 * * *", # 2 AM daily
parameters={"sources": ["https://docs.myapp.com", "https://api.myapp.com/kb"]},
)from prefect import flow, task
from prefect.concurrency.sync import concurrency
import asyncio
from openai import AsyncOpenAI
@task(retries=3, retry_delay_seconds=60)
async def process_batch(items: list[dict], model: str = "gpt-4o-mini") -> list[dict]:
"""Process a batch of items through LLM with rate limit protection."""
client = AsyncOpenAI()
async with concurrency("openai-api", occupy=len(items)): # rate limit
tasks = [
client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": item["prompt"]}],
max_tokens=256,
)
for item in items
]
responses = await asyncio.gather(*tasks, return_exceptions=True)
results = []
for item, response in zip(items, responses):
if isinstance(response, Exception):
results.append({**item, "error": str(response), "output": None})
else:
results.append({**item, "output": response.choices[0].message.content})
return results
@flow(name="batch-llm-inference")
async def batch_inference_flow(input_file: str, output_file: str, batch_size: int = 50):
import json
items = [json.loads(line) for line in open(input_file)]
batches = [items[i:i+batch_size] for i in range(0, len(items), batch_size)]
all_results = []
for batch in batches:
results = await process_batch(batch)
all_results.extend(results)
with open(output_file, "w") as f:
for result in all_results:
f.write(json.dumps(result) + "\n")
return len(all_results)from airflow.decorators import dag, task
from airflow.providers.cncf.kubernetes.operators.pod import KubernetesPodOperator
from datetime import datetime
from kubernetes.client import models as k8s
@dag(
dag_id="llm_fine_tuning",
schedule="@weekly",
start_date=datetime(2025, 1, 1),
catchup=False,
tags=["ai", "training"],
)
def llm_fine_tuning_dag():
@task
def prepare_dataset() -> str:
"""Download and preprocess training data."""
# ... data prep logic
return "s3://my-bucket/training-data/2025-03-01/"
train = KubernetesPodOperator(
task_id="train_model",
name="llm-training-job",
namespace="ml",
image="nvcr.io/nvidia/pytorch:24.05-py3",
cmds=["accelerate", "launch", "-m", "axolotl.cli.train", "/config/config.yaml"],
resources=k8s.V1ResourceRequirements(
limits={"nvidia.com/gpu": "4", "memory": "320Gi"},
requests={"nvidia.com/gpu": "4"},
),
node_selector={"nvidia.com/gpu.product": "A100-SXM4-80GB"},
volumes=[...],
volume_mounts=[...],
get_logs=True,
is_delete_operator_pod=True,
)
@task
def evaluate_model(dataset_path: str) -> dict:
"""Run evals; fail pipeline if quality drops."""
metrics = run_evals()
if metrics["accuracy"] < 0.85:
raise ValueError(f"Model quality too low: {metrics}")
return metrics
@task
def deploy_model(metrics: dict):
"""Push merged model to HF Hub and update vLLM config."""
update_serving_config(new_model="org/fine-tuned-v2")
dataset = prepare_dataset()
train.set_upstream(dataset)
eval_result = evaluate_model(dataset)
eval_result.set_upstream(train)
deploy_model(eval_result)
llm_fine_tuning_dag()from dagster import asset, AssetExecutionContext, define_asset_job, ScheduleDefinition
@asset(description="Raw documents fetched from knowledge sources")
def raw_documents(context: AssetExecutionContext) -> list[dict]:
context.log.info("Fetching documents...")
return fetch_all_documents()
@asset(
deps=[raw_documents],
description="Chunked and embedded document vectors",
)
def document_embeddings(context: AssetExecutionContext, raw_documents) -> int:
chunks = process_and_embed(raw_documents)
context.log.info(f"Generated {len(chunks)} embeddings")
upsert_to_qdrant(chunks)
return len(chunks)
@asset(
deps=[document_embeddings],
description="RAG system quality metrics",
)
def rag_quality_metrics(context: AssetExecutionContext) -> dict:
metrics = evaluate_rag_system()
context.add_output_metadata({"ragas_score": metrics["ragas_score"]})
return metrics
# Schedule: refresh embeddings nightly
nightly_refresh = ScheduleDefinition(
job=define_asset_job("rag_refresh_job", [raw_documents, document_embeddings]),
cron_schedule="0 1 * * *",
)concurrency limits in Prefect or pool slots in Airflow to respect external rate limits.© BagelHole, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in devops/ai/ai-pipeline-orchestration of BagelHole/DevOps-Security-Agent-Skills.
Open the folder on GitHubat commit 0365f57
AI Pipeline Orchestration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AI Pipeline Orchestration this skillBagelHole/DevOps-Security-Agent-Skills | 1.2k | — | ~2.3k | Automated safety check: Pass | MIT | |
| AI Pipeline Orchestrationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| AI Data Engineeringancoleman/ai-design-components | 525 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Migrating Dagster To Airflowastronomer/agents | 451 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Engineering Data Pipelinestelagod/code-abyss | 244 | — | ~236 | Automated safety check: Pass | MIT | |
| Senior Data Engineerborghei/Claude-Skills | 891 | — | ~1.4k | Automated safety check: Pass | MIT |
sickn33/agentic-awesome-skills
Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.
ancoleman/ai-design-components
Data pipelines, feature stores, and embedding generation for AI/ML systems.
astronomer/agents
Guide for migrating Dagster projects to Apache Airflow 3 on Astro.
telagod/code-abyss
Data engineering knowledge reference covering Airflow, Dagster, Kafka Streams, Flink, dbt, and data quality patterns.
borghei/Claude-Skills
Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka.
wshobson/agents
Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.
BagelHole/DevOps-Security-Agent-Skills
Manage secrets and PKI with HashiCorp Vault. An agent skill from BagelHole/DevOps-Security-Agent-Skills.
BagelHole/DevOps-Security-Agent-Skills
Handle security incidents with IR playbooks and procedures. An agent skill from BagelHole/DevOps-Security-Agent-Skills.
BagelHole/DevOps-Security-Agent-Skills
Deploy, scale, and manage Kubernetes workloads. An agent skill from BagelHole/DevOps-Security-Agent-Skills.
BagelHole/DevOps-Security-Agent-Skills
Apply CIS benchmarks and secure Linux servers. An agent skill from BagelHole/DevOps-Security-Agent-Skills.
BagelHole/DevOps-Security-Agent-Skills
Set up metrics collection and visualization with Prometheus and Grafana.
BagelHole/DevOps-Security-Agent-Skills
Scan systems and dependencies for CVEs and security vulnerabilities.
Works with
Categories
Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster. AI Pipeline Orchestration is an agent skill from BagelHole/DevOps-Security-Agent-Skills. Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.
AI Pipeline Orchestration fits situations like: tasks that involve Data pipelines and ETL; tasks that involve Fine-tuning.
Run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a claude-code`. Or copy the skill folder (devops/ai/ai-pipeline-orchestration in BagelHole/DevOps-Security-Agent-Skills) into .claude/skills/ai-pipeline-orchestration in your project. Claude Code loads it when a task matches its description.
Run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a codex`. Or copy the skill folder (devops/ai/ai-pipeline-orchestration in BagelHole/DevOps-Security-Agent-Skills) into .agents/skills/ai-pipeline-orchestration in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill ai-pipeline-orchestration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-pipeline-orchestration, .gemini/skills/ai-pipeline-orchestration, .github/skills/ai-pipeline-orchestration and .opencode/skills/ai-pipeline-orchestration in your project.
Going by SKILL.md and its folder, AI Pipeline Orchestration needs the command-line tools its instructions call (prefect and pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
AI Pipeline Orchestration is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with AI Pipeline Orchestration: AI Pipeline Orchestration (sickn33/agentic-awesome-skills, 47k stars), AI Data Engineering (ancoleman/ai-design-components, 525 stars), Migrating Dagster To Airflow (astronomer/agents, 451 stars) and Engineering Data Pipelines (telagod/code-abyss, 244 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
BagelHole (a GitHub user) maintains it in BagelHole/DevOps-Security-Agent-Skills, which has 1,152 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on May 22, 2026.
Source: BagelHole/DevOps-Security-Agent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.