Agent skill

ML Pipeline Creation

by seb1n in seb1n/awesome-ai-agent-skills

Design, implement, and validate reproducible machine-learning pipelines spanning data preparation, training, evaluation, registry, and deployment gates.

MITAuto-check passedDevOps & Cloud

Install ML Pipeline Creation

skills CLI
$ npx skills add seb1n/awesome-ai-agent-skills --skill ml-pipeline-creation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seb1n/awesome-ai-agent-skills ml-pipeline-creation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ai-ml-operations/ml-pipeline-creation .claude/skills/ml-pipeline-creation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-pipeline-creation
GitHub stars
206
Token cost
~1.5k tokens
SKILL.md length
628 words
Files
1
Skills in repo
101
Repo updated
First seen
Licence
MIT

At a glance

Design, implement, and validate reproducible machine-learning pipelines spanning data preparation, training, evaluation, registry, and deployment gates.

  • Works in 7 steps: A dependency graph of pipeline stages… → A versioned pipeline definition or… → Explicit schemas for every stage input… → …
  • The user requests an ML pipeline
  • SKILL.md covers Required Inputs, Output Contract, Workflow and Example, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

ML Pipeline Creation is an agent skill from seb1n/awesome-ai-agent-skills. Design, implement, and validate reproducible machine-learning pipelines spanning data preparation, training, evaluation, registry, and deployment gates. Use when the user requests an ML pipeline, needs to turn model scripts into an orchestrated workflow, or provides pipeline components that must be connected safely.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering MLOps and Machine learning. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.

When your agent uses it

  • The user requests an ML pipeline
  • Needs to turn model scripts into an orchestrated workflow
  • Provides pipeline components that must be connected safely

Example prompts

  • “/ml-pipeline-creation”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. A dependency graph of pipeline stages and artifacts
  2. A versioned pipeline definition or implementation
  3. Explicit schemas for every stage input and output
  4. Data, model, and environment versioning rules
  5. Evaluation and promotion gates with failure behavior
  6. Observability, retry, backfill, and rollback procedures
  7. A verification record showing how the pipeline was tested

What it can do on your machine

Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Pipeline Creation loads about 1.5k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 628 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 628 words, ~1,506 tokens.

Download SKILL.mdSave it as .claude/skills/ml-pipeline-creation/SKILL.md (or your agent's skills folder).
name
ml-pipeline-creation
description
Design, implement, and validate reproducible machine-learning pipelines spanning data preparation, training, evaluation, registry, and deployment gates. Use when the user requests an ML pipeline, needs to turn model scripts into an orchestrated workflow, or provides pipeline components that must be connected safely.
license
MIT
metadata.author
awesome-ai-agent-skills contributors
metadata.version
1.0.0

ML Pipeline Creation

Build reproducible ML workflows whose inputs, outputs, lineage, and promotion criteria are explicit. Prefer the project's existing orchestrator and conventions; do not introduce a platform merely to demonstrate one.

Required Inputs

  • Business objective and measurable model acceptance criteria
  • Data sources, ownership, sensitivity, and expected refresh cadence
  • Existing preprocessing, training, evaluation, and serving code
  • Target environments and available orchestration or CI system
  • Compute, cost, latency, reproducibility, and compliance constraints

If critical details are missing, state assumptions and design a platform-neutral pipeline before selecting an implementation.

Output Contract

Produce:

  1. A dependency graph of pipeline stages and artifacts
  2. A versioned pipeline definition or implementation
  3. Explicit schemas for every stage input and output
  4. Data, model, and environment versioning rules
  5. Evaluation and promotion gates with failure behavior
  6. Observability, retry, backfill, and rollback procedures
  7. A verification record showing how the pipeline was tested

Workflow

  1. Inspect the environment. Identify the repository language, dependency manager, existing orchestration system, model framework, artifact store, and deployment path. Reuse established tools where possible.
  2. Define the contract. Record the objective, data snapshot rules, target metric, baseline, acceptance threshold, resource budget, and deployment constraints. Separate offline evaluation from production health metrics.
  3. Model the DAG. Represent ingestion, validation, splitting, transformation, training, evaluation, registration, and deployment as idempotent stages. Declare every artifact rather than relying on undeclared files or mutable global state.
  4. Implement reproducibility. Pin dependencies, seed stochastic operations where appropriate, version code and data, capture parameters, and store immutable artifacts with provenance. Prevent train/validation leakage by fitting transformations only on training data.
  5. Add quality gates. Validate schemas before training, compare metrics with a baseline, fail closed on missing or invalid artifacts, and require explicit approval before production promotion when consequences are material.
  6. Design operations. Define retries only for transient failures, make reruns idempotent, specify backfill boundaries, emit structured logs and metrics, and document rollback to the last known-good model.
  7. Test incrementally. Run unit tests for components, a small deterministic end-to-end fixture, and a staging or dry-run execution. Confirm that a failed stage cannot silently publish a model.

Example

For a batch classifier, define the artifact flow explicitly:

yaml
pipeline: customer-churn-training
inputs:
  raw_snapshot: data/raw/churn-2026-08-01.parquet
stages:
  - name: prepare-data
    inputs: [raw_snapshot]
    outputs: [train_set, validation_set, test_set, feature_schema]
  - name: train-model
    inputs: [train_set, feature_schema, training_config]
    outputs: [model, training_metrics]
  - name: evaluate-model
    inputs: [model, validation_set, test_set, baseline_metrics]
    outputs: [evaluation_report, promotion_decision]
  - name: register-model
    condition: promotion_decision == "pass"
    inputs: [model, evaluation_report]
    outputs: [registered_model_version]

Require prepare-data to emit every declared split. Reject the run if test_set is absent rather than letting evaluation consume an undeclared path.

Show full SKILL.md (246 more words)Show less

Safety Boundaries

  • Treat datasets, credentials, model artifacts, and logs as potentially sensitive.
  • Never copy production data into development without authorization and required de-identification.
  • Do not deploy, replace a registered model, or alter production infrastructure without explicit approval.
  • Present destructive migration or cleanup plans before execution and preserve a rollback path.
  • Flag fairness, privacy, security, or regulatory review requirements instead of claiming compliance from pipeline execution alone.

Verification

  • Re-run the fixture twice and confirm identical stage contracts and expected deterministic outputs.
  • Force one stage to fail and verify downstream stages do not execute.
  • Verify artifact hashes, code revision, parameters, data version, and evaluation results are traceable from the registered model.
  • Test retry and backfill behavior without duplicating records or overwriting immutable artifacts.
  • Confirm the deployment gate rejects a model below threshold and accepts a known-good fixture.

Best Practices

  • Keep components small, idempotent, and independently testable.
  • Separate pipeline orchestration from model business logic.
  • Store configuration as versioned data; do not bury thresholds in code.
  • Prefer immutable artifacts and explicit lineage over mutable “latest” paths.
  • Monitor data quality and model behavior after deployment, not only during training.

Edge Cases

  • Streaming data: Use event-time semantics, checkpointing, and replay-safe sinks.
  • Non-deterministic training: Record seeds and environment details, then validate within an agreed tolerance.
  • Large backfills: Bound the date range, estimate cost, and test one partition before scaling.
  • Schema drift: Quarantine incompatible data and require a reviewed schema migration.
  • Partial promotion: Keep registry, serving configuration, and monitoring changes transactional or explicitly reversible.

© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in ai-ml-operations/ml-pipeline-creation of seb1n/awesome-ai-agent-skills.

Open the folder on GitHubat commit 75865a5

Compare with similar skills

ML Pipeline Creation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Pipeline Creation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Pipeline Creation this skillseb1n/awesome-ai-agent-skills206—~1.5kAutomated safety check: PassMIT
Build ML Pipelineprobabl-ai/skills138—~4.4kAutomated safety check: PassBSD-3-Clause
ML Pipeline Workflowwshobson/agents40k12 repos~1.8kAutomated safety check: PassMIT
Editomegaml/omegaml108—~206Automated safety check: PassApache-2.0
ML Pipeline ExpertJeffallan/claude-skills12k—~1.9kAutomated safety check: PassMIT
Machine Learning Ops ML Pipelineaiskillstore/marketplace4337 repos~2.6kAutomated safety check: PassNone

Similar skills

  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    138 GitHub stars~4.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • ML Pipeline Workflow

    wshobson/agents

    Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.

    40k GitHub starsUsed in 12 repos~1.8k tokens
    DevOps & CloudAuto-check passed
  • Edit

    omegaml/omegaml

    how to use the edit command properly

    108 GitHub stars~206 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub stars~1.9k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed
  • Machine Learning Ops ML Pipeline

    aiskillstore/marketplace

    Design and implement a complete ML pipeline for: $ARGUMENTS. An agent skill from aiskillstore/marketplace.

    433 GitHub starsUsed in 7 repos~2.6k tokens
    DevOps & CloudAuto-check passed
  • Model Deployment

    secondsky/claude-skills

    Deploy ML models with FastAPI, Docker, Kubernetes. An agent skill from secondsky/claude-skills.

    227 GitHub stars~2.4k tokensUpdated 12 days ago
    DevOps & CloudAuto-check passed

More from seb1n/awesome-ai-agent-skills

All 101 skills in this repo
  • Agent Red Teaming

    seb1n/awesome-ai-agent-skills

    Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.

    206 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Eu AI Act Readiness

    seb1n/awesome-ai-agent-skills

    Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…

    206 GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Human In The Loop

    seb1n/awesome-ai-agent-skills

    Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • MCP Server Building

    seb1n/awesome-ai-agent-skills

    Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • PDF Processing

    seb1n/awesome-ai-agent-skills

    Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Skill Supply Chain Audit

    seb1n/awesome-ai-agent-skills

    Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.

    206 GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about ML Pipeline Creation

What does ML Pipeline Creation do?

Design, implement, and validate reproducible machine-learning pipelines spanning data preparation, training, evaluation, registry, and deployment gates. ML Pipeline Creation is an agent skill from seb1n/awesome-ai-agent-skills. Design, implement, and validate reproducible machine-learning pipelines spanning data preparation, training, evaluation, registry, and deployment gates.

When should I use ML Pipeline Creation?

ML Pipeline Creation fits situations like: the user requests an ML pipeline; needs to turn model scripts into an orchestrated workflow; provides pipeline components that must be connected safely.

How do I install ML Pipeline Creation in Claude Code?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill ml-pipeline-creation -a claude-code`. Or copy the skill folder (ai-ml-operations/ml-pipeline-creation in seb1n/awesome-ai-agent-skills) into .claude/skills/ml-pipeline-creation in your project. Claude Code loads it when a task matches its description.

How do I install ML Pipeline Creation in Codex?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill ml-pipeline-creation -a codex`. Or copy the skill folder (ai-ml-operations/ml-pipeline-creation in seb1n/awesome-ai-agent-skills) into .agents/skills/ml-pipeline-creation in your project. Codex loads it when a task matches its description.

Can I use ML Pipeline Creation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill ml-pipeline-creation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-pipeline-creation, .gemini/skills/ml-pipeline-creation, .github/skills/ml-pipeline-creation and .opencode/skills/ml-pipeline-creation in your project.

What does ML Pipeline Creation need to run?

SKILL.md names no scripts, command-line tools or credentials: ML Pipeline Creation is instructions for the agent only.

Does ML Pipeline Creation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML Pipeline Creation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Pipeline Creation use?

ML Pipeline Creation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Pipeline Creation use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Pipeline Creation?

Skills that share tags, products or a category with ML Pipeline Creation: Build ML Pipeline (probabl-ai/skills, 138 stars), ML Pipeline Workflow (wshobson/agents, 40k stars), Edit (omegaml/omegaml, 108 stars) and ML Pipeline Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Pipeline Creation?

seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 101 skills in this directory. The repository was last updated on August 9, 2026.

Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.