ML Pipeline Workflow
wshobson/agents
Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.
Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.
$ npx skills add Jeffallan/claude-skills --skill ml-pipeline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Jeffallan/claude-skills ml-pipeline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-pipeline .claude/skills/ml-pipeline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ml-pipeline" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/ml-pipeline into .claude/skills/ml-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-pipeline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Jeffallan/claude-skills/tree/main/skills/ml-pipelineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Jeffallan/claude-skills --skill ml-pipeline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Jeffallan/claude-skills ml-pipeline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ml-pipeline .agents/skills/ml-pipeline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ml-pipeline" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/ml-pipeline into .agents/skills/ml-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-pipeline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Jeffallan/claude-skills --skill ml-pipeline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Jeffallan/claude-skills ml-pipeline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ml-pipeline .cursor/skills/ml-pipeline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ml-pipeline" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/ml-pipeline into .cursor/skills/ml-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-pipeline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Jeffallan/claude-skills.git --path skills/ml-pipeline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Jeffallan/claude-skills --skill ml-pipeline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Jeffallan/claude-skills ml-pipeline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ml-pipeline .gemini/skills/ml-pipeline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ml-pipeline" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/ml-pipeline into .gemini/skills/ml-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-pipeline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Jeffallan/claude-skills ml-pipelineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Jeffallan/claude-skills --skill ml-pipeline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ml-pipeline .github/skills/ml-pipeline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ml-pipeline" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/ml-pipeline into .github/skills/ml-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-pipeline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Jeffallan/claude-skills --skill ml-pipeline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Jeffallan/claude-skills ml-pipeline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ml-pipeline .opencode/skills/ml-pipeline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ml-pipeline" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/ml-pipeline into .opencode/skills/ml-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-pipeline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ml-pipelineDesigns ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.
The agent maps data flow and stages, runs schema and distribution checks that halt and report on failure before any training, builds feature transformations and feature stores, configures distributed training and hyperparameter tuning, logs metrics, parameters and artifacts so runs can be compared, and ends with evaluation gates plus A/B testing or shadow deployment before a model is promoted.
Reference files cover feature engineering with Feast and data validation, training pipelines, experiment tracking with MLflow and Weights & Biases plus a model registry, orchestration with Kubeflow Pipelines, Airflow and Prefect, and model validation. Templates include MLflow logging, a Kubeflow pipeline component and a Great Expectations style data check. The rules ask for versioning data, code and models with DVC, Git tags and the registry, and for pinned dependencies and random seeds.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comsynergetic.solutionsjeffallan.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ML Pipeline Expert loads about 1.9k tokens when it runs, and up to ~31k if it reads all its reference files. Until then it costs about 154 tokens; SKILL.md has 394 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 394 words, ~1,852 tokens.
.claude/skills/ml-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Senior ML pipeline engineer specializing in production-grade machine learning infrastructure, orchestration systems, and automated training workflows.
Load detailed guidance based on context:
| Topic | Reference | Load When |
|---|---|---|
| Feature Engineering | references/feature-engineering.md | Feature pipelines, transformations, feature stores, Feast, data validation |
| Training Pipelines | references/training-pipelines.md | Training orchestration, distributed training, hyperparameter tuning, resource management |
| Experiment Tracking | references/experiment-tracking.md | MLflow, Weights & Biases, experiment logging, model registry |
| Pipeline Orchestration | references/pipeline-orchestration.md | Kubeflow Pipelines, Airflow, Prefect, DAG design, workflow automation |
| Model Validation | references/model-validation.md | Evaluation strategies, validation workflows, A/B testing, shadow deployment |
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, f1_score
import numpy as np
# Pin random state for reproducibility
SEED = 42
np.random.seed(SEED)
mlflow.set_experiment("my-classifier-experiment")
with mlflow.start_run():
# Log all hyperparameters — never hardcode silently
params = {"n_estimators": 100, "max_depth": 5, "random_state": SEED}
mlflow.log_params(params)
model = RandomForestClassifier(**params)
model.fit(X_train, y_train)
preds = model.predict(X_test)
# Log metrics
mlflow.log_metric("accuracy", accuracy_score(y_test, preds))
mlflow.log_metric("f1", f1_score(y_test, preds, average="weighted"))
# Log and register the model artifact
mlflow.sklearn.log_model(model, artifact_path="model",
registered_model_name="my-classifier")from kfp.v2 import dsl
from kfp.v2.dsl import component, Input, Output, Dataset, Model, Metrics
@component(base_image="python:3.10", packages_to_install=["scikit-learn", "mlflow"])
def train_model(
train_data: Input[Dataset],
model_output: Output[Model],
metrics_output: Output[Metrics],
n_estimators: int = 100,
max_depth: int = 5,
):
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
import pickle, json
df = pd.read_csv(train_data.path)
X, y = df.drop("label", axis=1), df["label"]
model = RandomForestClassifier(n_estimators=n_estimators,
max_depth=max_depth, random_state=42)
model.fit(X, y)
with open(model_output.path, "wb") as f:
pickle.dump(model, f)
metrics_output.log_metric("train_samples", len(df))
@dsl.pipeline(name="training-pipeline")
def training_pipeline(data_path: str, n_estimators: int = 100):
train_step = train_model(n_estimators=n_estimators)
# Chain additional steps (validate, register, deploy) hereimport great_expectations as ge
def validate_training_data(df):
"""Run schema and distribution checks. Raise on failure — never skip."""
gdf = ge.from_pandas(df)
results = gdf.expect_column_values_to_not_be_null("label")
results &= gdf.expect_column_values_to_be_between("feature_1", 0, 1)
if not results["success"]:
raise ValueError(f"Data validation failed: {results['result']}")
return df # safe to proceed to trainingAlways:
Never:
When implementing a pipeline, provide:
MLflow, Kubeflow Pipelines, Apache Airflow, Prefect, Feast, Weights & Biases, Neptune, DVC, Great Expectations, Ray, Horovod, Kubernetes, Docker, S3/GCS/Azure Blob, model registry patterns, feature store architecture, distributed training, hyperparameter optimization
Maintained by @jeffallan, Principal Consultant at Synergetic Solutions
© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/ml-pipeline of Jeffallan/claude-skills.
Open the folder on GitHubat commit 1be15d8
ML Pipeline Expert next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ML Pipeline Expert this skillJeffallan/claude-skills | 12k | — | ~1.9k | Automated safety check: Pass | MIT | |
| ML Pipeline Workflowwshobson/agents | 40k | 12 repos | ~1.8k | Automated safety check: Pass | MIT | |
| ML Pipeline Automationsecondsky/claude-skills | 227 | — | ~3.2k | Automated safety check: Pass | MIT | |
| Implementing Mlopsancoleman/ai-design-components | 525 | — | ~9.2k | Automated safety check: Pass | MIT | |
| AI Data Engineeringancoleman/ai-design-components | 525 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Experiment Tracking Setuprevfactory/harness-100 | 1.3k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 |
wshobson/agents
Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.
secondsky/claude-skills
Automate ML workflows with Airflow, Kubeflow, MLflow. An agent skill from secondsky/claude-skills.
ancoleman/ai-design-components
Strategic guidance for operationalizing machine learning models from experimentation to production.
ancoleman/ai-design-components
Data pipelines, feature stores, and embedding generation for AI/ML systems.
revfactory/harness-100
Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology.
Orchestra-Research/AI-Research-SKILLs
Tracks ML experiments, versions models in the MLflow registry and covers deployment and reproducibility, with autologging for common frameworks.
Jeffallan/claude-skills
Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.
Jeffallan/claude-skills
Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.
Jeffallan/claude-skills
Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.
Jeffallan/claude-skills
Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.
Jeffallan/claude-skills
Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
Works with
Categories
Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates. The agent maps data flow and stages, runs schema and distribution checks that halt and report on failure before any training, builds feature transformations and feature stores, configures distributed training and hyperparameter tuning, logs metrics, parameters and artifacts so runs can be compared, and ends with evaluation gates plus A/B testing or shadow deployment before a model is promoted.
ML Pipeline Expert fits situations like: building a training pipeline with orchestrated stages; setting up experiment tracking and a model registry; defining a feature store schema; adding data validation and evaluation gates before deployment.
Run `npx skills add Jeffallan/claude-skills --skill ml-pipeline -a claude-code`. Or copy the skill folder (skills/ml-pipeline in Jeffallan/claude-skills) into .claude/skills/ml-pipeline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Jeffallan/claude-skills --skill ml-pipeline -a codex`. Or copy the skill folder (skills/ml-pipeline in Jeffallan/claude-skills) into .agents/skills/ml-pipeline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill ml-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-pipeline, .gemini/skills/ml-pipeline, .github/skills/ml-pipeline and .opencode/skills/ml-pipeline in your project.
SKILL.md names no scripts, command-line tools or credentials: ML Pipeline Expert is instructions for the agent only.
SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
ML Pipeline Expert is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 29k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with ML Pipeline Expert: ML Pipeline Workflow (wshobson/agents, 40k stars), ML Pipeline Automation (secondsky/claude-skills, 227 stars), Implementing Mlops (ancoleman/ai-design-components, 525 stars) and AI Data Engineering (ancoleman/ai-design-components, 525 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,802 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.
Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.