Agent skill

ML Pipeline Guide

by wentorai in wentorai/research-plugins

Build and deploy reproducible production ML pipelines for research

MITAuto-check passedDevOps & Cloud

Install ML Pipeline Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill ml-pipeline-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins ml-pipeline-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/ai-ml/ml-pipeline-guide .claude/skills/ml-pipeline-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-pipeline-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
295 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Build and deploy reproducible production ML pipelines for research

  • Tasks that involve MLOps
  • SKILL.md covers Overview, Pipeline Architecture, Experiment Tracking with MLflow and Data Versioning with DVC, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

ML Pipeline Guide is an agent skill from wentorai/research-plugins. Build and deploy reproducible production ML pipelines for research

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering MLOps. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve MLOps

Example prompts

  • “/ml-pipeline-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, yaml, bash and makefile).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • mlflow.org
    • dvc.org
    • hydra.cc
    • drivendata.github.io
    • madewithml.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Pipeline Guide loads about 2.3k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 295 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 295 words, ~2,330 tokens.

Download SKILL.mdSave it as .claude/skills/ml-pipeline-guide/SKILL.md (or your agent's skills folder).
name
ml-pipeline-guide
description
Build and deploy reproducible production ML pipelines for research

ML Pipeline Guide

Overview

Machine learning research increasingly demands reproducible, end-to-end pipelines that go beyond a single training script. A research ML pipeline encompasses data ingestion, feature engineering, model training, evaluation, experiment tracking, and artifact management. Without a structured pipeline, research results become difficult to reproduce, ablation studies become error-prone, and collaborators cannot build on prior work.

This guide covers the practical tools and patterns for building ML pipelines in an academic research context. The focus is on reproducibility, experiment tracking, and the transition from notebook prototyping to structured experiments. The patterns use MLflow, DVC, and standard Python tooling -- chosen because they are open source, widely adopted in published research, and require minimal infrastructure.

Unlike industry MLOps guides that emphasize deployment at scale, this guide prioritizes the research workflow: running many experiments, tracking what changed between runs, and producing results that reviewers can verify.

Pipeline Architecture

A research ML pipeline typically has five stages:

Data Ingestion → Feature Engineering → Training → Evaluation → Artifact Storage
     │                  │                 │            │              │
     ├── raw data       ├── transforms    ├── model    ├── metrics    ├── models
     ├── splits         ├── features      ├── logs     ├── plots      ├── configs
     └── metadata       └── cache         └── ckpts    └── tables     └── reports
Directory Structure for Reproducible Research
project/
├── configs/
│   ├── base.yaml           # Default hyperparameters
│   ├── experiment_001.yaml  # Experiment-specific overrides
│   └── sweep.yaml          # Hyperparameter search space
├── data/
│   ├── raw/                # Immutable original data
│   ├── processed/          # Cleaned and transformed
│   └── splits/             # Train/val/test splits (versioned)
├── src/
│   ├── data/               # Data loading and preprocessing
│   ├── features/           # Feature engineering
│   ├── models/             # Model definitions
│   ├── training/           # Training loops
│   └── evaluation/         # Metrics and visualization
├── experiments/            # MLflow/W&B experiment logs
├── notebooks/              # Exploratory analysis only
├── tests/                  # Unit tests for pipeline components
├── Makefile                # Reproducible commands
├── requirements.txt        # Pinned dependencies
└── dvc.yaml                # Data version control pipeline

Experiment Tracking with MLflow

python
import mlflow
import mlflow.pytorch
from pathlib import Path

def run_experiment(config: dict):
    """Run a single experiment with full tracking."""
    mlflow.set_experiment(config["experiment_name"])

    with mlflow.start_run(run_name=config.get("run_name")):
        # Log configuration
        mlflow.log_params({
            "model": config["model_name"],
            "learning_rate": config["lr"],
            "batch_size": config["batch_size"],
            "epochs": config["epochs"],
            "optimizer": config["optimizer"],
            "seed": config["seed"],
        })

        # Log environment
        mlflow.log_param("python_version", sys.version)
        mlflow.log_param("torch_version", torch.__version__)
        mlflow.log_param("cuda_version", torch.version.cuda)

        # Training
        model = build_model(config)
        for epoch in range(config["epochs"]):
            train_loss = train_one_epoch(model, train_loader, optimizer)
            val_loss, val_metrics = evaluate(model, val_loader)

            mlflow.log_metrics({
                "train_loss": train_loss,
                "val_loss": val_loss,
                **{f"val_{k}": v for k, v in val_metrics.items()},
            }, step=epoch)

        # Log final model
        mlflow.pytorch.log_model(model, "model")

        # Log artifacts (plots, configs)
        mlflow.log_artifact(config_path)
        save_evaluation_plots(model, test_loader, "plots/")
        mlflow.log_artifacts("plots/")

        return val_metrics

Data Versioning with DVC

yaml
# dvc.yaml -- Pipeline definition
stages:
  prepare_data:
    cmd: python src/data/prepare.py --config configs/base.yaml
    deps:
      - src/data/prepare.py
      - data/raw/
    outs:
      - data/processed/
    params:
      - configs/base.yaml:
          - data.split_ratio
          - data.random_seed

  extract_features:
    cmd: python src/features/extract.py --config configs/base.yaml
    deps:
      - src/features/extract.py
      - data/processed/
    outs:
      - data/features/
    params:
      - configs/base.yaml:
          - features

  train:
    cmd: python src/training/train.py --config configs/base.yaml
    deps:
      - src/training/train.py
      - src/models/
      - data/features/
    outs:
      - models/
    metrics:
      - metrics.json:
          cache: false
    plots:
      - plots/training_curve.csv:
          x: epoch
          y: loss
bash
# Reproduce the full pipeline
dvc repro

# Compare experiments
dvc metrics diff

# Push data to remote storage
dvc push

Configuration Management with Hydra

python
import hydra
from omegaconf import DictConfig, OmegaConf

@hydra.main(config_path="configs", config_name="base", version_base=None)
def main(cfg: DictConfig):
    print(OmegaConf.to_yaml(cfg))

    model = build_model(
        name=cfg.model.name,
        hidden_dim=cfg.model.hidden_dim,
        num_layers=cfg.model.num_layers,
    )

    train(
        model=model,
        lr=cfg.training.lr,
        epochs=cfg.training.epochs,
        batch_size=cfg.training.batch_size,
    )

# Override from command line:
# python train.py training.lr=1e-4 model.hidden_dim=512
# python train.py --multirun training.lr=1e-3,1e-4,1e-5
yaml
# configs/base.yaml
model:
  name: resnet50
  hidden_dim: 256
  num_layers: 4

training:
  lr: 1e-3
  epochs: 100
  batch_size: 32
  optimizer: adamw
  weight_decay: 0.01

data:
  dataset: cifar10
  split_ratio: [0.8, 0.1, 0.1]
  random_seed: 42
  augmentation: true

Feature Engineering Patterns

python
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
import joblib

def build_feature_pipeline(numeric_cols: list, categorical_cols: list) -> Pipeline:
    """Build a reproducible feature engineering pipeline."""
    numeric_transformer = Pipeline([
        ("imputer", SimpleImputer(strategy="median")),
        ("scaler", StandardScaler()),
    ])

    categorical_transformer = Pipeline([
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)),
    ])

    preprocessor = ColumnTransformer([
        ("num", numeric_transformer, numeric_cols),
        ("cat", categorical_transformer, categorical_cols),
    ])

    return preprocessor

# Save and load for reproducibility
preprocessor.fit(X_train)
joblib.dump(preprocessor, "artifacts/preprocessor.pkl")
# Later: preprocessor = joblib.load("artifacts/preprocessor.pkl")

Makefile for Reproducibility

makefile
.PHONY: setup data train evaluate all clean

setup:
	pip install -r requirements.txt
	dvc pull

data:
	python src/data/prepare.py --config configs/base.yaml

train:
	python src/training/train.py --config configs/base.yaml

evaluate:
	python src/evaluation/evaluate.py --config configs/base.yaml

all: setup data train evaluate

sweep:
	python src/training/train.py --multirun \
		training.lr=1e-3,1e-4,1e-5 \
		model.hidden_dim=128,256,512

clean:
	rm -rf outputs/ multirun/ __pycache__/

Best Practices

  • Never modify raw data. All transformations should be scripted and reproducible.
  • Pin every dependency version including CUDA, cuDNN, and OS-level libraries.
  • Separate configuration from code. Use YAML/JSON configs, not hardcoded values.
  • Track experiments from day one. Retrofitting experiment tracking is painful.
  • Write tests for data preprocessing. Shape mismatches and silent data corruption are common.
  • Use Makefile or dvc repro so any collaborator can reproduce results with one command.
  • Version your data alongside your code using DVC, Git-LFS, or cloud storage with manifests.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/ai-ml/ml-pipeline-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

ML Pipeline Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Pipeline Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Pipeline Guide this skillwentorai/research-plugins2981 repos~2.3kAutomated safety check: PassMIT
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0
SkyPilot Multi-Cloud OrchestrationOrchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT
Model Garden Deploymentgoogle/skills21k—~5kAutomated safety check: PassApache-2.0
Register ModelSunshow/droidgear127—~2.5kAutomated safety check: PassMIT
Build ML Pipelineprobabl-ai/skills138—~4.4kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • SkyPilot Multi-Cloud Orchestration

    Orchestra-Research/AI-Research-SKILLs

    Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

    13k GitHub starsUsed in 4 repos~2.4k tokens
    DevOps & CloudAuto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Register Model

    Sunshow/droidgear

    Register a new AI model in DroidGear's model registry by fetching specs from models.dev.

    127 GitHub stars~2.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    138 GitHub stars~4.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • ML Pipeline Workflow

    wshobson/agents

    Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.

    40k GitHub starsUsed in 12 repos~1.8k tokens
    DevOps & CloudAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Categories

Questions about ML Pipeline Guide

What does ML Pipeline Guide do?

Build and deploy reproducible production ML pipelines for research. ML Pipeline Guide is an agent skill from wentorai/research-plugins.

When should I use ML Pipeline Guide?

ML Pipeline Guide fits situations like: tasks that involve MLOps.

How do I install ML Pipeline Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill ml-pipeline-guide -a claude-code`. Or copy the skill folder (skills/domains/ai-ml/ml-pipeline-guide in wentorai/research-plugins) into .claude/skills/ml-pipeline-guide in your project. Claude Code loads it when a task matches its description.

How do I install ML Pipeline Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill ml-pipeline-guide -a codex`. Or copy the skill folder (skills/domains/ai-ml/ml-pipeline-guide in wentorai/research-plugins) into .agents/skills/ml-pipeline-guide in your project. Codex loads it when a task matches its description.

Can I use ML Pipeline Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill ml-pipeline-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-pipeline-guide, .gemini/skills/ml-pipeline-guide, .github/skills/ml-pipeline-guide and .opencode/skills/ml-pipeline-guide in your project.

What does ML Pipeline Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: ML Pipeline Guide is instructions for the agent only. Our summary lists: Python 3.

Does ML Pipeline Guide access the network?

SKILL.md names 5 domains. As links in the text: mlflow.org, dvc.org, hydra.cc, drivendata.github.io and madewithml.com. This is read from the text; nothing was executed.

Is ML Pipeline Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Pipeline Guide use?

ML Pipeline Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Pipeline Guide use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Pipeline Guide?

Skills that share tags, products or a category with ML Pipeline Guide: SageMaker Production Defaults (huggingface/skills, 11k stars), SkyPilot Multi-Cloud Orchestration (Orchestra-Research/AI-Research-SKILLs, 13k stars), Model Garden Deployment (google/skills, 21k stars) and Register Model (Sunshow/droidgear, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Pipeline Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.