MANDATORY whenever a task involves training, fine-tuning, tuning, or evaluating a machine-learning model on data (tabular, time series, text, images — any modality).

Apache-2.0Auto-check passedData & Analytics

Install ML Pipeline

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill ml-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins ml-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/jananthan30/ml-pipeline/skills/ml-pipeline .claude/skills/ml-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-pipeline
GitHub stars
1.3k
Token cost
~3.2k tokens
SKILL.md length
1,588 words
Files
2 (incl. references)
Skills in repo
714
Repo updated
First seen
Licence
Apache-2.0

At a glance

MANDATORY whenever a task involves training, fine-tuning, tuning, or evaluating a machine-learning model on data (tabular, time series, text, images — any modality).

  • Works in 5 steps: What was done — the steps completed, 1–2… → What was found — key findings, with the… → Decisions made and why — e.g. "dropped… → …
  • Tasks that involve MLOps
  • SKILL.md covers The pipeline (strict order —…, Phase gates — explicit…, Progress tracking (survives… and Tools: marimo notebooks +…, plus 3 more sections
  • Calls git

What it does

ML Pipeline is an agent skill from hashgraph-online/awesome-codex-plugins. MANDATORY whenever a task involves training, fine-tuning, tuning, or evaluating a machine-learning model on data (tabular, time series, text, images — any modality). Enforces a strict 16-step pipeline that starts with inspecting the raw data, gates each phase behind the user's explicit permission, and produces marimo notebooks with matplotlib visuals so the user can see and understand every step. Never jump straight to model training.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/model-selection.md`).

It sits in Data & Analytics, covering MLOps, Fine-tuning and Data visualization. It works with marimo, Matplotlib and Git. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve MLOps
  • Tasks that involve Fine-tuning
  • Tasks that involve Data visualization

Example prompts

  • “/ml-pipeline”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. What was done — the steps completed, 1–2 sentences each.
  2. What was found — key findings, with the visuals that show them.
  3. Decisions made and why — e.g. "dropped 312 duplicate rows", "chose time-based split
  4. What comes next — the next phase's steps, in one short list. At Gate B this section
  5. The question — ask for explicit permission to continue. Wait for a clear yes.

What it can do on your machine

Read from SKILL.md and the folder at commit 9e7b281. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Pipeline loads about 3.2k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 1,588 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 9e7b281, republished under its Apache-2.0 licence (© hashgraph-online). 1,588 words, ~3,233 tokens.

Download SKILL.mdSave it as .claude/skills/ml-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ml-pipeline
description
MANDATORY whenever a task involves training, fine-tuning, tuning, or evaluating a machine-learning model on data (tabular, time series, text, images — any modality). Enforces a strict 16-step pipeline that starts with inspecting the raw data, gates each phase behind the user's explicit permission, and produces marimo notebooks with matplotlib visuals so the user can see and understand every step. Never jump straight to model training.

ML Pipeline Discipline

Hard rule: no model is trained until every earlier step in the pipeline is done and the user has explicitly approved the phase gates before it. "Train a model on this data" is a request to start the pipeline at step 1, not at step 10.

The pipeline (strict order — never reorder, never skip silently)

RAW DATA
  → 1. Data inspection
  → 2. Exploratory data analysis (EDA)
  → 3. Define the prediction problem
  ──────────── GATE A: user approval ────────────
  → 4. Data cleaning
  → 5. Data engineering
  → 6. Train / validation / test split
  → 7. Feature engineering
  → 8. Preprocessing
  ──────────── GATE B: user approval ────────────
  → 9. Baseline model
  → 10. Model training
  → 11. Hyperparameter tuning
  → 12. Model evaluation
  ──────────── GATE C: user approval ────────────
  → 13. Error analysis
  → 14. Final test (test set touched ONCE)
  → 15. Deployment (only if user asks)
  → 16. Monitoring + retraining plan
  ──────────── GATE D: wrap-up report ───────────

Phase gates — explicit permission, every time

At each gate, STOP and give the user, in plain non-jargon language:

  1. What was done — the steps completed, 1–2 sentences each.
  2. What was found — key findings, with the visuals that show them.
  3. Decisions made and why — e.g. "dropped 312 duplicate rows", "chose time-based split because the data has dates".
  4. What comes next — the next phase's steps, in one short list. At Gate B this section includes the Model rationale: block (format under Progress tracking).
  5. The question — ask for explicit permission to continue. Wait for a clear yes. Silence, ambiguity, or "hmm" is not a yes. If the user redirects, incorporate it.

If the user says "skip ahead" or "just train it": explain in 2–3 sentences which steps are missing and the concrete risk (usually leakage or garbage-in), then ask once for explicit override confirmation. If they confirm, proceed and record the override in PIPELINE.md.

Progress tracking (survives across sessions)

On first use in a project, create ml_pipeline/PIPELINE.md — a checklist of the 16 steps with status (todo / in progress / done / approved-gate / overridden), one line of results per finished step, and dated gate approvals. Update it after every step. On any new session, read it first and resume from the first unfinished step — never restart, never skip ahead of it.

Record each gate approval on its own line in exactly this shape — the enforcement hook parses it:

- Gate A: approved 2026-09-16
- Gate B: approved 2026-09-17
- Override: Gate B - user approved skipping to training 2026-09-17 - reason: <why>

A line that says a gate is todo, pending, or not yet approved never counts as approval.

After recording a gate approval in a git repository, checkpoint it: git add -A && git commit -m "ml-pipeline: Gate X approved" && git tag -f gate-X. Every approved phase is then a reproducible point to return to.

Before recording Gate B: approved, PIPELINE.md must contain a model rationale in exactly this shape (the hook checks the five bullets exist; the user judges their content):

Model rationale:
- traits: binary, 1,428 rows (small), minority 11.4%, has_datetime, no groups
- baseline: majority-class dummy + logistic regression (class_weight=balanced)
- candidates: gradient boosting with small trees; regularized logistic
- ruled out: neural nets (small data); k-NN (mixed scales, weak on tabular); random split (temporal)
- metric: PR-AUC primary, recall at fixed precision secondary

Tools: marimo notebooks + matplotlib visuals

  • The workbench is a marimo notebook, not loose scripts. Keep notebooks in ml_pipeline/: 01_eda.py (steps 1–3), 02_prep.py (steps 4–8), 03_model.py (steps 9–12), 04_eval.py (steps 13–16). In Claude Code, drive them live with the marimo-pair skill so the user watches the work happen. In other harnesses, write the notebook files and tell the user to open them with marimo edit <file>.
  • Every step that looks at data produces matplotlib figures (seaborn on top is fine). Also save each figure to ml_pipeline/figures/<step>_<name>.png so gates can reference them even without a live notebook.
  • Explain every figure in 1–2 plain sentences: what it shows and why it matters for the next decision. A figure without an explanation is not done.

What each step must produce

  1. Data inspection — copy the plugin's guard library to ml_pipeline/guard.py (its path is given at session start; in Codex/Kimi it is skills/ml-pipeline/lib/mlpipeline_guard.py next to this file), load the raw data read-only, and run guard.profile(df, target=..., time_col=..., group_col=...). Read ml_pipeline/data_profile.json before anything else: its leakage_suspects and traits decide the split, the metric, and the model family. Report rows × columns, column kinds, missing values, duplicates, and every leakage suspect. For images or text, build a manifest table first (path, label, size or length) and profile that. No modification yet.
  2. EDA — guard.eda_figures(df, target=..., time_col=...) writes the required figures (01_missingness, 02_target_balance, 02_distributions, 02_correlations, 02_temporal_coverage when a date column exists) with explanations computed from the data; add any others with guard.fig(step, name, figure, explanation). If matplotlib is missing, install it — text descriptions are not figures and the gate will not accept them. Output: the figures, your interpretation of each, and a short list of hypotheses and problems spotted.
  3. Define the prediction problem — write a short contract: target (exact definition, units), prediction unit and population, prediction time/horizon, information actually available at prediction time, objective, evaluation metric, constraints. The user must approve this contract at Gate A — it controls everything after. Consult references/model-selection.md with the profile's traits when choosing the metric and the candidate model families.
  4. Data cleaning — missing values, duplicates, invalid/impossible values, inconsistent categories, unit/format issues. Report before/after counts for every rule. Document every rule in PIPELINE.md. Prefer preserving data over deleting. Never use information from the future or from the test rows to decide a cleaning rule.
  5. Data engineering — joins/integration, aggregation, time alignment to an index date, business rules, one-row-per-prediction-unit feature table, data-quality checks (row counts, uniqueness, ranges).
  6. Split before any fitting — copy the plugin's guard library to ml_pipeline/guard.py (its path is given at session start; in Codex/Kimi it is skills/ml-pipeline/lib/mlpipeline_guard.py next to this file) and split with it: train, val, test = guard.split(df, target=..., time_col=<col> if the data is temporal, group_col=<col> if the same entity appears in multiple rows). It drops exact duplicates, refuses a random split when a datetime column exists, keeps groups together, checks that no row lands in two splits, and freezes a fingerprint of the test set. The test set is touched exactly once, at step 14, through guard.final_test().
  7. Feature engineering — design features on the training set's statistics only, then apply the same transformations to validation/test.
  8. Preprocessing — scalers, encoders, imputers fit on train only, wrapped in a pipeline object so train and inference can never diverge.
  9. Baseline first — a dummy predictor (majority class / mean) AND one simple model (logistic/linear regression or small tree). Record their metrics. Every later model is judged against this line; a complex model that can't beat it gets reported as such.
  10. Model training — train candidate models on train, compare on validation. Log every run's config and score.
  11. Hyperparameter tuning — on validation/cross-validation only. The test set is never part of tuning.
  12. Model evaluation — the contract's metric plus supporting views: confusion matrix and ROC/PR curves for classification, residual plots for regression, always compared to the baseline. Figures required.
  13. Error analysis — worst predictions, performance by slice/subgroup, calibration, where the model fails and a hypothesis for why.
  14. Final test — score = guard.final_test(model.predict, test, target=..., metric_fn=...). It verifies the frame is the frozen test set, refuses a second call, and logs the result to PIPELINE.md. Report the number honestly, even if it is worse than validation. No going back to tune on it — if the result forces changes, agree a new test strategy with the user and record - Override: final test - user approved a second evaluation YYYY-MM-DD - reason: <why>.
  15. Deployment — only when the user asks. Save the full pipeline artifact (preprocessing + model together), verify a reloaded artifact reproduces predictions, document the inference input contract.
  16. Monitoring + retraining — write down: what drift to watch (input and prediction distributions), what metric threshold triggers retraining, and how retraining reuses this same pipeline from step 1.
Show full SKILL.md (435 more words)Show less

Enforcement in Claude Code (hook)

Installed as a Claude Code plugin, hooks/guard_training.py runs before every Bash, Write, Edit, and notebook tool call, scans the code for fitting and training calls, and denies:

  • model training until PIPELINE.md has Gate B: approved … or an Override: Gate B … line. With no ml_pipeline/PIPELINE.md at all, training is denied with a pointer to step 1.
  • any fitting — scalers, encoders, imputers included — until Gate A: approved ….
  • fitting on the test split (X_test, df_test, test_*) — always, at every gate.
  • evaluating on the test split — .predict/.score or a metric function whose arguments name X_test, df_test, test_* — until Gate C: approved … (or Override: Gate C …).
  • recording a gate before its artifacts exist — Gate A needs data_profile.json, one 01_*.png and two 02_*.png; Gate B needs a 04_*.png and the Model rationale: block; Gate C a 12_*.png; Gate D a 13_*.png and the Step 14 final test: line. Every PNG must be larger than 1 KB and have its - Figure <file>: line. The denial lists exactly what is missing.
  • red flags the profile can prove — neural-net-small-data, random-split-temporal, group-split, accuracy-imbalanced, resample-before-split (see references/model-selection.md). Override one only after the user agreed: - Override: red flag <name> - user approved <what> YYYY-MM-DD - reason: <why>.

A PostToolUse hook also warns (never blocks) when a result looks too good to be honest (accuracy/AUC/F1 ≥ 0.98) or when train_test_split() is called without stratify= or on data that mentions dates. To prove a pipeline does not leak, run it on examples/canary/canary.csv: honest test accuracy cannot exceed 0.80, and mlpipeline_canary.verdict(score, n_test) says whether a number is above the ceiling.

A denial is not an obstacle to route around: do the missing steps, get the user's explicit approval, record the gate line, and retry. The hook fails open on its own errors, and ML_PIPELINE_ENFORCE=0 switches it off for projects that are not ML pipelines. Codex and Kimi have no hook support, so there the same rules apply as instructions only.

Non-negotiables

  • Selection leakage is the one that matters most: the test set is never used to pick a model, a seed, a feature, or a threshold. Validation only.
  • Every gate is recorded with its figures on disk and their explanations in PIPELINE.md. A figure that was described but not drawn does not exist.
  • Test set is used exactly once. No tuning, no peeking, no "just checking".
  • All fitting (cleaning statistics, features, preprocessing, models) uses training data only.
  • Temporal data gets temporal splits; repeated entities get group splits.
  • Baseline before any complex model; every result is reported relative to it.
  • Failures and disappointing numbers are reported plainly — never hidden or reframed.
  • Chat explanations stay beginner-friendly; the code stays production-grade.

© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/jananthan30/ml-pipeline/skills/ml-pipeline of hashgraph-online/awesome-codex-plugins.

  • SKILL.md
  • references/model-selection.md

Open the folder on GitHubat commit 9e7b281

Compare with similar skills

ML Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Pipeline this skillhashgraph-online/awesome-codex-plugins1.3k—~3.2kAutomated safety check: PassApache-2.0
Plot ML Figureprobabl-ai/skills138—~785Automated safety check: PassBSD-3-Clause
Paper FiguresEvoScientist/EvoSkills4761 repos~4.4kAutomated safety check: PassApache-2.0
Analytics Data AnalysisMindrally/skills269—~1.6kAutomated safety check: PassApache-2.0
Scientific Toolkit SkillzLanqing/codex-claude-academic-skills4.7k—~1.2kAutomated safety check: PassMIT
Senior Data Scientistborghei/Claude-Skills886—~1.7kAutomated safety check: PassMIT

Similar skills

  • Plot ML Figure

    probabl-ai/skills

    Pick how to write a figure before custom plot code. An agent skill from probabl-ai/skills.

    138 GitHub stars~785 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Paper Figures

    EvoScientist/EvoSkills

    A skill your agent uses to produce standalone, publication-ready PNG graphics and reproducible matplotlib scripts from tabular data (CSVs or DataFrames).

    476 GitHub starsUsed in 1 repo~4.4k tokens
    Data & AnalyticsAuto-check passed
  • Analytics Data Analysis

    Mindrally/skills

    Best practices for analytics, data analysis, and visualization using Python, pandas, matplotlib, seaborn, and Jupyter notebooks.

    269 GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.7k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    borghei/Claude-Skills

    A skill your agent uses when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model…

    886 GitHub stars~1.7k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Matplotlib

    zLanqing/codex-claude-academic-skills

    Low-level plotting library for full customization. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 17 repos~2.9k tokens
    Data & AnalyticsAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 714 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.3k GitHub stars~922 tokensUpdated today
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.3k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.3k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.3k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Calle

    hashgraph-online/awesome-codex-plugins

    Use CALL-E from Codex through the calle CLI. An agent skill from hashgraph-online/awesome-codex-plugins.

    1.3k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.3k GitHub stars~618 tokensUpdated today
    Auto-check passed

Questions about ML Pipeline

What does ML Pipeline do?

MANDATORY whenever a task involves training, fine-tuning, tuning, or evaluating a machine-learning model on data (tabular, time series, text, images — any modality). ML Pipeline is an agent skill from hashgraph-online/awesome-codex-plugins. MANDATORY whenever a task involves training, fine-tuning, tuning, or evaluating a machine-learning model on data (tabular, time series, text, images — any modality).

When should I use ML Pipeline?

ML Pipeline fits situations like: tasks that involve MLOps; tasks that involve Fine-tuning; tasks that involve Data visualization.

How do I install ML Pipeline in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill ml-pipeline -a claude-code`. Or copy the skill folder (plugins/jananthan30/ml-pipeline/skills/ml-pipeline in hashgraph-online/awesome-codex-plugins) into .claude/skills/ml-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install ML Pipeline in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill ml-pipeline -a codex`. Or copy the skill folder (plugins/jananthan30/ml-pipeline/skills/ml-pipeline in hashgraph-online/awesome-codex-plugins) into .agents/skills/ml-pipeline in your project. Codex loads it when a task matches its description.

Can I use ML Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill ml-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-pipeline, .gemini/skills/ml-pipeline, .github/skills/ml-pipeline and .opencode/skills/ml-pipeline in your project.

What does ML Pipeline need to run?

Going by SKILL.md and its folder, ML Pipeline needs the command-line tools its instructions call (git).

Does ML Pipeline access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is ML Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Pipeline use?

ML Pipeline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Pipeline use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to ML Pipeline?

Skills that share tags, products or a category with ML Pipeline: Plot ML Figure (probabl-ai/skills, 138 stars), Paper Figures (EvoScientist/EvoSkills, 476 stars), Analytics Data Analysis (Mindrally/skills, 269 stars) and Scientific Toolkit Skill (zLanqing/codex-claude-academic-skills, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Pipeline?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,255 GitHub stars. The repository holds 714 skills in this directory. The repository was last updated on October 9, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.