Agent skill

Evaluate ML Pipeline

by probabl-ai in probabl-ai/skills

Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

BSD-3-ClauseAuto-check passedDevOps & Cloud

Install Evaluate ML Pipeline

skills CLI
$ npx skills add probabl-ai/skills --skill evaluate-ml-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install probabl-ai/skills evaluate-ml-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluate-ml-pipeline .claude/skills/evaluate-ml-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-ml-pipeline
GitHub stars
138
Token cost
~4.3k tokens
SKILL.md length
1,971 words
Files
9 (incl. references)
Skills in repo
23
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

  • Works in 7 steps: python -m skore_skills status. If… → History-dependent pipeline (backward… → python -m skore_skills evaluate consent… → …
  • Code calls crossvalscore
  • SKILL.md covers Human-facing prose, Procedure, The call and After put, plus 4 more sections
  • Runs Python scripts from its folder; calls python, git and uv

What it does

Evaluate ML Pipeline is an agent skill from probabl-ai/skills. Evaluate one learner with skore.evaluate. The splitter and any non-default score are already on the DataOp from build-ml-pipeline. This skill writes the call without splitter=, persists the report, and records the locator. When the locked split is the training table and test table that shipped with the data, fit on the training table, then pass splitter="prefit" with only the test table. When model-ml-pipeline dispatched this turn, return the locator. A standalone turn also owns the evaluate-stage close. TRIGGER…

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `evals/evals.json`, `references/custom-checks.md` and `references/custom-metrics.md`).

It sits in DevOps & Cloud, covering MLOps. It works with Python. The repository describes itself as: Tabular Data Science Skills for guardrailing AI Agents. The licence is BSD-3-Clause.

When your agent uses it

  • Code calls crossvalscore
  • Classificationreport
  • .skb.crossvalidate
  • A handwritten metric print

Example prompts

  • “prefit”
  • “/evaluate-ml-pipeline”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. python -m skore_skills status. If status.setup.pending
  2. History-dependent pipeline (backward shift, lag, rolling
  3. python -m skore_skills evaluate consent --stem .
  4. G-SKORE-MODE. Read status.policy.skore_mode. If it is
  5. Emit Pre-flight, then 1–3 sentences: this is local
  6. Write the call in experiments/NN_*.py only. See the call
  7. After put has stored the report, copy templates/snapshot.py

What it can do on your machine

Read from SKILL.md and the folder at commit f273d39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • git
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate ML Pipeline loads about 4.3k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 245 tokens; SKILL.md has 1,971 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~245
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from probabl-ai/skills at commit f273d39, republished under its BSD-3-Clause licence (© probabl-ai). 1,971 words, ~4,321 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-ml-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
evaluate-ml-pipeline
description
Evaluate one learner with `skore.evaluate`. The splitter and any non-default score are already on the DataOp from `build-ml-pipeline`. This skill writes the call without `splitter=`, persists the report, and records the locator. When the locked split is the training table and test table that shipped with the data, fit on the training table, then pass `splitter="prefit"` with only the test table. When `model-ml-pipeline` dispatched this turn, return the locator. A standalone turn also owns the evaluate-stage close. TRIGGER when code calls `cross_val_score`, `cross_validate`, `classification_report`, `.skb.cross_validate`, or a handwritten metric print, or the user asks to score, evaluate, or run CV on one learner. Narrative reads of a persisted report belong to `audit-ml-pipeline`. HOW TO USE: before any evaluation call. Resolve G-SKORE-MODE, read the stops, and emit Pre-flight before code. Confirm symbols with `python -m skore_skills api get`.

Evaluate ML Pipeline

Score one learner and persist the report. The pipeline, the splitter, and any non-default score are already declared. skore.evaluate is the entry point. Do not hand-roll cross_val_score, cross_validate, classification_report, or metric prints.

A SkrubLearner does not implement sklearn's fit(X, y). cross_val_score raises. Call skore.evaluate(learner, data={...}) and omit splitter=, unless translation.splitter is prefit.

Human-facing prose

Details: setup-workspace references/human_facing_prose.md. Experiment markdown and # comments describe this evaluation. Questions and the close use the same data-science language: not skill ids, G-* names, or the wrapper CLI. <!-- results-embed: … --> is a site marker.

Procedure

  1. python -m skore_skills status. If status.setup.pending is non-empty and status.skills.setup-ml-project is true, load setup-ml-project and stop. Do not start this skill. When it returns, continue. Do not load it again on this turn. If that skill is not installed, name the pending pieces in one line and stop. Do not invent git init, scaffold, or env init. If status.setup.env or status.setup.workspace is declined, stop in one line. A declined git or editable is not asked again; continue. Then python -m skore_skills frame show without --revise. Anything other than proceed loads frame-ml-problem and stops. translation null → stop. Do not write an evaluation call.
  2. History-dependent pipeline (backward shift, lag, rolling window, target shift, or a join with side history): if tests/smoke/test_<stem>.py is missing or pytest is red, stop. Route to build-ml-pipeline. Do not write skore.evaluate. Documented n/a only when there is no history-dependent step.
  3. python -m skore_skills evaluate consent --stem <stem>.
    • ask, and this turn has not already answered Evaluate from build-ml-pipeline: quote JSON context and present Evaluate (Recommended) / Modify / Stop. Stop. "Run evaluation" is not consent on a first run.
    • The user already answered Evaluate this turn, or JSON is proceed (a real locator already exists): continue.
    • stop — smoke file missing. Route to build.
  4. G-SKORE-MODE. Read status.policy.skore_mode. If it is already set, keep it. If unset, ask local (recommended) / hub / mlflow, then python -m skore_skills policy set skore_mode <mode>. local: mkdir reports (exist_ok); do not write reports/README.md. hub or mlflow: do not create reports/. Then load add-python-package if installed; it runs env add-skore --mode <mode> --execute. Do not spell skore[...] here. Do not run env add here. Constructors: references/g_skore_mode.md. A switch or a migration loads sync-ml-reports when that skill is installed.
  5. Emit Pre-flight, then 1–3 sentences: this is local full-dataset evaluation with the locked scheme, report metrics, the skore checks, and project.put. The checks are computed in this run and stored with the report. When translation.splitter is prefit, say the model is fitted on the training table and scored on the shipped test table. Name the stem, the fold count when known, and experiments/<stem>.py plus scratch/results/<stem>/. Timing depends on rows, folds, the learner, and the checks. Do not invent minutes. If consent is still pending, preview and stop.
  6. Write the call in experiments/NN_*.py only. See the call shapes below. python -m skore_skills style after the edit. The experiment ends at report.checks.summarize(), project.put, and a bare report.
  7. After put has stored the report, copy templates/snapshot.py to scratch/results/<stem>/snapshot.py and run it. Then End of turn. Do not add snapshot writes to the experiment file.

Every Python probe goes to scratch/<ts>_<short>.py and runs with the composed-dev Python from env verify. No inline python -c. No warnings.filterwarnings unless the user asks. An unconfirmed signature stays unwritten. End that probe with BLOCKED: <class> signature needs an API lookup that cannot run this turn (<why>). A cache hit is a satisfied lookup.

The call

skore.evaluate in experiments/NN_*.py. Omit splitter= so skore reuses the DataOp cv and split_kwargs, unless translation.splitter is prefit. Passing any other splitter= drops split_kwargs. Omitted splitter= with no DataOp cv is an 80/20 holdout, correct only when translation.report is EstimatorReport and translation.splitter is null. Wiring: references/metadata-routing.md.

When translation.splitter is prefit, fit on the training table only, then score the test table. Confirm fit and evaluate with api get. Do not pass the training table into evaluate. Omitting splitter= here draws a new 80/20 split of the test table. Do not concatenate the two tables.

python
learner = build_learner()
fitted = learner.fit({"table": train_df})
report = skore.evaluate(
    fitted,
    data={"table": test_df},
    splitter="prefit",
)

An estimator whose fit is (X, y) uses the same rule:

python
fitted_model = LogisticRegression().fit(X_train, y_train)
report = skore.evaluate(
    fitted_model, X_test, y_test, splitter="prefit"
)

X / y and data stay mutually exclusive.

  • SkrubLearner — skore.evaluate(learner, data={...}). Keys are the skrub.var names. Interop: references/skrub_interop.md.
  • An estimator whose fit is (X, y) — skore.evaluate(estimator, X, y). Still omit splitter= when the locked cv is already the evaluation scheme. The prefit call above is the exception: the estimator is already fitted, and X and y are the test table.

The cv on mark_as_X is KFold, GroupKFold, or the date-based class from build. If translation names a cv and the marker has none, return to build-ml-pipeline. Do not wire split_kwargs here. Empty split_kwargs plus a possible group column → return to build-ml-pipeline. Do not default to KFold. translation.splitter prefit is not a cv on the marker.

No Stratified* for class imbalance. It compresses across-fold variance.

The headline is translation.metric. A name on the skore default list in build-ml-pipeline needs no scorer. Any other name must already be .skb.with_scoring(...) on the prediction DataOp. If it is not, return to build-ml-pipeline before skore.evaluate. Do not call report.metrics.add. When the scorer is attached, that name is a row in report.metrics.summarize().frame(). skore.evaluate has no scoring= argument (references/custom-metrics.md). If the name is attached and still missing from that frame, the predictor class is wrong: it must be the mixin, then BaseEstimator (RegressorMixin or ClassifierMixin first). Return to build-ml-pipeline. BaseEstimator alone, and BaseEstimator before the mixin, both fail.

An extra check the user asks for after the lock: references/custom-checks.md. Subclass skore.Check at module level in experiments/NN_*.py, then report.checks.add(...) after evaluate and before summarize. add extends SKD checks. Do not invent a check. Do not register one from audit/. Confirm Check and checks.add with api get.

Every evaluation calls report.checks.summarize() with no arguments after evaluate and after any checks.add, then project.put. That stores the check results on the report. Do not pass fast_mode or ignore. Confirm checks.summarize with api get.

python
report = skore.evaluate(...)
report.checks.summarize()
project.put(STEM, report)
report

Escalate past evaluate only when the dispatcher is too coarse (references/reports.md): EstimatorReport for one held-out fit, CrossValidationReport for per-fold artifacts. Holdout uses EstimatorReport. This loop scores that one learner with one skore.evaluate, one report.checks.summarize(), and one project.put.

CV is necessary but not sufficient for any pipeline with history-dependent features. skore.evaluate materializes the graph once with one env-dict and splits indices. The smoke test exercises a fresh env-dict at predict time. A passing smoke test is still required before the caller may flip status to done. Do not edit JOURNAL.md History or the design-note Status to done from this skill.

skore.evaluate(...) and project.put(...) live only in experiments/NN_*.py. A scratch probe, an audit file, or a notebook that calls them duplicates the report under the same key. Read a stored report with project.summarize() then project.get(id). get(key) raises KeyError because get is by id. Do not re-run evaluate to paper over that.

Plots that skore does not already draw load plot-ml-figure when that skill is installed.

Show full SKILL.md (831 more words)Show less

After put

Project.put returns None. Read project.summarize().frame(), take the newest row for the key, and form one locator. Copy templates/snapshot.py to scratch/results/<stem>/snapshot.py (gitignored). Substitute the Project init from experiments/<stem>.py, the report id, and that locator. Run the script with the composed-dev Python from env verify, before loop locator and loop artifacts. Do not put these writes in experiments/<stem>.py. Run python -m skore_skills loop locator --stem <stem> and paste JSON locator verbatim. Missing locator is n/a — backend did not expose a locator.

  • local — local workspace: [reports/](../reports/) · id: <id>, plus the absolute reports/ path. Do not link a private file.
  • hub — the exact Consult your report at … URL: [Open report](<url>) · hub · id: <id>. If Skore emits no URL, link the project landing page as Open project.
  • mlflow — an exact View run … URL when present. If absent and tracking_uri is HTTP(S), use /#/experiments/<experiment-id>/runs/<run-id>. For file:, sqlite:, databricks, or any other URI without an emitted URL: mlflow · tracking: <uri> · experiment: <id> · run: <run-id>. Do not invent a browser link.

The script writes report._repr_html_() to scratch/results/<stem>/report.html and repr(report) to report.txt. It regenerates the Method viewer from the stored learner: report.estimator_ on EstimatorReport, report.reports_[0].estimator_ on CrossValidationReport. Confirm eval on DataOp.skb.report with api get. learner.report forwards it. When eval is a parameter, call learner.report with eval=False, open=False, overwrite=True, and output_dir set to scratch/results/<stem>/pipeline. No environment. That does not fit. Do not call full_report, and do not call report without eval=False. If eval is absent, or Graphviz still fails after one add-python-package retry for skrub, write pipeline.html from _repr_html_ or sklearn.utils.estimator_html_repr.

End of turn

When model-ml-pipeline dispatched this turn, pass JSON locator up and return. Do not run loop notebooks, loop artifacts, audit, record-outcome, convert, site, or git end-turn.

Otherwise this skill owns the close. First run python -m skore_skills loop artifacts --stem <stem>. stop / evaluate_incomplete → name the missing file and do not audit. audit → load audit-ml-pipeline when installed. record → skip audit. record is that command's action, not the notebook gate. Do not convert in the audit skill.

Then this order. Do not reorder it. loop notebooks continues only on skip.

  1. python -m skore_skills loop notebooks --stem <stem>. Treat JSON action as authoritative. Do not record-outcome, site build, or git end-turn while action is convert. not_evaluated does not convert.
  2. On convert, run python -m skore_skills notebook convert <source> for every sources entry, with --html when html is true. Converting only audit/<stem>.py is not the close. The unfitted-snapshot ban does not apply: convert the experiment script even though this turn wrote skore.evaluate. If experiments/<stem>.py was in sources, re-run scratch/results/<stem>/snapshot.py and python -m skore_skills loop locator --stem <stem>. Re-run step 1 until action is skip. Convert re-executes the script. If convert fails because ipywidgets is missing, load add-python-package for it (agent) and convert again. Missing jupytext / nbclient / nbconvert → one line naming add-python-package.
  3. record-outcome, only once step 1 is skip, and before site build. If audit ran, pass its digest and G-AUDIT-FINDING into manage-ml-backlog when that skill is installed, with the locator from step 2 when the experiment script was converted. If audit did not run, call record-outcome with that locator, the user's headline, if any, and n/a — audit not run. Missing backlog skill → one line. Do not write History here. Never mark done while smoke is red.
  4. Checkpoint, below. Link report.html only when policy.site is true.
  5. site build only when policy.site is true and export-ml-site is installed: python -m skore_skills site build. It embeds audit/<stem>.nb.html under ## Notebooks; do not add <!-- results-embed: audit -->. If site build errors with mkdocs-material is required, load add-python-package for mkdocs-material (agent) and build once more. Do not pixi add / uv add. Name a build error; do not fail the turn.
  6. python -m skore_skills git end-turn --stage evaluate. If JSON action is invoke, load persist-ml-git when installed and stop. Otherwise load triage-ml-task when installed. Do not run git commit here.
Checkpoint

Standalone only. 2–6 sentences of the result, grounded in the headline or, when audit was skipped, report.txt. If audit ran, ground the story in its Checks and Metrics. Do not invent a metric.

Links: [report.html](<workspace>/report.html) and html/<stem>.html only when policy.site is true. Otherwise [journal/<stem>.md](journal/<stem>.md).

Tokens after the narrative: JSON locator first, then G-AUDIT-FINDING (n/a — audit not run when skipped).

Stops

  • Pending setup (status.setup.pending non-empty) → load setup-ml-project. A declined git is not asked again. A declined env or workspace stops.
  • Smoke missing or red on a history-dependent pipeline → build.
  • Empty split_kwargs plus a possible group column → return to build-ml-pipeline. Do not default to KFold.
  • import skore fails → G-SKORE-MODE if unset, then add-python-package. Do not drop back to cross_val_score.
  • Hyperparameter search, serving, and multi-run tracking are out of scope.
  • A missing package loads add-python-package when installed. Do not run env add here.

Pre-flight — emit before any code

Pre-flight (evaluate-ml-pipeline):
- [ ] sklearn, skrub, skore import
- [ ] frame show is proceed; DataOp cv matches translation
      (or holdout / prefit, and the marker has no cv)
- [ ] skore_mode is set (local | hub | mlflow)
- [ ] evaluate consent is proceed, or the user answered Evaluate
- [ ] Call site is experiments/NN_*.py
- [ ] Snapshots are scratch/results/<stem>/snapshot.py
      (not cells in the experiment file)
- [ ] skore.evaluate omits splitter=
      (or splitter="prefit" and only the test table is passed)
- [ ] report.checks.summarize() then project.put
      (no fast_mode, no ignore; after any checks.add)
- [ ] Smoke: passing | n/a (no history-dependent step) | STOP

References

  • references/metadata-routing.md — the locked cv stays on the DataOp; evaluate omits splitter=.
  • references/skrub_interop.md — env-dict versus (X, y).
  • references/g_skore_mode.md — Project constructors.
  • references/reports.md — when evaluate is too coarse.
  • references/custom-metrics.md — a non-default metric via with_scoring.
  • references/custom-checks.md — a check the user asked for.
  • templates/snapshot.py — agent-only post-put snapshot. Copy to scratch/results/<stem>/snapshot.py.

© probabl-ai, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/evaluate-ml-pipeline of probabl-ai/skills.

  • SKILL.md
  • evals/evals.json
  • references/custom-checks.md
  • references/custom-metrics.md
  • references/g_skore_mode.md
  • references/metadata-routing.md
  • references/reports.md
  • references/skrub_interop.md
  • templates/snapshot.py

Open the folder on GitHubat commit f273d39

Compare with similar skills

Evaluate ML Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate ML Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate ML Pipeline this skillprobabl-ai/skills138—~4.3kAutomated safety check: PassBSD-3-Clause
Azure AI ML Pymicrosoft/skills3.1k5 repos~2.2kAutomated safety check: PassMIT
SageMaker IAM Role Preflighthuggingface/skills11k1 repos~1.8kAutomated safety check: PassApache-2.0
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
MLflow Experiment TrackingOrchestra-Research/AI-Research-SKILLs13k2 repos~3.9kAutomated safety check: PassMIT
Oci Data Scienceoracle/accelerated-data-science125—~2.1kAutomated safety check: PassUPL-1.0

Similar skills

  • Azure AI ML Py

    microsoft/skills

    Official

    Azure Machine Learning SDK v2 for Python. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~2.2k tokens
    DevOps & CloudAuto-check passed
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    DevOps & CloudAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • MLflow Experiment Tracking

    Orchestra-Research/AI-Research-SKILLs

    Tracks ML experiments, versions models in the MLflow registry and covers deployment and reproducibility, with autologging for common frameworks.

    13k GitHub starsUsed in 2 repos~3.9k tokens
    DevOps & CloudAuto-check passed
  • Oci Data Science

    oracle/accelerated-data-science

    Official

    OCI Data Science service patterns including Jobs, Pipelines, Model Catalog, authentication, and the ADS SDK beyond AQUA.

    125 GitHub stars~2.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Py

    crazyguitar/pysheeet

    Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

    8.2k GitHub stars~886 tokensUpdated 2 days ago
    DevelopmentAuto-check passed

More from probabl-ai/skills

All 23 skills in this repo
  • Add Python Package

    probabl-ai/skills

    Add a Python dependency through the project env manager, or ask the user to install it when env.managed is false.

    138 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    138 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Setup Workspace

    probabl-ai/skills

    Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.

    138 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Audit ML Pipeline

    probabl-ai/skills

    Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

    138 GitHub stars~9.4k tokensUpdated yesterday
    Auto-check passed
  • Explore ML Data

    probabl-ai/skills

    Owns data understanding before any model is designed. An agent skill from probabl-ai/skills.

    138 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Export ML Notebook

    probabl-ai/skills

    Convert a jupytext percent %% Python file into an executed .ipynb with cell outputs.

    138 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Evaluate ML Pipeline

What does Evaluate ML Pipeline do?

Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills. Evaluate ML Pipeline is an agent skill from probabl-ai/skills.evaluate.

When should I use Evaluate ML Pipeline?

Evaluate ML Pipeline fits situations like: code calls crossvalscore; classificationreport; .skb.crossvalidate; A handwritten metric print.

How do I install Evaluate ML Pipeline in Claude Code?

Run `npx skills add probabl-ai/skills --skill evaluate-ml-pipeline -a claude-code`. Or copy the skill folder (skills/evaluate-ml-pipeline in probabl-ai/skills) into .claude/skills/evaluate-ml-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate ML Pipeline in Codex?

Run `npx skills add probabl-ai/skills --skill evaluate-ml-pipeline -a codex`. Or copy the skill folder (skills/evaluate-ml-pipeline in probabl-ai/skills) into .agents/skills/evaluate-ml-pipeline in your project. Codex loads it when a task matches its description.

Can I use Evaluate ML Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add probabl-ai/skills --skill evaluate-ml-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-ml-pipeline, .gemini/skills/evaluate-ml-pipeline, .github/skills/evaluate-ml-pipeline and .opencode/skills/evaluate-ml-pipeline in your project.

What does Evaluate ML Pipeline need to run?

Going by SKILL.md and its folder, Evaluate ML Pipeline needs Python for the scripts in its folder and the command-line tools its instructions call (python, git and uv). Our summary lists: Python 3.

Does Evaluate ML Pipeline access the network?

SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evaluate ML Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate ML Pipeline use?

Evaluate ML Pipeline is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate ML Pipeline use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.4k tokens, read only when the agent opens those files.

What are the alternatives to Evaluate ML Pipeline?

Skills that share tags, products or a category with Evaluate ML Pipeline: Azure AI ML Py (microsoft/skills, 3.1k stars), SageMaker IAM Role Preflight (huggingface/skills, 11k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars) and MLflow Experiment Tracking (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate ML Pipeline?

probabl-ai (a GitHub organization) maintains it in probabl-ai/skills, which has 138 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 8, 2026.

Source: probabl-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.