Agent skill

Build ML Pipeline

by probabl-ai in probabl-ai/skills

Declare the pipeline from data source to predictor as a skrub DataOps graph.

BSD-3-ClauseAuto-check passedDevOps & Cloud

Install Build ML Pipeline

skills CLI
$ npx skills add probabl-ai/skills --skill build-ml-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install probabl-ai/skills build-ml-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/build-ml-pipeline .claude/skills/build-ml-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
build-ml-pipeline
GitHub stars
138
Token cost
~4.4k tokens
SKILL.md length
1,886 words
Files
8 (incl. references)
Skills in repo
23
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Declare the pipeline from data source to predictor as a skrub DataOps graph.

  • Works in 7 steps: python -m skore_skills status. Pending… → Emit 1–3 sentences before the first… → Emit Pre-flight. Tick a box only with… → …
  • Editing any link from data source to predictor (loaders
  • SKILL.md covers Human-facing prose, Procedure, Entry contracts and Rule 1 — Skrub DataOps is the…, plus 7 more sections
  • Calls python and git

What it does

Build ML Pipeline is an agent skill from probabl-ai/skills. Declare the pipeline from data source to predictor as a skrub DataOps graph. Stateless steps use .skb.applyfunc. Stateful steps use .skb.apply. After the graph exists, attach the locked splitter and any non-default score. No fit, tune, persistence, or skore.evaluate. TRIGGER when writing or editing any link from data source to predictor (loaders, preprocessing, features, composition, the final estimator), including a pure-Python function on that path, a step added or reordered, a bare sklearn.Pipeline as the…

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `evals/evals.json`, `references/common_patterns.md` and `references/custom-splitter.md`).

It sits in DevOps & Cloud, covering MLOps and Machine learning. It works with Python and scikit-learn. The repository describes itself as: Tabular Data Science Skills for guardrailing AI Agents. The licence is BSD-3-Clause.

When your agent uses it

  • Editing any link from data source to predictor (loaders
  • The final estimator)
  • Including a pure-Python function on that path
  • A bare sklearn.Pipeline as the top-level

Example prompts

  • “/build-ml-pipeline”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. python -m skore_skills status. Pending setup → S0. Then
  2. Emit 1–3 sentences before the first write: this is local
  3. Emit Pre-flight. Tick a box only with evidence from this turn.
  4. Declare build_learner (Rule 1–3). Sources, features, and
  5. Unfitted snapshot: references/snapshot.md.
  6. When experiments/NN_*.py exists for this stem, load
  7. On smoke run proceed, the checkpoint below, then

What it can do on your machine

Read from SKILL.md and the folder at commit 77bb26c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Build ML Pipeline loads about 4.4k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 198 tokens; SKILL.md has 1,886 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~198
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from probabl-ai/skills at commit 77bb26c, republished under its BSD-3-Clause licence (© probabl-ai). 1,886 words, ~4,393 tokens.

Download SKILL.mdSave it as .claude/skills/build-ml-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
build-ml-pipeline
description
Declare the pipeline from data source to predictor as a skrub DataOps graph. Stateless steps use `.skb.apply_func`. Stateful steps use `.skb.apply`. After the graph exists, attach the locked splitter and any non-default score. No fit, tune, persistence, or `skore.evaluate`. TRIGGER when writing or editing any link from data source to predictor (loaders, preprocessing, features, composition, the final estimator), including a pure-Python function on that path, a step added or reordered, a bare `sklearn.Pipeline` as the top-level, or a request to build a pipeline, classifier, or regressor. HOW TO USE: before the first declarative line and on every structural edit. Read the stops and emit Pre-flight before code. Confirm new names with `python -m skore_skills api get`.
metadata.role
helper
metadata.modelTier
big

Build ML Pipeline

Declare a skrub DataOps graph, then attach the locked splitter and any non-default score. Smoke is the last step here. Do not fit, tune, persist, or call skore.evaluate.

Human-facing prose

Details: setup-workspace references/human_facing_prose.md. Experiment markdown, design-note Method text, and # comments describe this pipeline. Questions and the checkpoint use the same data-science language: not skill ids, G-* names, or the wrapper CLI. <!-- results-embed: … --> is a site marker. Authoring hints stay in this skill.

X marker is .skb.mark_as_X(). Predict grid is the rows to score. Cross-row step reads other rows (lag, rolling, group-agg, side join). Layers 1 / 2 / 3 are sources, the grid plus marker, and features after the marker.

Procedure

  1. python -m skore_skills status. Pending setup → S0. Then python -m skore_skills design consent --stem <stem>. JSON action is authoritative. ask / stop → do not declare ("build it" is not approval). proceed continues. Then python -m skore_skills frame show without --revise. Anything other than proceed loads frame-ml-problem and stops. A proceed whose translation is null stops: there is no splitter translation.
  2. Emit 1–3 sentences before the first write: this is local preparation of an unfitted build_learner, Method cells, and a pipeline snapshot, then smoke. It does not train or run full-dataset evaluation. Do not invent a duration. Say it once. If approval is pending, preview only.
  3. Emit Pre-flight. Tick a box only with evidence from this turn.
  4. Declare build_learner (Rule 1–3). Sources, features, and the estimator first. Then attach cv and any non-default score from translation, still before .skb.make_learner(). python -m skore_skills api get for every new symbol. python -m skore_skills style after edits. Probes go to scratch/ via the composed-dev Python from env verify. No inline python -c. No warning filters unless the user asks.
  5. Unfitted snapshot: references/snapshot.md.
  6. When experiments/NN_*.py exists for this stem, load smoke-test-ml-pipeline only if that skill is installed. Missing skill → one line; do not invent the pytest file. Then python -m skore_skills smoke run --stem <stem>. JSON stop / red / smoke_missing → fix the graph here. Do not loosen the assertion. Do not evaluate. Do not claim pytest is green without this command.
  7. On smoke run proceed, the checkpoint below, then python -m skore_skills evaluate consent --stem <stem>.
    • ask — quote JSON context (question, experiment, smoke, persisted_report) in 2–4 lines: what the answer authorizes, the design question, and that this stem has no persisted report yet. Then AskUserQuestion, Evaluate (Recommended) / Modify / Stop, and stop. Do not write skore.evaluate. Modify → edit, then smoke run again. Stop → end.
    • proceed — a report already exists. Load evaluate-ml-pipeline if installed, or return to model-ml-pipeline when that skill called this one.
    • stop — smoke file missing. Do not evaluate. If the user answers Evaluate, load evaluate-ml-pipeline when installed. smoke run proceed is not that answer.
Checkpoint

After smoke run proceed, before the Evaluate menu. 2–6 sentences: what was declared, that smoke is green, the learner. Ground them in Method. Do not invent a metric. No locator yet.

Then links. When site build ran this turn: [report.html](<workspace>/report.html) and html/<stem>.html. Otherwise [journal/<stem>.md](journal/<stem>.md) and [experiments/<stem>.py](experiments/<stem>.py).

Entry contracts

Implement the approved design. Do not silently upgrade it.

  • Locked baseline — the one token in the journal baseline cell. dummy is DummyClassifier / DummyRegressor in the normal DataOps graph; it checks that the path runs. logistic, seasonal_naive, group_mean, and production are that one comparison model. api get the class.
  • EDA-backed — only Method-cited findings. A missing choice stops for a question. A temporal finding is the locked cv, not a license for three layers, lags, or AlignXy unless Method names those steps.
  • Backlog / discussion — Method is the boundary.

On a feature, transform, or leakage question, load research-ml-practice if it is installed. Abstract the problem class. AskUserQuestion allow_multiple on declare rows that do not violate a stop. The question's last line is exactly: Select each one you want. Write nothing after that line. measure revisits EDA and does not edit data_analysis.py. The splitter still comes from frame show. evaluate names evaluate-ml-pipeline. confirm asks the user. Missing skill → one line.

Rule 1 — Skrub DataOps is the pipeline entry point

Root at skrub.var(...), not a bare sklearn.Pipeline. skrub.X / skrub.y are not roots (S4). If the user asks for sklearn.Pipeline or build_pipeline(), do not import Pipeline, even as an inner estimator. Redirect to skrub.var and build_learner returning predictions.skb.make_learner(). Do not illustrate the refusal with a Pipeline([...]) constructor.

python
import skrub
from sklearn.ensemble import HistGradientBoostingRegressor

from <pkg>.data import TARGET_COL, load_raw


def build_learner(data_dir_preview=None):
    """Return the unfit learner (skrub SkrubLearner)."""
    data_dir = (
        skrub.var("data_dir", value=str(data_dir_preview))
        if data_dir_preview is not None
        else skrub.var("data_dir")
    )
    data = data_dir.skb.apply_func(load_raw)
    X = data.drop(columns=[TARGET_COL]).skb.mark_as_X()
    y = data[TARGET_COL].skb.mark_as_y()
    predictions = X.skb.apply(
        HistGradientBoostingRegressor(random_state=0), y=y
    )
    return predictions.skb.make_learner()

translation.task multioutput-regression predicts every column in translation.targets together. Drop all of them from X. y stays a DataFrame, one column per output:

python
TARGET_COLS = ["load", "temp"]
X = data.drop(columns=TARGET_COLS).skb.mark_as_X()
y = data[TARGET_COLS].skb.mark_as_y()
predictions = X.skb.apply(
    HistGradientBoostingRegressor(random_state=0), y=y
)

Use an estimator that fits a numeric target matrix, such as HistGradientBoostingRegressor, RandomForestRegressor, or Ridge. When the chosen estimator fits only one output, wrap it with MultiOutputRegressor. Confirm both symbols with api get. Do not use MultiOutputClassifier for this task, and do not leave a target column in X. One column still uses TARGET_COL as above. Details: references/common_patterns.md.

A quick traditional baseline uses skrub.tabular_pipeline (or TableVectorizer) plus one task-appropriate estimator. Confirm both with api get. Do not hand-tune columns or search hyperparameters here. Other shapes: references/common_patterns.md.

Rule 2 — Mark X early; featurize after

Per-row math and fit-time encoders: marker on the loaded frame. Any cross-row step: marker upstream of that step. Code for the three layers, including the loader-baked horizon refusal: references/layer_examples.md. Read it before proposing Layer 2.

value= is preview only. Expose data_dir_preview=None on build_learner. Do not bake a relative path into pipeline.py.

Splitter and scoring, after the graph

Read translation from the frame show that returned proceed. Attach on the existing X marker, then .skb.make_learner(). skrub requires cv= whenever split_kwargs is set. Integer cv is not a splitter. Evaluate omits splitter= so skore reuses this cv (evaluate-ml-pipeline/references/metadata-routing.md).

That cv copies the locked deployment: a scored fit contains only rows that deployment would already have seen. The shapes below are the usual ones. When the deployment is a different structure, open references/custom-splitter.md and write split for that setting.

  • scheme date_time — open the time series section of references/custom-splitter.md. cv= is the project-local class that section describes. Build one of those splitters per entry in translation.horizons. Each one is one predictor. split_kwargs carries the timestamp values. The timestamp column is the one the EDA or the text shipped with the data already names. Ask which column holds the timestamps only when those sources do not name one. time_role covariate keeps that column in the features. sort_key drops it from the features and still passes it in split_kwargs.
  • splitter GroupKFold — cv=GroupKFold(n_splits=<folds>) and split_kwargs={"groups": data["<translation.groups>"]}.
  • splitter KFold — cv=KFold(n_splits=<folds>) and empty split_kwargs.
  • splitter prefit — the data already has a training table and a test table. No cv. The graph loads the training table only. Do not add the test table as a second source, and do not concatenate the two files. Both tables share the schema the loader expects, including the target. The experiment binds the test table later. Do not call train_test_split.
  • Holdout (report EstimatorReport and splitter null) or translation null — no cv. This holdout is one split drawn from a single table. It is not the shipped train/test pair.

Scoring uses translation.metric. For multioutput-regression, skore already shows each output of MSE, RMSE, MAE, and R² (multioutput="raw_values"). The locked comparison is still one number, so attach .skb.with_scoring(...) for that name with an explicit aggregation: uniform_average unless the user named variance_weighted or a weight per output. That aggregate is the headline. Leave the per-output metrics in the report. For any other task, if the name is one skore already reports (regression: MSE, RMSE, MAE, R²; binary: accuracy, precision, recall, F1, ROC-AUC; multiclass: macro and micro variants; multilabel: per-label and averages), do not attach a scorer. Otherwise .skb.with_scoring(...) in this same late step. The callable shape is evaluate-ml-pipeline/references/custom-metrics.md. Do not pass scoring= to skore.evaluate.

Do not write skore.evaluate(...) here. Do not call train_test_split from pipeline code.

Show full SKILL.md (607 more words)Show less

Rule 3 — Attach: stateless function, stateful estimator

  • .skb.apply_func(fn) — output depends only on the current row and constants.
  • .skb.apply(estimator) — learns on training, reapplies on test.
  • skrub.deferred — rare; only when combining several DataOps and no skrub joiner fits. Default is apply_func. Details: references/source-binding.md.

A custom class is the mixin, then BaseEstimator. The mixin is the first base. That order matters: BaseEstimator first hides the mixin, because BaseEstimator.__sklearn_tags__ is resolved before it and does not call through. A transformer is TransformerMixin, BaseEstimator. A predictor is RegressorMixin, BaseEstimator or ClassifierMixin, BaseEstimator. BaseEstimator alone has no score, so skrub never exposes SkrubLearner.score and a with_scoring name never appears in summarize(). Details: references/common_patterns.md.

Would the output change on the training subset versus the whole frame? Yes → .skb.apply. Means, medians, quantiles, vocabularies, target encoding, TF-IDF: stateful.

STOP — target encoding / apply_func. When the user asks for
`def target_encode` + `.skb.apply_func`: refuse. Do not paste
the leaky function body and then the fix. Propose sklearn
TargetEncoder (or TransformerMixin, then BaseEstimator) via
`.skb.apply`. Name `api get` for the signature.

Reproducibility

done History rows stay runnable. Details: references/reproducibility_mechanics.md.

  • Option 1 — a default-preserving flag on the existing function. Small append. The default keeps prior callers unchanged (include_calendar_features: bool = False).
  • Option 2 — a new function called only from the new experiment.
  • Option 3 — branch the module. Last resort.

Three or more flags, or a flag that changes an existing caller's default → Option 2, or stop. After the change, pytest all of tests/smoke/.

Stops

S0. Setup still pending

status first. If status.setup.pending is non-empty and status.skills.setup-ml-project is true, load setup-ml-project and stop. Do not start this skill. When it returns, continue. Do not load it again on this turn. If that skill is not installed, name the pending pieces in one line and stop. Do not invent git init, scaffold, or env init. If status.setup.env or status.setup.workspace is declined, stop in one line. A declined git or editable is not asked again; continue. Missing design: design consent; ask / stop stay here. Missing data contract: explain and stop.

S0b. Modeling decisions are not locked
  • Rule: status.modeling_decisions must be locked, and python -m skore_skills frame show must return proceed with a non-null translation, before any declaration.
  • Recovery: load frame-ml-problem when installed. A null translation stops with no pipeline. Do not invent the table.
S1. Missing dependency

import skrub / sklearn failure, or a DataOp HTML stub, loads add-python-package for skrub and scikit-learn. Do not env add here. Do not substitute sklearn.Pipeline.

S2. Symbol from memory is forbidden

Every new skrub, sklearn, or skore name comes from api get or a matching cache read this turn.

S4. skrub.X / skrub.y are not graph roots

Root on skrub.var("<source>", value=preview). An existing skrub.X graph: show the alternative and ask. Do not auto-rewrite. Catalogue: references/source-binding.md.

Refuse: skrub.X / skrub.y are not graph roots (S4).
They bake the marker at the source and defeat Layer 1.

Alternative (refactor — ask before rewriting):
  data = skrub.var("data_dir", value=preview).skb.apply_func(load_raw)
  X = data.drop(columns=[TARGET_COL]).skb.mark_as_X()
  y = data[TARGET_COL].skb.mark_as_y()
S5. Late mark_as_X when any feature is cross-row

The marker sits upstream of every cross-row step. Symptom: len(predictions) != n_predict_grid_rows, feature_steps=[], or a wrapper whose job is to filter NaNs the pipeline produced. Recovery: references/layer_examples.md. Do not loosen smoke.

A loader that computes y = col.shift(-H) and then marks X is S5 and S6. Horizon lives in Layer 2. Do not invent a wrapper estimator that shifts, dropnas, or filters nulls.

S6. Layer 1 doesn't know the question

Loaders describe what data exists. Horizon, lag, window, and task filters belong in Layer 2+. If an external consumer could not derive the step without knowing the task, push it past Layer 1.

Pre-flight — emit before any code

Pre-flight (build-ml-pipeline):
- [ ] sklearn, skrub, skore import
      Evidence: scratch/<ts>_check_tier1.py. Not inline python -c.
- [ ] api get for new skrub and sklearn symbols this turn
- [ ] Each skrub.var is a source id, not a baked path
- [ ] mark_as_X placement (loaded frame, or predict grid if cross-row)
- [ ] Layer 1 has no horizon, lag, or task filter
- [ ] cv on mark_as_X matches translation
      (date class | GroupKFold | KFold | no cv on holdout
      | prefit: training table only, test table not joined)
- [ ] Non-default translation.metric uses with_scoring
      before make_learner (n/a for a listed skore default;
      required for a multi-output regression aggregate)
- [ ] data_dir_preview=None; no path literal in pipeline.py

Re-emit it with evidence before the final message.

References

  • references/snapshot.md — unfitted Method report, experiment cells, optional site build.
  • references/layer_examples.md — three layers. Read before proposing Layer 2.
  • references/source-binding.md — identifier versus materialized roots.
  • references/reproducibility_mechanics.md — Option 1 / 2 / 3.
  • references/common_patterns.md — tabular shapes with code, including a custom predictor (mixin, then BaseEstimator).
  • references/custom-splitter.md — the split copies the locked deployment. The time series section is the date splitter when translation.scheme is date_time.
  • evaluate-ml-pipeline/references/metadata-routing.md — evaluate omits splitter=.
  • evaluate-ml-pipeline/references/custom-metrics.md — a score that is not a skore default.

© probabl-ai, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in skills/build-ml-pipeline of probabl-ai/skills.

  • SKILL.md
  • evals/evals.json
  • references/common_patterns.md
  • references/custom-splitter.md
  • references/layer_examples.md
  • references/reproducibility_mechanics.md
  • references/snapshot.md
  • references/source-binding.md

Open the folder on GitHubat commit 77bb26c

Compare with similar skills

Build ML Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Build ML Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Build ML Pipeline this skillprobabl-ai/skills138—~4.4kAutomated safety check: PassBSD-3-Clause
MLflow Experiment TrackingOrchestra-Research/AI-Research-SKILLs13k2 repos~3.9kAutomated safety check: PassMIT
Editomegaml/omegaml108—~206Automated safety check: PassApache-2.0
Oci Data Scienceoracle/accelerated-data-science125—~2.1kAutomated safety check: PassUPL-1.0
Time Series Analytics Useropen-edge-platform/edge-ai-libraries171—~3.1kAutomated safety check: PassApache-2.0
SwanLab Experiment TrackingOrchestra-Research/AI-Research-SKILLs13k—~2.4kAutomated safety check: PassMIT

Similar skills

  • MLflow Experiment Tracking

    Orchestra-Research/AI-Research-SKILLs

    Tracks ML experiments, versions models in the MLflow registry and covers deployment and reproducibility, with autologging for common frameworks.

    13k GitHub starsUsed in 2 repos~3.9k tokens
    DevOps & CloudAuto-check passed
  • Edit

    omegaml/omegaml

    how to use the edit command properly

    108 GitHub stars~206 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Oci Data Science

    oracle/accelerated-data-science

    Official

    OCI Data Science service patterns including Jobs, Pipelines, Model Catalog, authentication, and the ADS SDK beyond AQUA.

    125 GitHub stars~2.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    171 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • SwanLab Experiment Tracking

    Orchestra-Research/AI-Research-SKILLs

    Shows how to log ML runs, configs, metrics and media with SwanLab and view them in cloud, local or self-hosted dashboards.

    13k GitHub stars~2.4k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • ML Engineer

    RightNow-AI/openfang

    Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps

    18k GitHub stars~987 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from probabl-ai/skills

All 23 skills in this repo
  • Add Python Package

    probabl-ai/skills

    Add a Python dependency through the project env manager, or ask the user to install it when env.managed is false.

    138 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Evaluate ML Pipeline

    probabl-ai/skills

    Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

    138 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Setup Workspace

    probabl-ai/skills

    Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.

    138 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Audit ML Pipeline

    probabl-ai/skills

    Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

    138 GitHub stars~9.6k tokensUpdated yesterday
    Auto-check passed
  • Explore ML Data

    probabl-ai/skills

    Owns data understanding before any model is designed. An agent skill from probabl-ai/skills.

    138 GitHub stars~6.4k tokensUpdated yesterday
    Auto-check passed
  • Export ML Notebook

    probabl-ai/skills

    Write a jupytext percent %% Python file out as an .ipynb. An agent skill from probabl-ai/skills.

    138 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed

Questions about Build ML Pipeline

What does Build ML Pipeline do?

Declare the pipeline from data source to predictor as a skrub DataOps graph. Build ML Pipeline is an agent skill from probabl-ai/skills. Declare the pipeline from data source to predictor as a skrub DataOps graph.

When should I use Build ML Pipeline?

Build ML Pipeline fits situations like: editing any link from data source to predictor (loaders; the final estimator); including a pure-Python function on that path; A bare sklearn.Pipeline as the top-level.

How do I install Build ML Pipeline in Claude Code?

Run `npx skills add probabl-ai/skills --skill build-ml-pipeline -a claude-code`. Or copy the skill folder (skills/build-ml-pipeline in probabl-ai/skills) into .claude/skills/build-ml-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Build ML Pipeline in Codex?

Run `npx skills add probabl-ai/skills --skill build-ml-pipeline -a codex`. Or copy the skill folder (skills/build-ml-pipeline in probabl-ai/skills) into .agents/skills/build-ml-pipeline in your project. Codex loads it when a task matches its description.

Can I use Build ML Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add probabl-ai/skills --skill build-ml-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-ml-pipeline, .gemini/skills/build-ml-pipeline, .github/skills/build-ml-pipeline and .opencode/skills/build-ml-pipeline in your project.

What does Build ML Pipeline need to run?

Going by SKILL.md and its folder, Build ML Pipeline needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Build ML Pipeline access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Build ML Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Build ML Pipeline use?

Build ML Pipeline is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Build ML Pipeline use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.7k tokens, read only when the agent opens those files.

What are the alternatives to Build ML Pipeline?

Skills that share tags, products or a category with Build ML Pipeline: MLflow Experiment Tracking (Orchestra-Research/AI-Research-SKILLs, 13k stars), Edit (omegaml/omegaml, 108 stars), Oci Data Science (oracle/accelerated-data-science, 125 stars) and Time Series Analytics User (open-edge-platform/edge-ai-libraries, 171 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Build ML Pipeline?

probabl-ai (a GitHub organization) maintains it in probabl-ai/skills, which has 138 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 9, 2026.

Source: probabl-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.