Agent skill

Audit ML Pipeline

by probabl-ai in probabl-ai/skills

Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

BSD-3-ClauseAuto-check passedDevOps & Cloud

Install Audit ML Pipeline

skills CLI
$ npx skills add probabl-ai/skills --skill audit-ml-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install probabl-ai/skills audit-ml-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/audit-ml-pipeline .claude/skills/audit-ml-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-ml-pipeline
GitHub stars
138
Token cost
~9.6k tokens
SKILL.md length
4,474 words
Files
7 (incl. references)
Skills in repo
23
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

  • Works in 9 steps: Module docstring (markdown cell) — what… → Imports (code cell) — import skore, from… → Open the Project (bare-expression cell)… → …
  • A completed run needs an audit
  • SKILL.md covers Human-facing prose, Next-step pointers, Where things live — visual map and Read-only contract, plus 17 more sections
  • Runs Python scripts from its folder; calls python, git and uv

What it does

Audit ML Pipeline is an agent skill from probabl-ai/skills. Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/. Do not execute it. Run materialize.py once. Never call skore.evaluate or project.put. TRIGGER when a completed run needs an audit, the user asks to audit, show, or re-audit an experiment, a re-run overwrote the same put() key, or they want a narrative of a past experiment without journal/ideas/ files. STOP when status.setup.pending is non-empty (load setup-ml-project), or when there is no…

Its SKILL.md is about 9.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `evals/evals.json`, `references/cell_anatomy.md` and `references/failure_modes.md`).

It sits in DevOps & Cloud, covering MLOps, Audit readiness and Data analysis. The repository describes itself as: Tabular Data Science Skills for guardrailing AI Agents. The licence is BSD-3-Clause.

When your agent uses it

  • A completed run needs an audit
  • The user asks to audit
  • Re-audit an experiment
  • A re-run overwrote the same put() key

Example prompts

  • “/audit-ml-pipeline”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Module docstring (markdown cell) — what this file is, the
  2. Imports (code cell) — import skore, from import ....
  3. Open the Project (bare-expression cell) — `project =
  4. List reports — summary = project.summarize(); then summary.
  5. Load the report — set REPORT_ID from the URL printed by
  6. Persisted report — substitute the exact normalized locator
  7. Checks summary — checks = report.checks.summarize(), then
  8. Metrics summary —
  9. Available report accessors — not a notebook cell. Copy

What it can do on your machine

Read from SKILL.md and the folder at commit 9e1b987. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • git
    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, uv and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit ML Pipeline loads about 9.6k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 252 tokens; SKILL.md has 4,474 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~252
When it runs · the whole SKILL.md, loaded when a task matches
~9.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from probabl-ai/skills at commit 9e1b987, republished under its BSD-3-Clause licence (© probabl-ai). 4,474 words, ~9,636 tokens.

Download SKILL.mdSave it as .claude/skills/audit-ml-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
audit-ml-pipeline
description
Read-only audit of one persisted skore report: `audit/NN_<stem>.py` (jupytext percent), 1:1 with `experiments/` and `journal/`. Do not execute it. Run `materialize.py` once. Never call `skore.evaluate` or `project.put`. TRIGGER when a completed run needs an audit, the user asks to audit, show, or re-audit an experiment, a re-run overwrote the same `put()` key, or they want a narrative of a past experiment without `journal/ideas/` files. STOP when `status.setup.pending` is non-empty (load `setup-ml-project`), or when there is no approved design, no report, or no agent feature: explain and send the user to setup, `evaluate-ml-pipeline`, or `model-ml-pipeline`. Also stop for raw-data exploration or sourcing a future experiment. HOW TO USE: confirm journal, experiment, smoke, and report; place the file from `templates/audit.py`; copy `templates/materialize.py` to scratch; execute once; derive G-AUDIT-FINDING; then the continue-or-close gate. Resolve skore symbols with `api get`.
metadata.role
helper
metadata.modelTier
medium

Audit ML Pipeline

Per-experiment, human-readable narrative of a skore report. The # %% file is the human notebook and is not executed. materialize.py writes the digest and the HTML viewers in one run. Read-only against the skore Project.

Human-facing prose

Details: setup-workspace references/human_facing_prose.md. Audit markdown and # comments describe this report's findings — not the skills framework, the CLI, or the command that produced an output. Questions, replies, and the close narrative use the same data-science language — not skill ids, G-* names, or the wrapper CLI. Trailing locator tokens stay index strings. Do not put API tutorials, version floors, or locator recipes in audit/<stem>.py. <!-- results-embed: … --> is a site marker. Authoring hints stay in this skill. style is ruff only.

Next-step pointers

Came here from…After audit, next is…
model-ml-pipeline (implement loop)→ Return. This audit execution may already materialize its notebook. The caller runs python -m skore_skills loop notebooks --stem <stem> as a freshness gate before record-outcome, then site build and git end-turn --stage implement.
User free-text ("audit 02", "re-audit 04")→ Surface metrics, then own the close (see § End of turn)
Re-run of an existing experiment→ Re-execute the existing audit file; surface diff if metrics changed

The implement loop dispatches audit after evaluate and before record-outcome. The digest carries the checks summary and metrics summary. G-AUDIT-FINDING is a separate normalized finding derived from that digest. manage-ml-backlog record-outcome consumes both and never dispatches audit back.

Where things live — visual map

PathDurabilityWho writes itWhat it holds
audit/<NN>_<short_name>.pyDurable (in git)This skill, once per experimentThe bare-expression cells. Source of truth. Not executed by the agent. The site notebook is this file with empty outputs
scratch/audit/<stem>/audit.mdEphemeral (gitignored)materialize.pyrepr of the checks display and the metrics frame. audit finding parses this
scratch/results/<stem>/materialize.pyEphemeral (gitignored)Evaluate, from evaluate-ml-pipeline/templates/materialize.pyThe only skore.evaluate and project.put. Writes report.html, report.txt, id.txt, checks.html, metrics.html, and the Method viewer
scratch/results/<stem>/report.html report.txt locator.txt id.txt pipeline/ or pipeline.htmlEphemeral (gitignored)Evaluate materialize.py, then the agent writes locator.txt from id.txtFull-report viewer, text fallback, locator, DataOp report or graph fallback
scratch/audit/<stem>/materialize.pyEphemeral (gitignored)This skill, from templates/materialize.pyRe-opens the report. Writes checks, metrics, audit.md, and extra viewers, plus accessors.txt
scratch/audit/<stem>/accessors.txtEphemeral (gitignored)materialize.pyhelp() trees. Additional report view labels come from the Displays groups here, not from the notebook
scratch/results/<stem>/checks.html metrics.html and extra <slug>.html / .pngEphemeral (gitignored)materialize.pyPer-item viewers the site embeds under ## Results. The digest already carries the text, so no extra .txt is written here
Execution digestCaptured from CLI output or the digest pathCLIStreamed or written digest — the agent reads this directly

Mnemonic: audit/ is source (in git); scratch/audit/ and stdout are output. Never copy audit/<stem>.py under scratch/audit/. materialize.py is the gitignored agent script. Never commit anything under scratch/audit/.

Read-only contract

The central rule. Surfaced as the first Stop condition below.

Allowed in audit/<stem>.py:

  • skore.Project(...) — open the project this experiment wrote to.
  • project.summarize() — list (key, id) pairs.
  • project.get(id) — load a specific report by id.
  • Every report.* accessor.
  • Imports from <pkg> (read-only inspection).

Forbidden in audit/<stem>.py:

  • skore.evaluate(...) — duplicates the report under the same key and pollutes summarize().
  • project.put(...) — same.
  • Snapshot writes (write_text of HTML, locator.txt, report.txt) and help() loops in audit/<stem>.py. Those live in scratch/audit/<stem>/materialize.py, which may write scratch/results/<stem>/ and accessors.txt.
  • Writes to data/, reports/, or src/<pkg>/. The audit notebook is a viewer.
  • Mutation of the loaded report that survives the cell (e.g. monkey-patching skore symbols).

The runner renders every cell's source + last-expression repr + stdout to the digest. A forbidden call surfaces in the digest (as a put row in a later summarize() cell, or as a **error:** section). The contract is visible, not invisible.

Sibling read-only consumers (different output shapes, same discipline): scratch/<ts>_*.py probes. review-ml-experiment reads the digest as text and does not open the Project. See evaluate-ml-pipeline § Stop conditions for the consumer rule.

Stop conditions — read before anything else

  • Setup still pending. Run python -m skore_skills status first. If status.setup.pending is non-empty and status.skills.setup-ml-project is true, load setup-ml-project and stop. Do not start this skill. When it returns, continue. Do not load it again on this turn. If that skill is not installed, name the pending pieces in one line and stop. Do not invent git init, scaffold, or env init. If status.setup.env or status.setup.workspace is declined, stop in one line. A declined git or editable is not asked again; continue.
  • No report → STOP. Four-way pairing is hard: run python -m skore_skills design consent --stem <stem> (ask / stop → do not audit) + experiments/NN_*.py + smoke pytest passed + a report under that key in the Project. If the report is missing, explain and stop. Do not skore.evaluate / project.put. Route to evaluate-ml-pipeline or model-ml-pipeline. Direct "audit 02" uses this same gate.
  • Read-only against the skore Project. See § Read-only contract. Never skore.evaluate(...) or project.put(...) in an audit file.
  • project.get(...) is by id, not key. For hub mode, read the id from the URL printed by project.put(): https://…/<workspace>/<project>/<type-plural>/<N> → id is skore:report:<type-singular>:<N> (URL segment is plural; id uses the singular — drop the trailing s, e.g. cross-validations → cross-validation, estimators → estimator). Hardcode REPORT_ID in the audit file — no summarize() traversal needed. For local mode, read the "id" column of project.summarize() for the matching key row. A KeyError from get("<stem>") means the lookup shape is wrong (get is by id), not that the report is missing.
  • Symbol from memory is forbidden. Any skore / skrub / sklearn symbol must come from python -m skore_skills api get this turn. Cache hits under scratch/api/skore/<version>/ count (Shape 0); inline memory does not.
  • Agent feature missing → STOP and delegate. If ipython / ipython aren't importable, do NOT fabricate audit outputs by writing print() calls as a workaround. Do NOT type pixi add ... / uv add ... yourself — install is owned by add-python-package § Agent feature. Request via agent tools (ruff / ipython / ipykernel) (binary: install / skip); resume only when add-python-package returns "ready".
  • Bare expressions, not print(). The runner captures each cell's last bare expression via result.result and renders its repr. Wrapping in print(repr(...)) lands in stdout instead of the output section; mixed and harder to scan. Use bare expressions; statement-only cells (variable binding) are fine.
  • One audit file per experiment stem (four-way pairing). No audit_NN_<short_name>_v2.py. When an experiment is re-run, the audit file is overwritten in place — same stem, same audit.
  • Executed artifacts go to scratch/audit/<stem>/, NOT into audit/. Durable artifact is audit/<stem>.py; the rendered digest is ephemeral.
  • audit/<stem>.py does not write snapshot files. HTML, locator.txt, and help() trees are materialize.py. No writes to data/ or reports/.
  • Don't filter warnings in audit cells. No warnings.filterwarnings(...) unless the user explicitly asks — the runner streams cell stderr into the digest and that's signal. See python -m skore_skills style § Stop conditions.
  • Harness "no clarifying questions" hints do NOT waive agent tools (ruff / ipython / ipykernel). Install gate fires regardless.
  • Post-hoc audit — required before ending the turn. Walk every pre-flight row; surface unfilled Evidence cells.
  • The post-audit gate is mandatory. After the first digest and every refresh, ask Additional report view / Custom query / Custom plot / Close audit in that exact order unless the user already explicitly closed the audit. The one execution may materialize a policy-enabled notebook and a stale-only site preview before this gate. No close-time conversion, git close, record-outcome, or dispatcher return before Close.

Forbidden shortcuts

ShortcutWhy it's wrong
report = project.get(REPORT_ID); print(repr(report))Runner captures bare expressions via result.result, not stdout. print(repr(...)) mixes stdout and output sections. Use report on its own line
checks.frame() instead of the bare checks DisplayThe Display __repr__ already groups issues and tips with their codes and documentation URLs. .frame() flattens that into a table the agent then has to re-read, and drops the severity grouping the review mines
project.get(KEY) raised KeyError → re-run evaluate + put "to refresh"Lookup shape is wrong (get is by id, not key). Hub: read the id from the URL printed by put(). Local: read summary["id"] for the matching key row. Never re-run evaluate + put to recover
Write pixi add --feature agent ipython directly from this skillInstall commands owned by add-python-package. This skill requests; it does not install
Dump the audit .py into scratch/audit/<stem>/The notebook stays in audit/. materialize.py is the only .py that belongs under scratch/audit/<stem>/, and it is gitignored
Register a Jupyter kernel "to be safe"Current runner is in-process; no kernel. Registering creates an orphan kernelspec
Add a fix-up cell that mutates data/ or reports/Audit files are read-only. State mutations belong in a scratch/<ts>_*.py probe or the experiment script
Substitute <SKORE_PROJECT_INIT> in audit/<stem>.py without reading experiments/<stem>.py firstAudit must open the same Project. Always Read experiments/<stem>.py this turn and copy the literal Project init block byte-identical (modulo formatting)
Hub mode: put skore.login(mode="hub") after skore.Project(...)Project(...) constructor authenticates at init time; without prior login, fails. Order is fixed: login first, Project second
Implement-loop audit → write scratch probe first to "double-check metrics"The audit IS the metric-extraction step. Scratch probes for metrics are the anti-pattern this dispatch replaces
write_text or a help() loop inside audit/<stem>.pyThe notebook is the human view. Copy templates/materialize.py to scratch/audit/<stem>/materialize.py for HTML, the locator file, and accessors.txt

Pre-flight — emit before any audit-file write or execution

Pre-flight (audit-ml-pipeline):
- [ ] Experiment stem confirmed: <NN_short_name>
      Evidence: journal/NN_<short_name>.md exists AND state is approved
                | "n/a — user invoked re-audit on existing stem"
- [ ] Four-way pairing complete:
        journal/NN_<short_name>.md       — design note (state approved or done)
        experiments/NN_<short_name>.py   — script
        tests/smoke/test_NN_<short_name>.py — smoke test (passing)
        audit/NN_<short_name>.py         — about to be written / refreshed
      Evidence: ls / Glob on each path
- [ ] Report present in skore Project under key=<NN_short_name>
      Evidence: project.summarize() this turn; row with
                key == "<NN_short_name>" appears.
                "Run finished, put() landed" is NOT sufficient.
- [ ] Agent feature available:
        run `ipython -c "print(0)"` and `ipython --version`
        through the project's composed dev environment
      Evidence: tool output of each
                | JOURNAL.md Status `agent feature: installed`
                Missing → STOP, delegate to add-python-package agent tools (ruff / ipython / ipykernel)
- [ ] API CLI consulted for skore symbols used:
      Project, summarize, get, report.checks.summarize, report.metrics.summarize
      Evidence: Read scratch/api/skore/<version>/<topic>.md (this turn)
                | Write the same (this turn)
                | "n/a — cache hit, file already on disk + Read this turn"
- [ ] Template copy + substitution decided:
        <pkg> → package name from src/<pkg>/
        <NN>_<short_name> → experiment stem
        <SKORE_PROJECT_INIT> → literal block copied from experiments/<stem>.py
      Evidence: Read experiments/<stem>.py this turn for the Project init block;
                Read templates/audit.py this turn before Write audit/<stem>.py
- [ ] Read-only contract acknowledged: audit file contains
      summarize / get / report.* only — no evaluate, no put
      Evidence: explicit grep / Read confirmation of the drafted file
- [ ] Execution command shape confirmed:
        python scratch/audit/<stem>/materialize.py
        notebooks true: notebook fill audit/<stem>.py [--html]
      Evidence: command emitted in the response before running
- [ ] G-AUDIT-FINDING derived from the executed digest
      Evidence: issue/tip counts + ordered codes + optional metric
                context | explicit clean-checks value
- [ ] Post-audit gate answered: additional view | query | plot | close
- [ ] Pre-flight re-emitted with evidence before final message.
      Evidence: this checklist appears in the end-of-turn summary.

Before execution

After the four-way pairing, report lookup, agent feature, and API checks are satisfied, emit 1–3 natural sentences immediately before writing or running audit/<stem>.py. Say that this is local read-only report materialization, not retraining: it opens the persisted Skore report, renders checks / metrics / available views, and writes the digest under scratch/audit/<stem>/ plus HTML viewers under scratch/results/<stem>/. Name the exact policy-matched execution command.

Describe cost from facts: report loading and requested view rendering, not model fits or full CV. Unless an observed duration is already available, say timing depends on report size and selected views and do not invent minutes. Emit this preview once for the initial audit; refresh it only when Additional report view / Custom query / Custom plot materially changes the work. The run pauses at the existing post-audit gate before close.

Questions about what the report means are LLM narrative work over the digest. For a fired check, read its documentation URL and apply that page's recommendation. They do not call evaluate or put. If any mandatory gate is pending, preview the possible audit but do not write or execute it.

Audit file contract — overview

The audit file is jupytext percent format (# %%). Filename: audit/NN_<short_name>.py — stem matches the experiment exactly. Template: templates/audit.py.

Substitutions
PlaceholderReplaced with
<pkg>The importable package name (from src/<pkg>/)
<NN>_<short_name>The experiment stem (e.g. 02_target_transform)
<SKORE_PROJECT_INIT>The full Project init block (including any preceding skore.login(...) call for hub mode), copied byte-identical from experiments/<stem>.py
<project-name>The name= argument from experiments/<stem>.py (read it; don't invent)
<hub-workspace>Hub-mode only. Copy from the workspace= argument in experiments/<stem>.py
<REPORT_LOCATOR>The normalized post-put locator handed off by evaluate. If absent, derive it from the selected id and backend rules below; never guess a URL

<SKORE_PROJECT_INIT> and <project-name> are the most error-prone substitutions: the audit must open the same Project the experiment wrote to. Always Read experiments/<stem>.py this turn to lift the literal init block; never reconstruct from memory of the skore mode: decision alone.

Cell sequence (what each cell does)

Brief outline; full anatomy with concrete examples → references/cell_anatomy.md.

  1. Module docstring (markdown cell) — what this file is, the read-only rule.
  2. Imports (code cell) — import skore, from <pkg> import ....
  3. Open the Project (bare-expression cell) — project = skore.Project(...); then project on its own line.
  4. List reports — summary = project.summarize(); then summary.
  5. Load the report — set REPORT_ID from the URL printed by project.put() (hub: "skore:report:<type-singular>:<N>" — URL path segment is plural, id uses singular, e.g. cross-validations → cross-validation, estimators → estimator; local and mlflow: read summary["id"] for the matching key row), then report = project.get(REPORT_ID), then report as the last expression. Do not write HTML or locator.txt in this cell. materialize.py does that (confirm _repr_html_ with api get).
  6. Persisted report — substitute the exact normalized locator from evaluate. For a direct audit, use the selected REPORT_ID and policy.skore_mode: local links ../reports/; Hub uses the exact saved put URL (or a clearly labeled project link); MLflow uses an emitted run URL or records tracking URI + experiment id + run id. If no authoritative locator exists, write n/a — backend did not expose a locator. Do not call put, inspect private storage, or invent a frontend URL.
  7. Checks summary — checks = report.checks.summarize(), then checks as the last expression. materialize.py writes checks.html. Its repr groups the walk by severity; every issue / tip line ends with the documentation URL. Read that page and apply its recommendation (custom CSTM* checks may have none).
  8. Metrics summary — metrics = report.metrics.summarize().frame(verbose_name=True, flat_index=False), then metrics last. materialize.py writes metrics.html from that frame. That frame is the only metrics table. Do not also paste its values into the design note.
  9. Available report accessors — not a notebook cell. Copy templates/materialize.py to scratch/audit/<stem>/materialize.py and run it. It calls help() on report.metrics, report.checks, and any other namespace that exists (inspection, data, …) and writes the trees to accessors.txt. help() prints its tree and returns None. Skip a namespace that is absent. Do not call plot accessors here. Do not put the loop in audit/<stem>.py.

That's the core template. Leave checks as a bare Display: editors render _repr_html_, the runner records the repr, and .frame() drops the severity grouping. Metrics are the frame above, not that Display and not a second markdown table. Extra Display cells are appended only after the user picks Additional report view. Details: → references/cell_anatomy.md.

The digest is the review's canonical source

The rendered digest at scratch/audit/<stem>/audit.md is the single source of truth that review-ml-experiment reads to write journal/ideas/<stem>-<slug>.md and one Ideas row in journal/JOURNAL.md. That skill reads the digest as text, walks Issues: then Tips: in ## Checks summary, and does not re-open the Project, call report.*, or write scratch/<ts>_*.py probes for metric extraction. manage-ml-backlog later triages those files. A promoted idea leaves the Ideas table and becomes a Backlog row.

The contract stays narrow: persisted-report locator + checks + metrics summary. The review must not walk extra Display headings. Do not put ROC / confusion-matrix / importance cells in the core template. After Close-audit extras, those cells use their own ## <title> headings. If the user explicitly asks for a custom figure that accessors.txt did not list, load plot-ml-figure if installed before writing the cell; never replace a skore Display. Save PNG (or HTML) and leave the figure visible; never plt.close in the audit notebook.

G-AUDIT-FINDING

After every successful audit execution, run python -m skore_skills audit finding --stem <stem> (or pass the digest path). Paste JSON finding verbatim. Do not rewrite codes, counts, or the metric clause. Missing digest → JSON finding is n/a — audit digest unavailable.

Extra Display cells do not change G-AUDIT-FINDING (re-run the command; it still reads only checks + metrics).

Return G-AUDIT-FINDING verbatim with the digest and G-REPORT-LOCATOR from python -m skore_skills loop locator --stem <stem>. manage-ml-backlog writes the finding to the design note's Status block. Recompute it after every added view, query, or plot because the digest is overwritten.

Continue auditing

After the initial digest and every successful refresh, unless the user already said to close, AskUserQuestion with one pick in this exact order. None is recommended or preselected.

Gate context. Ahead of the question, state in 2–4 lines what the answer authorizes and the facts it rests on — echoed inline from this turn's digest (the checks and metrics just read, the Display names in accessors.txt) — plus what each option does. A file link is an addition, never the context.

LabelContract
Additional report viewMenu labels are exactly the names under the Displays group of this turn's help() trees in scratch/audit/<stem>/accessors.txt. Confirm the picked name with python -m skore_skills api get. The list is task-dependent — a regression report has no roc. Omit anything the trees do not list; never show remembered Display names. Ask one pick before editing.
Custom queryWait for one concrete read-only question about the loaded report. Append only the minimal accessor cells needed to answer it, in audit/<stem>.py and in the materialize <CALL> block. Prefer a name from the trees when it answers the question.
Custom plotOnly when the trees have no Display for this chart. Load plot-ml-figure only if status.skills.plot-ml-figure is true; else one-line skip and return to this gate. Append the read-only plot to both files. materialize.py saves the figure under scratch/results/<stem>/.
Close auditContinue to the existing dispatched or direct close.

Additional report view / Custom query / Custom plot all edit the same durable audit/<stem>.py, below ## Core audit complete, and the <CALL> block of materialize.py. After an edit: run style, re-run materialize.py once, notebook fill when notebooks are on, derive G-AUDIT-FINDING again with python -m skore_skills audit finding --stem <stem> (from checks + metrics only), then run site build --if-stale when site policy is true and re-present this same gate. Do not run the close-time loop notebooks, run git end-turn, record-outcome, or return to the dispatcher before Close audit. The audit remains read-only: no evaluate, no put, and no workspace-data mutation.

Show full SKILL.md (1,698 more words)Show less
Extra Display cells

Slug = the accessor name accessors.txt listed under Displays. Do not invent slugs from docs memory.

  1. Markdown cell ## <human title> — not ## Checks summary or ## Metrics summary.
  2. Code cell: call the accessor; confirm _repr_html_ with api get on the returned Display. Last expression: the bare Display.
  3. In materialize.py only, append that accessor to EXTRA. It writes scratch/results/<stem>/<slug>.html from _repr_html_() when it exists. Only when it does not, save <slug>.png so the site can still embed a figure. Re-run materialize.py.
  4. Failed or inapplicable accessors stay in scratch/ probes. Do not add a Results subsection for them.

manage-ml-backlog turns each extra <slug> viewer (other than report, checks, metrics) into a ### Results heading with <!-- results-embed: <slug> -->, summarizing it from the digest cell that produced it.

Execution contract — one execution

Before the first execution for a stem, run python -m skore_skills review consent --stem <stem>.

  • stop — name the missing report.html and stop.
  • audit — write audit/<stem>.py and execute it once. Check results were stored with the report; this read uses them.
  • proceed — digest exists. Do not execute unless the user explicitly asked to re-audit. A re-audit repeats the one policy-matched execution.

Copy templates/materialize.py to scratch/audit/<stem>/materialize.py, substitute the Project init, report id, and locator, and run it once with the composed-dev Python from env verify. That run writes the HTML viewers, accessors.txt, and audit.md. Do not cells run the audit file and do not notebook convert it. When policy.notebooks is true, then run:

bash
python -m skore_skills notebook fill audit/<stem>.py [--html]

Add --html only when site is also true. Fill writes an output-free notebook and does not write the digest. When notebooks is null or false, do not change policy and do not fill. Then run:

bash
python -m skore_skills audit finding --stem <stem>

The notebook does not contain the materialize writes. If site policy is true and export-ml-site is installed, run python -m skore_skills site build --if-stale after materialize and before that gate, then link report.html and html/<stem>.html.

Paste JSON finding verbatim. Read audit.md; do not derive it from the notebook.

Executed notebook

When notebook policy is true, notebook fill above owns the output-free notebook on direct and dispatched paths. The caller runs python -m skore_skills loop notebooks --stem <stem> and obeys convert by filling, not by kernel-executing. Naming notebook fill is not that close. A current source fingerprint makes that gate skip; a changed source fills again as a safety net.

On a direct close, run the notebook gate in § End of turn before record-outcome. The audit has no page of its own. Site build embeds audit/<stem>.nb.html under the design note's ## Notebooks section, after the evaluation notebook, whenever that file exists. Do not add <!-- results-embed: audit -->. A marker does not create the viewer, and a missing .nb.html omits it.

Re-execution semantics
  • Re-running an experiment (overwriting put() under the same key) → re-execute the matching audit file. manage-ml-backlog § 4 fires this on every record-outcome.
  • Editing the audit file's source (adding a metric accessor) → re-execute. The digest is regenerable.
  • scratch/audit/<stem>/ is overwritten on every execution. No version history; the source .py + git history is the audit trail.

Four-way stem-pairing rule

Extends setup-workspace's pairing rule from three artifacts to four:

journal/NN_<short_name>.md           — design note
experiments/NN_<short_name>.py       — script
tests/smoke/test_NN_<short_name>.py  — smoke test
audit/NN_<short_name>.py             — audit  ← this skill

Identical stems, 1:1. By the time the experiment shows done in journal/JOURNAL.md, all four exist.

Dispatching in and out

Called from
CallerWhen
review-ml-experimentWhen review consent is audit, before idea files
model-ml-pipelineDoes not load this skill; it loads review-ml-experiment
evaluate-ml-pipelineStandalone evaluate may run audit before its close
User free-text"audit experiment 02", "show me what 03", "re-audit 04" — resolves directly
Calls into
CalleeWhy
python -m skore_skills audit findingAfter audit execution. Paste JSON finding verbatim
python -m skore_skills loop locatorG-REPORT-LOCATOR from locator.txt or the audit file
python -m skore_skills loop artifactsDirect-audit close: record before the notebook gate
python -m skore_skills loop notebooksBefore record-outcome. Obey convert; skip continues
add-python-packageWhen ipython is missing
manage-ml-backlog (record-outcome mode)End of turn on a direct free-text audit — hands over the digest so the History row and design-note Status block get written
python -m skore_skills styleAfter writing / editing audit/<stem>.py — bundled ruff.toml carries audit/** per-file ignores. Ruff only; it does not rewrite comments. Do not write workflow/process prose in the audit file

End of turn

Dispatched (model-ml-pipeline, evaluate-ml-pipeline, or manage-ml-backlog this turn), after Close audit: run python -m skore_skills audit finding --stem <stem> and python -m skore_skills loop locator --stem <stem>. Return the digest, JSON finding, JSON locator, and an optional headline to that caller. Stop. Do not run record-outcome, the close-time loop notebooks, another site build, git end-turn, or triage here. The live execution/preview above is not repeated. When the caller is model-ml-pipeline or evaluate-ml-pipeline, tell it to run python -m skore_skills loop notebooks --stem <stem> and obey convert with notebook fill before record-outcome. That gate continues on skip. record is loop artifacts, not the notebook gate. Then site build and that caller's git end-turn (--stage implement from model, --stage evaluate from evaluate). Naming the fill commands is not the close. Filling only the audit file does not finish that close. Do not paste the direct-audit close as a preview of what the dispatcher will run.

Direct free-text audit, after Close audit: this skill owns the close. Do not start it before Close audit. Run python -m skore_skills loop artifacts --stem <stem> first (record expected). record is that command's action, not the notebook gate. Then this order. Do not reorder it. loop notebooks continues only on skip.

  1. python -m skore_skills loop notebooks --stem <stem>. Treat JSON action as authoritative. Do not record-outcome, site build, or git end-turn while action is convert. not_evaluated does not convert.
  2. On convert, run notebook fill for every experiments/ or audit/ source, with --html when html is true. Do not notebook convert those files. Fill does not put. Re-run the matching materialize.py first only when that human call changed since the run that wrote the current artifacts. Re-run step 1 until action is skip. Missing jupytext / nbformat → one-line skip naming add-python-package; do not fail the audit, do not pixi add.
  3. Load manage-ml-backlog in record-outcome mode only if status.skills.manage-ml-backlog is true and hand it the digest, G-AUDIT-FINDING, and that locator. Missing skill → one-line skip; do not write History here. That mode records and returns; it does not re-dispatch this skill. Never mark done while smoke is red.
  4. site build only when policy.site is true and export-ml-site is installed: python -m skore_skills site build --if-stale. If it errors with mkdocs-material is required, load add-python-package for mkdocs-material (agent) and build once more. Do not pixi add / uv add. Name a build error; do not fail the audit.
  5. User-facing close. The message is 2–6 sentences from Checks + Metrics. Do not invent a metric or paste the digest. Link [report.html](<workspace>/report.html) and html/<stem>.html only when step 4 succeeded. Otherwise link [journal/<stem>.md](journal/<stem>.md). Then JSON locator from step 2 verbatim (local: also the absolute reports/ path), then G-AUDIT-FINDING. Index strings, not the narrative. Dispatched audit never writes this message.
  6. python -m skore_skills git end-turn --stage evaluate. The audit continues the evaluate stage; there is no audit stage. If JSON action is invoke, load persist-ml-git only if status.skills.persist-ml-git is true and stop; that skill returns to triage. If persist is missing, name the pending staged paths and stop. Otherwise load triage-ml-task only if status.skills.triage-ml-task is true; else stop. Do not run git commit in this skill.

Failure modes and recovery

Quick lookup; detailed recovery steps in references/failure_modes.md.

SymptomCauseFix
project.get(key) raises KeyError / TypeErrorLookup by key, not id; local vs hub shape differs→ references/failure_modes.md § "project.get(key) raises"
ModuleNotFoundError: No module named 'IPython'Agent feature not installedDelegate to add-python-package; never pip install here
Cell renders as <Display object at 0x…>skore too old to give the Display a text reprAdd .frame() to that cell only; leave the rest bare
AttributeError for a report.* accessorSymbol from memory; skore version drift→ references/failure_modes.md § "AttributeError"
RuntimeError: No report under key=...put() landed in a different Project→ references/failure_modes.md § "wrong Project"
Report differs across runs with unchanged sourceNon-deterministic step / different data sliceNot a bug here; surface to user
Hub mode: skore.login() auth errorToken expired / first-time login→ references/failure_modes.md § "skore.login fails"
Hub mode: TypeError: workspace kwargHub form left local-mode kwarg→ references/failure_modes.md § "TypeError workspace"
Hub mode: report missing in summarize() after put()Wrong hub workspace OR no read access→ references/failure_modes.md § "report missing"

What this skill does NOT do

  • Open or write the skore Project's reports (evaluate-ml-pipeline).
  • Install ipython (add-python-package owns).
  • Write journal/ideas/ files or Ideas rows (review-ml-experiment owns both).
  • Write or edit journal/NN_*.md or journal/JOURNAL.md directly. At end of turn, dispatch manage-ml-backlog record-outcome mode instead — that skill owns every journal write.
  • Run pytest / smoke tests (smoke-test-ml-pipeline).
  • Render commits or PRs.
  • Decide which metrics matter — the cells are filled per task, but judgment about what matters is the user's, in the design note.

Companion skills

SkillRelationship
manage-ml-backlogDownstream record-outcome consumer; never dispatches audit from record-outcome mode
review-ml-experimentLoop caller. Parses the digest as text and writes one idea file and one Ideas row per candidate. Never opens the Project
evaluate-ml-pipelineProducer side. The experiment file holds the call; materialize.py is the only run
setup-workspaceWorkspace layout; four-way stem pairing
add-python-packageAgent feature install (agent tools (ruff / ipython / ipykernel)). This skill requests; that skill installs
python -m skore_skills api getskore symbol lookups, including help and extra Display methods. Cache hits first
plot-ml-figureCustom plot gate only when accessors.txt has no Display
python -m skore_skills styleruff after writing/editing audit/<stem>.py
choose-python-library / python -m skore_skills env stackAgent tools (ipython, ipykernel) live under the agent feature

Templates and assets

  • templates/audit.py — per-experiment audit file skeleton. Copy
    • substitute; don't rewrite from memory. Bare displays only.
  • templates/materialize.py — the only audit run. Copy to scratch/audit/<stem>/materialize.py. Writes the HTML viewers, accessors.txt, and audit.md.

Need a package?

When an import is missing, load add-python-package if status.skills.add-python-package is true. That skill owns env add and the unmanaged ask. Do not run env add here. If the skill is not installed, name the package and stop.

References (load on demand)

  • references/cell_anatomy.md — concrete cell examples (right / wrong shapes), core sequence, why checks stay a bare Display and metrics use frame(verbose_name=True, flat_index=False), extra Display snapshot contract, bare-expression rules.
  • references/runner_internals.md — leftover runner internals (IPython, Agg). Prefer --help / the package docstring.
  • references/failure_modes.md — detailed recovery for every symptom in § Failure modes.

© probabl-ai, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in skills/audit-ml-pipeline of probabl-ai/skills.

  • SKILL.md
  • evals/evals.json
  • references/cell_anatomy.md
  • references/failure_modes.md
  • references/runner_internals.md
  • templates/audit.py
  • templates/materialize.py

Open the folder on GitHubat commit 9e1b987

Compare with similar skills

Audit ML Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit ML Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit ML Pipeline this skillprobabl-ai/skills138—~9.6kAutomated safety check: PassBSD-3-Clause
Tech Data PlaybookLeoYeAI/openclaw-master-skills2.2k—~6.8kAutomated safety check: PassMIT
ML Pipeline Workflowwshobson/agents40k12 repos~1.8kAutomated safety check: PassMIT
Editomegaml/omegaml108—~206Automated safety check: PassApache-2.0
ML Pipeline ExpertJeffallan/claude-skills12k—~1.9kAutomated safety check: PassMIT
Azure Storage File Datalake Pyaiskillstore/marketplace4334 repos~1.5kAutomated safety check: PassNone

Similar skills

  • Tech Data Playbook

    LeoYeAI/openclaw-master-skills

    World-Class Technology & Data Playbook. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~6.8k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • ML Pipeline Workflow

    wshobson/agents

    Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.

    40k GitHub starsUsed in 12 repos~1.8k tokens
    DevOps & CloudAuto-check passed
  • Edit

    omegaml/omegaml

    how to use the edit command properly

    108 GitHub stars~206 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub stars~1.9k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed
  • Azure Storage File Datalake Py

    aiskillstore/marketplace

    Azure Data Lake Storage Gen2 SDK for Python. An agent skill from aiskillstore/marketplace.

    433 GitHub starsUsed in 4 repos~1.5k tokens
    DevOps & CloudAuto-check passed
  • Monte Carlo Monitoring Advisor

    sickn33/agentic-awesome-skills

    Analyze data coverage, create monitors for warehouse tables and AI agents.

    47k GitHub starsUsed in 1 repo~5.2k tokens
    DevOps & CloudAuto-check passed

More from probabl-ai/skills

All 23 skills in this repo
  • Add Python Package

    probabl-ai/skills

    Add a Python dependency through the project env manager, or ask the user to install it when env.managed is false.

    138 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    138 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Evaluate ML Pipeline

    probabl-ai/skills

    Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

    138 GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Frame ML Problem

    probabl-ai/skills

    Record the problem, the deployment setting, the comparison metric, the baseline, and the fold count in the journal before model code.

    138 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Setup Workspace

    probabl-ai/skills

    Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.

    138 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Explore ML Data

    probabl-ai/skills

    Owns data understanding before any model is designed. An agent skill from probabl-ai/skills.

    138 GitHub stars~6.4k tokensUpdated today
    Auto-check passed

Questions about Audit ML Pipeline

What does Audit ML Pipeline do?

Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/. Audit ML Pipeline is an agent skill from probabl-ai/skills.py (jupytext percent), 1:1 with experiments/ and journal/.

When should I use Audit ML Pipeline?

Audit ML Pipeline fits situations like: A completed run needs an audit; the user asks to audit; re-audit an experiment; A re-run overwrote the same put() key.

How do I install Audit ML Pipeline in Claude Code?

Run `npx skills add probabl-ai/skills --skill audit-ml-pipeline -a claude-code`. Or copy the skill folder (skills/audit-ml-pipeline in probabl-ai/skills) into .claude/skills/audit-ml-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Audit ML Pipeline in Codex?

Run `npx skills add probabl-ai/skills --skill audit-ml-pipeline -a codex`. Or copy the skill folder (skills/audit-ml-pipeline in probabl-ai/skills) into .agents/skills/audit-ml-pipeline in your project. Codex loads it when a task matches its description.

Can I use Audit ML Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add probabl-ai/skills --skill audit-ml-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-ml-pipeline, .gemini/skills/audit-ml-pipeline, .github/skills/audit-ml-pipeline and .opencode/skills/audit-ml-pipeline in your project.

What does Audit ML Pipeline need to run?

Going by SKILL.md and its folder, Audit ML Pipeline needs Python for the scripts in its folder and the command-line tools its instructions call (python, git, uv and pip). Our summary lists: Python 3.

Does Audit ML Pipeline access the network?

SKILL.md contains no URLs. Its commands use git, uv and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Audit ML Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit ML Pipeline use?

Audit ML Pipeline is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit ML Pipeline use?

About 9.6k tokens (SKILL.md is roughly 39k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.1k tokens, read only when the agent opens those files.

What are the alternatives to Audit ML Pipeline?

Skills that share tags, products or a category with Audit ML Pipeline: Tech Data Playbook (LeoYeAI/openclaw-master-skills, 2.2k stars), ML Pipeline Workflow (wshobson/agents, 40k stars), Edit (omegaml/omegaml, 108 stars) and ML Pipeline Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit ML Pipeline?

probabl-ai (a GitHub organization) maintains it in probabl-ai/skills, which has 138 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 10, 2026.

Source: probabl-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.