Agent skill

Manage ML Backlog

by probabl-ai in probabl-ai/skills

Canonical backlog loop step. An agent skill from probabl-ai/skills.

BSD-3-ClauseAuto-check passedDevOps & Cloud

Install Manage ML Backlog

skills CLI
$ npx skills add probabl-ai/skills --skill manage-ml-backlog -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install probabl-ai/skills manage-ml-backlog --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/manage-ml-backlog .claude/skills/manage-ml-backlog && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
manage-ml-backlog
GitHub stars
138
Token cost
~3.4k tokens
SKILL.md length
1,932 words
Files
2
Skills in repo
23
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Canonical backlog loop step. An agent skill from probabl-ai/skills.

  • Works in 3 steps: When the user has not named a B, present… → Turn the selected row's Item + Source… → Return that proposal to…
  • The user asks what to try next
  • SKILL.md covers Human-facing prose, Model-entry selection mode, Record-outcome mode and Procedure, plus 2 more sections
  • Calls python, git and uv

What it does

Manage ML Backlog is an agent skill from probabl-ai/skills. Canonical backlog loop step. Record an experiment outcome in History and triage idea files. Summarize open, discarded, and aside ideas in the JOURNAL Ideas table. Move a row into Backlog when it is promoted, and keep the idea file. Also supports the model-entry selection mode: show real B<N rows supplied by the deterministic CLI and consume one into a proposal. Trigger after audit, when a run finishes, when the user asks what to try next or to triage idea files, or when model-ml-pipeline routes its Backlog choice…

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in DevOps & Cloud, covering MLOps. The repository describes itself as: Tabular Data Science Skills for guardrailing AI Agents. The licence is BSD-3-Clause.

When your agent uses it

  • The user asks what to try next
  • Triage idea files
  • Model-ml-pipeline routes its Backlog choice here

Example prompts

  • “/manage-ml-backlog”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. When the user has not named a B, present exactly those
  2. Turn the selected row's Item + Source into a Proposal. Ask only
  3. Return that proposal to model-ml-pipeline. Do not ask to

What it can do on your machine

Read from SKILL.md and the folder at commit f273d39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Manage ML Backlog loads about 3.4k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 1,932 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from probabl-ai/skills at commit f273d39, republished under its BSD-3-Clause licence (© probabl-ai). 1,932 words, ~3,367 tokens.

Download SKILL.mdSave it as .claude/skills/manage-ml-backlog/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
manage-ml-backlog
description
Canonical backlog loop step. Record an experiment outcome in History and triage idea files. Summarize open, discarded, and aside ideas in the JOURNAL Ideas table. Move a row into Backlog when it is promoted, and keep the idea file. Also supports the model-entry selection mode: show real B<N> rows supplied by the deterministic CLI and consume one into a proposal. Trigger after audit, when a run finishes, when the user asks what to try next or to triage idea files, or when model-ml-pipeline routes its Backlog choice here. This is cadence, not a methodology owner.

Manage ML Backlog

Replace iterate-as-cadence. Do not own setup, exploratory data analysis, build, smoke, evaluate, or audit methodology.

Human-facing prose

Details: setup-workspace references/human_facing_prose.md. JOURNAL rows, design-note Status / Results, and # comments describe this experiment's outcome — not the skills framework, the CLI, or the command that produced an output. Questions and replies use the same data-science language — not skill ids, G-* names, or the wrapper CLI. <!-- results-embed: … --> is a site marker. Authoring hints stay in this skill. style is ruff only.

The CLI writes JOURNAL with sections in order: Status, Data understanding, Modeling decisions, History, Ideas, Backlog. History, Ideas, and Backlog start as header-only tables. Column contracts stay here (Stem, Intent, Status, Headline result, Report, Design note; Question, Status, Experiment, Source; #, Item, Source). Do not put those contracts back into HTML comments in the file.

Model-entry selection mode

When model-ml-pipeline calls with the backlog array from python -m skore_skills model choices:

  1. When the user has not named a B<N>, present exactly those rows in their returned order and AskUserQuestion for one pick. One row, or a message that already names a B<N>, is the pick: do not ask again. Carry each row's Item and Source as the option context, and say in 2–4 lines what the pick authorizes (a design note to approve) and what it does not (no model code yet). A file link is an addition, never the context. Do not rescan into a different menu and do not add an idea.
  2. Turn the selected row's Item + Source into a Proposal. Ask only for missing shaping facts; do not invent a Method from a one-line item. Do not ask Yes / No on the proposal.
  3. Return that proposal to model-ml-pipeline. Do not ask to confirm it. After the model stage creates and populates the design note, remove only the selected Backlog row and add the planned History row. Preserve every other stable B<N> index.

This mode does not require a report/audit digest and does not run the outcome-recording procedure below. Empty Backlog is a routing error: return to model-ml-pipeline; do not fabricate B1.

Record-outcome mode

When model-ml-pipeline, evaluate-ml-pipeline, or audit-ml-pipeline calls at end of turn with the normalized G-REPORT-LOCATOR, optional headline, and G-AUDIT-FINDING. This is the only path that records an outcome without a full backlog turn. Audit may have been skipped; the locator remains required and the finding becomes n/a — audit not run. If the caller omitted the locator, run python -m skore_skills loop locator --stem <stem> and paste JSON locator. If it omitted the finding, run python -m skore_skills audit finding --stem <stem> (stop → n/a — audit not run / n/a — audit digest unavailable). Do not rephrase either string.

Run Procedure steps 1-3 and nothing else:

  1. Step 1 — python -m skore_skills status; require an approved stem.
  2. Step 2 — read journal/JOURNAL.md; scaffold the index if it is missing.
  3. Step 3 — update the matching History row and design-note Status block from the digest or user-supplied headline when available. Paste G-REPORT-LOCATOR and G-AUDIT-FINDING verbatim into their separate Status lines. Do not change State or Approved by user on. Also refresh the journal/JOURNAL.md Status rows Last experiment and Last result. Insert or replace ## Results between Status and Notebooks from digest text, not HTML. ### Metrics is a heading only.

Then return to the caller. Do not read journal/ideas/, do not edit the Ideas table, and do not open the idea-triage menu — the caller did not ask what to try next.

Do not dispatch audit-ml-pipeline in this mode; the digest is already in hand and dispatching would bounce back here. Do not run this skill's End of turn either: the caller owns convert / site / git end-turn and the User-facing close. This mode writes History; it does not replace the caller's chat close.

The Procedure guards still bind. Never mark done while smoke is red, and never invent a metric — no digest and no user-supplied value means use n/a, not a guess. Never construct a missing backend URL. Record n/a — backend did not expose a locator in both markdown destinations when the digest has no authoritative locator.

Procedure

  1. Run python -m skore_skills status. When recording a done outcome, run python -m skore_skills design consent --stem <stem>. ask / stop → do not mark done. Also require green smoke evidence and a normalized report locator. An audit digest is optional. G-AUDIT-FINDING is required as one of: the value returned by audit, n/a — audit not run, or n/a — audit digest unavailable.
  2. Read journal/JOURNAL.md History and Backlog. If the index is missing, run python -m skore_skills scaffold --journal. Do not write or paste the file. The CLI writes sections in order: Status, Data understanding, Modeling decisions, History, Ideas, and Backlog. If that command cannot run this turn, name it and stop. After the file exists, edit the existing History and Backlog tables (columns: Stem, Intent, Status, Headline result, Report, Design note; and #, Item, Source). Ideas columns are Question, Status, Experiment, Source; triage owns that table. A planned History row uses n/a in Report. Stable B<N> indices. Do not renumber on removal. Record-outcome does not edit Ideas.
  3. If recording a run: copy the headline metric from the audit digest or the user's value. Do not invent numbers. With no digest and no user headline, skip the headline in one line and leave the History status unchanged. Do not write done with headline n/a. Update the matching History row (planned → done only if smoke passed and a headline result exists). Headline metric remains the source for History and Last result; never substitute G-AUDIT-FINDING for performance. Copy the digest's persisted-report locator into the History Report cell and the design note's Persisted report Status line. Paste that string verbatim; do not paraphrase it as "normalized" or rewrite the Hub URL. If the digest has none, write n/a — backend did not expose a locator in both places; do not derive or guess a URL. Copy G-AUDIT-FINDING verbatim into the design note's Audit findings line. Audit skipped → n/a — audit not run; missing/errored digest → n/a — audit digest unavailable. The Status update copies Headline result, Persisted report, and Audit findings. Do not change State or Approved by user on. Then insert or replace ## Results in the design note, between ## Status and ## Notebooks. Summarize from the audit digest — its cell outputs carry repr(report), ## Checks summary, and ## Metrics summary as text. With no audit this turn, fall back to scratch/results/<stem>/report.txt, which evaluate writes. Write ### Report overview from the report text, then ### Checks when that section exists. ### Metrics is a heading only when ## Metrics summary exists: do not transcribe the metric table or its values. The site embeds scratch/results/<stem>/metrics.html under that heading. History and Last result still get the single headline metric. Evaluation-only (audit skipped): Report overview only — do not invent Checks or Metrics subsections. After Metrics, add one ### subsection per extra Display cell the audit appended, using a human title and a <!-- results-embed: <slug> --> comment with the accessor name as <slug> so site build can inject the viewer. Report overview, Checks, and each extra subsection are 2–4 sentences from that cell's output. Do not copy G-AUDIT-FINDING, do not parse *.html, and do not paste iframes (site build injects those). Do not add <!-- results-embed: audit -->. The audit notebook viewer is audit/<stem>.nb.html, which site build places under ## Notebooks. If no subsection has a source, skip the Results section.
  4. Idea triage, separate from record-outcome. Read journal/ideas/*.md and the Ideas table. If ## Ideas is missing, insert that header-only table between History and Backlog. Do not rewrite the other sections. A file's Triage line is open, promoted, discarded, or aside. A missing line is open. An open file whose Question and Source already match one Backlog row is set to promoted without asking, its Ideas row is removed, and no duplicate row is appended. Match both fields so two user ideas do not collapse. An open file with no Ideas row gets one: Question as plain text, Status open, Experiment from the file, Source copied verbatim. Ask the other open files: promote / discard / set aside. Write that value on the Triage line and keep the file. Promote removes that Ideas row and appends a stable B<N> row (Item from Question, Source copied verbatim) and sets promoted. Discard sets discarded on the file and the Ideas row and adds no Backlog row. Set aside sets aside on the file and the Ideas row and adds no Backlog row. A promoted idea is absent from Ideas. promoted, discarded, and aside leave the default queue; a later pass does not ask about them. Do not create a design note here. Do not delete an idea file. Changing Triage does not remove or renumber a Backlog row. An empty folder is a one-line skip: there are no idea files to triage, and it does not fabricate B1. When no file is open and tagged files remain, say in one line how many are promoted, discarded, and set aside, and offer to revisit. Do not retag until the user picks a file. On revisit, the same three choices apply. Promoting then appends B<N> only when that Question and Source are not already a row, and removes the Ideas row. The only other follow-up is offering to shape an idea or search the literature when those skills are installed. Do not load either skill, and do not start a search or a shaping menu, until the user picks one. Missing skill → one-line skip; do not invent that skill's search or shaping steps. After the user picks, that skill writes the idea file and returns here; triage the new file in this same mode. When the user picks an existing B<N> to draft, return that row to model-ml-pipeline, which can create its design-note shell with python -m skore_skills scaffold --journal --stem <NN_short_name>. Do not draft that template in this backlog turn.
Show full SKILL.md (311 more words)Show less

Stop conditions

  • Do not design or implement the next experiment in this turn.
  • Do not dispatch setup or audit by skill id. Returning a selected row or proposal to model-ml-pipeline is required. Do not ask Yes / No on that proposal.
  • In model-entry selection mode, do not invent a Backlog row or remove it before the paired design note exists.
  • Do not invent metrics.
  • Do not derive, shorten, or merge G-AUDIT-FINDING with the headline metric. Copy each into its owned field.
  • Do not change design-note State or Approved by user on. Record-outcome copies Headline result, Persisted report, and Audit findings.
  • Do not parse scratch/results/<stem>/*.html when writing ## Results. Summarize Report overview and Checks from the digest, or from report.txt on the evaluation-only path. ### Metrics is a heading only.
  • Do not paste a JOURNAL.md body or recreate the index from memory.
  • Do not mark done while smoke is red.
  • Do not delete an idea file. Triage writes its Triage line and moves or updates its Ideas row.
  • Design approval is owned by model-ml-pipeline; this skill only returns a proposal or selected Backlog row.

End of turn

If policy.site is true, export-ml-site is installed, run python -m skore_skills site build. Do not run notebook convert. Skip in one line otherwise. If site build errors with mkdocs-material is required, load add-python-package for mkdocs-material (agent) and build once more. Do not pixi add / uv add. If that skill is missing, or the retry still fails, name the error in one line. Name a build error; do not fail the backlog turn.

Run python -m skore_skills git end-turn --stage backlog. If JSON action is invoke, load persist-ml-git only if status.skills.persist-ml-git is true and stop; that skill returns to triage. If persist is missing, name the pending staged paths and stop. Otherwise load triage-ml-task only if status.skills.triage-ml-task is true; else stop. Do not run git commit in this skill.

© probabl-ai, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/manage-ml-backlog of probabl-ai/skills.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit f273d39

Compare with similar skills

Manage ML Backlog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Manage ML Backlog compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Manage ML Backlog this skillprobabl-ai/skills138—~3.4kAutomated safety check: PassBSD-3-Clause
ML Pipeline Workflowwshobson/agents40k12 repos~1.8kAutomated safety check: PassMIT
Editomegaml/omegaml107—~206Automated safety check: PassApache-2.0
ML Pipeline ExpertJeffallan/claude-skills12k—~1.9kAutomated safety check: PassMIT
Agent Platform Model Registrygoogle/skills21k—~2.6kAutomated safety check: PassApache-2.0
Machine Learning Ops ML Pipelineaiskillstore/marketplace4307 repos~2.6kAutomated safety check: PassNone

Similar skills

  • ML Pipeline Workflow

    wshobson/agents

    Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.

    40k GitHub starsUsed in 12 repos~1.8k tokens
    DevOps & CloudAuto-check passed
  • Edit

    omegaml/omegaml

    how to use the edit command properly

    107 GitHub stars~206 tokensUpdated today
    DevOps & CloudAuto-check passed
  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub stars~1.9k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed
  • Official

    Agent Platform Model Registry Management. An agent skill from google/skills.

    21k GitHub stars~2.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Machine Learning Ops ML Pipeline

    aiskillstore/marketplace

    Design and implement a complete ML pipeline for: $ARGUMENTS. An agent skill from aiskillstore/marketplace.

    430 GitHub starsUsed in 7 repos~2.6k tokens
    DevOps & CloudAuto-check passed
  • ML Pipeline Automation

    secondsky/claude-skills

    Automate ML workflows with Airflow, Kubeflow, MLflow. An agent skill from secondsky/claude-skills.

    227 GitHub stars~3.2k tokensUpdated 11 days ago
    DevOps & CloudAuto-check passed

More from probabl-ai/skills

All 23 skills in this repo
  • Add Python Package

    probabl-ai/skills

    Add a Python dependency through the project env manager, or ask the user to install it when env.managed is false.

    138 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    138 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Evaluate ML Pipeline

    probabl-ai/skills

    Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

    138 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Setup Workspace

    probabl-ai/skills

    Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.

    138 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Audit ML Pipeline

    probabl-ai/skills

    Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

    138 GitHub stars~9.4k tokensUpdated yesterday
    Auto-check passed
  • Explore ML Data

    probabl-ai/skills

    Owns data understanding before any model is designed. An agent skill from probabl-ai/skills.

    138 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed

Questions about Manage ML Backlog

What does Manage ML Backlog do?

Canonical backlog loop step. An agent skill from probabl-ai/skills. Manage ML Backlog is an agent skill from probabl-ai/skills. Canonical backlog loop step.

When should I use Manage ML Backlog?

Manage ML Backlog fits situations like: the user asks what to try next; triage idea files; model-ml-pipeline routes its Backlog choice here.

How do I install Manage ML Backlog in Claude Code?

Run `npx skills add probabl-ai/skills --skill manage-ml-backlog -a claude-code`. Or copy the skill folder (skills/manage-ml-backlog in probabl-ai/skills) into .claude/skills/manage-ml-backlog in your project. Claude Code loads it when a task matches its description.

How do I install Manage ML Backlog in Codex?

Run `npx skills add probabl-ai/skills --skill manage-ml-backlog -a codex`. Or copy the skill folder (skills/manage-ml-backlog in probabl-ai/skills) into .agents/skills/manage-ml-backlog in your project. Codex loads it when a task matches its description.

Can I use Manage ML Backlog in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add probabl-ai/skills --skill manage-ml-backlog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/manage-ml-backlog, .gemini/skills/manage-ml-backlog, .github/skills/manage-ml-backlog and .opencode/skills/manage-ml-backlog in your project.

What does Manage ML Backlog need to run?

Going by SKILL.md and its folder, Manage ML Backlog needs the command-line tools its instructions call (python, git and uv). Our summary lists: Python 3.

Does Manage ML Backlog access the network?

SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Manage ML Backlog safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Manage ML Backlog use?

Manage ML Backlog is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Manage ML Backlog use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Manage ML Backlog?

Skills that share tags, products or a category with Manage ML Backlog: ML Pipeline Workflow (wshobson/agents, 40k stars), Edit (omegaml/omegaml, 107 stars), ML Pipeline Expert (Jeffallan/claude-skills, 12k stars) and Agent Platform Model Registry (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Manage ML Backlog?

probabl-ai (a GitHub organization) maintains it in probabl-ai/skills, which has 138 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 8, 2026.

Source: probabl-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.