Agent skill

Frame ML Problem

by probabl-ai in probabl-ai/skills

Record the problem, the deployment setting, the comparison metric, the baseline, and the fold count in the journal before any model code.

BSD-3-ClauseAuto-check passedData & Analytics

Install Frame ML Problem

skills CLI
$ npx skills add probabl-ai/skills --skill frame-ml-problem -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install probabl-ai/skills frame-ml-problem --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/frame-ml-problem .claude/skills/frame-ml-problem && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
frame-ml-problem
GitHub stars
138
Token cost
~2.4k tokens
SKILL.md length
1,340 words
Files
10 (incl. references)
Skills in repo
23
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Record the problem, the deployment setting, the comparison metric, the baseline, and the fold count in the journal before any model code.

  • Works in 8 steps: Run python -m skore_skills status. If… → Run python -m skore_skills frame show.… → stop — say the JSON reason and stop. → …
  • The user asks which metric to compare on
  • SKILL.md covers Human-facing prose, Procedure and Stop conditions
  • Calls python and git

What it does

Frame ML Problem is an agent skill from probabl-ai/skills. Record the problem, the deployment setting, the comparison metric, the baseline, and the fold count in the journal before any model code. Ask every missing decision in one turn, from frame show. Does not write Python, estimator hyperparameters, or splitter constructors. TRIGGER when the user asks which metric to compare on, how new rows should be split, which baseline to use, or says a problem constraint changed. Not when they ask to run evaluation or CV. HOW TO USE: run python -m skoreskills frame show. Read…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `evals/evals.json`, `references/baseline.md` and `references/deployment.md`).

It sits in Data & Analytics. It works with Python. The repository describes itself as: Tabular Data Science Skills for guardrailing AI Agents. The licence is BSD-3-Clause.

When your agent uses it

  • The user asks which metric to compare on
  • How new rows should be split
  • Which baseline to use
  • Says a problem constraint changed

Example prompts

  • “/frame-ml-problem”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Run python -m skore_skills status. If status.setup.pending
  2. Run python -m skore_skills frame show. When the user is
  3. stop — say the JSON reason and stop.
  4. ask / uncovered — read references/fallback.md only and
  5. ask / missing_keys — read each distinct reference in
  6. When the user named one cell and Status is draft, do not
  7. ask / set — the table was already complete and Status is
  8. proceed — the table is locked, and the user is not changing

What it can do on your machine

Read from SKILL.md and the folder at commit f273d39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Frame ML Problem loads about 2.4k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 253 tokens; SKILL.md has 1,340 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~253
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from probabl-ai/skills at commit f273d39, republished under its BSD-3-Clause licence (© probabl-ai). 1,340 words, ~2,367 tokens.

Download SKILL.mdSave it as .claude/skills/frame-ml-problem/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
frame-ml-problem
description
Record the problem, the deployment setting, the comparison metric, the baseline, and the fold count in the journal before any model code. Ask every missing decision in one turn, from `frame show`. Does not write Python, estimator hyperparameters, or splitter constructors. TRIGGER when the user asks which metric to compare on, how new rows should be split, which baseline to use, or says a problem constraint changed. Not when they ask to run evaluation or CV. HOW TO USE: run `python -m skore_skills frame show`. Read each reference named in `questions` once, ask every key in `missing` in one message, and write every answered cell. If a key is still unanswered, ask it and stop. If the write fills every required cell, set Status to `locked`, say those choices are reused and can be changed by name, then follow `proceed`. Do not ask to confirm. If that command is missing, or the problem is not classification or regression, read `references/fallback.md` and do not invent the closed menu.

Frame ML Problem

Write ## Modeling decisions in journal/JOURNAL.md. The table is the contract. This skill does not declare a learner and does not evaluate one.

Human-facing prose

Details: setup-workspace references/human_facing_prose.md. Journal cells describe this dataset. Do not name the skills framework, the CLI, or a splitter class in the table. Questions use data-science language — not skill ids, G-* names, or the wrapper CLI.

Procedure

  1. Run python -m skore_skills status. If status.setup.pending is non-empty and status.skills.setup-ml-project is true, load setup-ml-project and stop. Do not start this skill. When it returns, continue. Do not load it again on this turn. If that skill is not installed, name the pending pieces in one line and stop. Do not invent git init, scaffold, or env init. If status.setup.env or status.setup.workspace is declined, stop in one line. A declined git or editable is not asked again; continue. If data_analysis is missing and explore-ml-data is installed, AskUserQuestion: explore first (default) or continue from facts the user stated. Explore loads explore-ml-data and stops. Do not invent dataset facts. Do not ask this again once data_analysis is present or skipped.
  2. Run python -m skore_skills frame show. When the user is changing a locked constraint and named one cell, do not add --revise and do not treat proceed as the end of the turn. Run python -m skore_skills frame clear --cell <key>, then frame show with no --revise, and ask the replacement in this turn the way step 5 asks a missing decision. If the message already states the new value, write it and ask only the decisions that are still empty. Do not ask Modify / Keep / Stop. The question names the decision and its current value (the comparison metric is MAE; which metric replaces it). Do not say cell, blank, clear, or reopen, and do not say the new value waits until next turn. Do not load model-ml-pipeline or git-close on this turn. When they are changing a constraint and did not name a cell, AskUserQuestion one pick among the filled decisions (skip n/a) and stop. Do not also ask for a typed answer. Do not --revise, do not frame clear, and do not edit the journal on that turn. JSON action is authoritative. Do not invent a menu. If the command is missing or exits without JSON, read references/fallback.md and follow it. Do not open another reference. Do not guess candidates.
  3. stop — say the JSON reason and stop.
  4. ask / uncovered — read references/fallback.md only and follow it. When that writes the uncovered cells, set Status to locked, say the reuse and change lines, and say there is no splitter translation. Do not ask Lock / Modify / Stop. Do not load build-ml-pipeline. Stop this turn.
  5. ask / missing_keys — read each distinct reference in questions once before asking. Do not open any other file under references/. Ask every key in missing in one message. For a question that has candidates, those are the options. When candidates is absent, ask for the value the reference describes. Draw on three sources, and only what they actually say: the EDA report, free-form text that came with the data if any is present (notes, a dictionary, or a README beside the raw files), and facts the user stated. If none of that text is present, do not invent it. When one of them already states the fact, quote it in the question. Write every Value cell the user answered in this turn. Do not rename Variable cells. Do not stop after the first cell. When the deployment makes other rows inapplicable, set those cells to n/a in the same edit. Horizon, gap, and time role are n/a unless deployment is time. Generalize-to is n/a unless deployment is groups. A fold count of 1 is one train/test split drawn from a single table. When the EDA report, the text shipped with the data, or the user already names a separate training table and test table, offer using that split in the folds question and write predefined if they choose it. Do not offer it otherwise. Do not write prefit in the table. Do not type Revised on; only frame clear writes that date. If any key in missing is still unanswered, set Status to draft once any decision cell is filled, ask those keys, and stop. Do not invent their values. Do not set Status to locked. Do not ask to confirm the table. If the write fills every required cell, set Status to locked in that same edit, not draft. Run python -m skore_skills frame show again in this turn. On proceed, say the reuse and change lines, then follow step 8. On ask / set, write Status locked only, say those lines, run frame show again, and follow step 8. If that frame show still returns missing_keys, the table was not complete: ask those keys and stop, and leave Status draft. Reuse and change lines, quoting JSON context in 2–4 lines: these choices are reused for the rest of the experiment so models stay comparable, and any one of them can be changed by naming it (for example the comparison metric). Do not say "lock" in those lines. Do not AskUserQuestion.
  6. When the user named one cell and Status is draft, do not treat set as accepting the table. Run python -m skore_skills frame clear --cell <key> for that cell and stop. Do not write the new value. Do not name any other cell as cleared. The command's JSON blanked list is the record. Status stays draft. frame clear stamps Revised on; do not type that date. The next frame show asks only keys that are still empty or invalid.
  7. ask / set — the table was already complete and Status is still draft. Write Status locked only. Say the reuse and change lines from step 5. Run frame show again and follow step 8. Do not AskUserQuestion. The user sentence that opened this screen is not a choice. ask / revise is not a user question. Do not present Modify / Keep / Stop. Ignore those choices. Clear the named decision and ask the replacement, as in step 2.
  8. proceed — the table is locked, and the user is not changing a named decision. If they are, step 2 already handled it. If translation is null, say that this lock has no splitter translation. Do not load build-ml-pipeline and do not return to model-ml-pipeline. Stop. If model-ml-pipeline dispatched this turn, return to that coordinator and stop. Do not start build, write a design note, or run the git close from here. If no experiment script exists and History has no running, done, or abandoned model row, and status.skills.model-ml-pipeline is true, load that skill and stop. Do not write a design note here. Do not git end-turn. Otherwise run python -m skore_skills git end-turn --stage implement. If JSON action is invoke, load persist-ml-git only if status.skills.persist-ml-git is true and stop. Otherwise load triage-ml-task only if that skill is installed.
Show full SKILL.md (185 more words)Show less

Stop conditions

  • Do not write Python, a pipeline, a test, or a design note.
  • Do not put a class name or a constructor argument in the journal. TimeSeriesSplit, KFold, GroupKFold, and gap= stay out of the table.
  • Do not re-ask a key that is absent from missing.
  • Do not add an option that is absent from candidates.
  • Do not open a reference the JSON did not name, except references/fallback.md when the command is missing.
  • A locked decision changes by clearing that decision, then filling it like any missing one. A complete fill sets Status to locked again. Do not ask Modify / Keep / Stop.
  • frame clear is the only journal edit that reopens a decision, and only for the decision the user named. It stamps Revised on; do not type that date. Do not rewrite experiments/, audit/, or a report in this skill.
  • After a decision is reopened, do not run an existing experiment script. Say that it still uses the previous splitter and metric. Do not say cell, blank, clear, or reopen. The next build or evaluate rewrites it after the table is locked again.

© probabl-ai, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in skills/frame-ml-problem of probabl-ai/skills.

  • SKILL.md
  • evals/evals.json
  • references/baseline.md
  • references/deployment.md
  • references/fallback.md
  • references/generalize-to.md
  • references/horizon-gap.md
  • references/metric-role.md
  • references/prediction-goal.md
  • references/validation.md

Open the folder on GitHubat commit f273d39

Compare with similar skills

Frame ML Problem next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Frame ML Problem compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Frame ML Problem this skillprobabl-ai/skills138—~2.4kAutomated safety check: PassBSD-3-Clause
Scikit LearnzLanqing/codex-claude-academic-skills4.7k16 repos~3.9kAutomated safety check: PassBSD-3-Clause
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0
Excel and CSV Data Analysisbytedance/deer-flow84k4 repos~2.2kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
Scientific Figure MakingChenLiu-1996/figures4papers8.3k—~557Automated safety check: PassCustom licence

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 16 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    84k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Figure Making

    ChenLiu-1996/figures4papers

    Covers publication-ready matplotlib figures for academic papers, slides, and reports—bars, trends, scatter, heatmaps, and multi-panel layouts—with this…

    8.3k GitHub stars~557 tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Diagnose and fix ModuleNotFoundError in Nuitka standalone binaries caused by missing implicit imports.

    15k GitHub stars~519 tokensUpdated today
    Data & AnalyticsAuto-check passed

More from probabl-ai/skills

All 23 skills in this repo
  • Add Python Package

    probabl-ai/skills

    Add a Python dependency through the project env manager, or ask the user to install it when env.managed is false.

    138 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    138 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Evaluate ML Pipeline

    probabl-ai/skills

    Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

    138 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Setup Workspace

    probabl-ai/skills

    Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.

    138 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Audit ML Pipeline

    probabl-ai/skills

    Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

    138 GitHub stars~9.4k tokensUpdated yesterday
    Auto-check passed
  • Explore ML Data

    probabl-ai/skills

    Owns data understanding before any model is designed. An agent skill from probabl-ai/skills.

    138 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Frame ML Problem

What does Frame ML Problem do?

Record the problem, the deployment setting, the comparison metric, the baseline, and the fold count in the journal before any model code. Frame ML Problem is an agent skill from probabl-ai/skills. Record the problem, the deployment setting, the comparison metric, the baseline, and the fold count in the journal before any model code.

When should I use Frame ML Problem?

Frame ML Problem fits situations like: the user asks which metric to compare on; how new rows should be split; which baseline to use; says a problem constraint changed.

How do I install Frame ML Problem in Claude Code?

Run `npx skills add probabl-ai/skills --skill frame-ml-problem -a claude-code`. Or copy the skill folder (skills/frame-ml-problem in probabl-ai/skills) into .claude/skills/frame-ml-problem in your project. Claude Code loads it when a task matches its description.

How do I install Frame ML Problem in Codex?

Run `npx skills add probabl-ai/skills --skill frame-ml-problem -a codex`. Or copy the skill folder (skills/frame-ml-problem in probabl-ai/skills) into .agents/skills/frame-ml-problem in your project. Codex loads it when a task matches its description.

Can I use Frame ML Problem in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add probabl-ai/skills --skill frame-ml-problem -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/frame-ml-problem, .gemini/skills/frame-ml-problem, .github/skills/frame-ml-problem and .opencode/skills/frame-ml-problem in your project.

What does Frame ML Problem need to run?

Going by SKILL.md and its folder, Frame ML Problem needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Frame ML Problem access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Frame ML Problem safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Frame ML Problem use?

Frame ML Problem is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Frame ML Problem use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Frame ML Problem?

Skills that share tags, products or a category with Frame ML Problem: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.7k stars), TimesFM Forecasting (google-research/timesfm, 34k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 84k stars) and Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Frame ML Problem?

probabl-ai (a GitHub organization) maintains it in probabl-ai/skills, which has 138 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 8, 2026.

Source: probabl-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.