Agent skill

Research ML Practice

by probabl-ai in probabl-ai/skills

Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus…

BSD-3-ClauseAuto-check passedData & Analytics

Install Research ML Practice

skills CLI
$ npx skills add probabl-ai/skills --skill research-ml-practice -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install probabl-ai/skills research-ml-practice --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research-ml-practice .claude/skills/research-ml-practice && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
research-ml-practice
GitHub stars
138
Token cost
~1.1k tokens
SKILL.md length
517 words
Files
3 (incl. references)
Skills in repo
23
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus…

  • Works in 5 steps: Intake. Infer modality from dtypes /… → Abstract the problem class before any… → Run the matching search loop… → …
  • Explore-ml-data
  • SKILL.md covers Human-facing prose, Sequence and Stop conditions
  • Calls git and uv

What it does

Research ML Practice is an agent skill from probabl-ai/skills. Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus existing EDA. Trigger when explore-ml-data or build-ml-pipeline load this skill. Not for routine profiling or a single API signature — use api get for symbols. Not a session owner — callers distill and write project files. HOW TO USE: skip intake fields the caller already supplied. Abstract the problem class before searching…

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `evals/evals.json` and `references/search.md`).

It sits in Data & Analytics, covering Data analysis, Machine learning and MLOps. The repository describes itself as: Tabular Data Science Skills for guardrailing AI Agents. The licence is BSD-3-Clause.

When your agent uses it

  • Explore-ml-data
  • Build-ml-pipeline load this skill

Example prompts

  • “/research-ml-practice”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Intake. Infer modality from dtypes / JOURNAL when
  2. Abstract the problem class before any query
  3. Run the matching search loop (references/search.md).
  4. Write scratch/research/ using the matching structure in
  5. Return to the caller: scratch path and a one- or two-sentence

What it can do on your machine

Read from SKILL.md and the folder at commit f273d39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Research ML Practice loads about 1.1k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 237 tokens; SKILL.md has 517 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~237
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from probabl-ai/skills at commit f273d39, republished under its BSD-3-Clause licence (© probabl-ai). 517 words, ~1,145 tokens.

Download SKILL.mdSave it as .claude/skills/research-ml-practice/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
research-ml-practice
description
Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus existing EDA. Trigger when explore-ml-data or build-ml-pipeline load this skill. Not for routine profiling or a single API signature — use `api get` for symbols. Not a session owner — callers distill and write project files. HOW TO USE: skip intake fields the caller already supplied. Abstract the problem class before searching (never the dataset proper name). Survey extra analyses or depth a named concern; search distinct angles, fetch primary sources, follow up if thin, write `scratch/research/`. Return the path plus one or two sentences. Never mutate raw data, pick the final learner, or write `data_analysis.md` / the design note. Chat stays path + those sentences even when a harness asks for a complete answer in the message.

Research ML Practice

Worker skill. Callers own the stage turn and the user-facing summary. Do not git end-turn.

Human-facing prose

Details: setup-workspace references/human_facing_prose.md. Chat stays path + one or two sentences of the finding. Do not narrate the skills framework or the wrapper CLI.

Sequence

  1. Intake. Infer modality from dtypes / JOURNAL when obvious; use a stated domain and stage (data_analysis | model) if the caller named them. Then pick a mode (references/search.md):
    • Caller passed a named concern (not the canned EDA extra-analysis question) → depth. Do not re-ask. Do not run the canned survey first.
    • Caller passed the canned extra-analysis survey or “survey extra analyses” → survey.
    • Neither → if JOURNAL and an EDA report exist, run the canned survey; else AskUserQuestion for a named concern, stating the modality and stage inferred this turn and that the answer only drives a literature search, no code. A file link is an addition, never the context. Do not start a depth search with an empty concern.
  2. Abstract the problem class before any query (references/search.md). JOURNAL and EDA are context, not the answer list and not search keywords for the table’s proper name.
  3. Run the matching search loop (references/search.md). Fetch primary pages. Follow up per promising extra if the first pass is thin or single-sourced.
  4. Write scratch/research/ using the matching structure in references/search.md (survey: survey-<slug>.md; depth: <slug>.md with lanes). There is no templates/ directory in this skill — copy the markdown skeleton from that reference. Gitignored.
  5. Return to the caller: scratch path and a one- or two-sentence finding. Name that candidates are laned (measure / declare / evaluate / confirm). Chat is path + those sentences only — no pasted headings, tables, or “Depth note — …” body. That is the complete user-facing deliverable even when a harness says to put the full answer in chat. If tools cannot search or write, stop there: still only path + sentences (name the intended scratch/research/ path). No hypotheses, diagnostics, planned-query bullets, or template headings in chat (that is the paste). Do not write data_analysis.md, the design note, or data/. The caller asks which extras to add.
Show full SKILL.md (175 more words)Show less

Stop conditions

  • Do not drop, impute, or remove outliers. Do not change the split.
  • Do not pick the final learner or architecture.
  • Do not treat a single blog as ground truth; say when sources disagree. Thresholds need two independent sources.
  • Do not pixi add / uv add / env add. If code needs a library, name add-python-package and return.
  • Do not run api get as a substitute for literature (symbols still go through api get in the caller).
  • Missing skill from a caller → that caller one-line skips.
  • Do not copy Open questions / EDA findings onto the extras list without a source.
  • Do not search the dataset proper name, sklearn.datasets, a Kaggle slug, or “baseline pipeline”. Do not return learners / Pipeline steps as EDA extras.
  • Never answer from memory when search ran. Do not ask the user to go look something up.
  • Do not paste the scratch markdown into chat (no “Depth note —”, no survey body). Path + 1–2 sentences only. If tools cannot search or write, stop after that. A no-tools harness does not license pasting the note.

© probabl-ai, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/research-ml-practice of probabl-ai/skills.

  • SKILL.md
  • evals/evals.json
  • references/search.md

Open the folder on GitHubat commit f273d39

Compare with similar skills

Research ML Practice next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Research ML Practice compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Research ML Practice this skillprobabl-ai/skills138—~1.1kAutomated safety check: PassBSD-3-Clause
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Code Generatorliangdabiao/claude-data-analysis-ultra-main290—~513Automated safety check: PassNone
Data Scientistdavila7/claude-code-templates32k8 repos~2.6kAutomated safety check: PassMIT
Scientific Toolkit SkillzLanqing/codex-claude-academic-skills4.7k—~1.2kAutomated safety check: PassMIT
Data Sciencetravisjneuman/.claude101—~2.3kAutomated safety check: PassMIT

Similar skills

  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Code Generator

    liangdabiao/claude-data-analysis-ultra-main

    Generates production-ready analysis code in Python, R, SQL. An agent skill from liangdabiao/claude-data-analysis-ultra-main.

    290 GitHub stars~513 tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Data Scientist

    davila7/claude-code-templates

    Expert data scientist for advanced analytics, machine learning, and statistical modeling.

    32k GitHub starsUsed in 8 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.7k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Data Science

    travisjneuman/.claude

    Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

    101 GitHub stars~2.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • ML Engineer

    RightNow-AI/openfang

    Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps

    18k GitHub stars~987 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from probabl-ai/skills

All 23 skills in this repo
  • Add Python Package

    probabl-ai/skills

    Add a Python dependency through the project env manager, or ask the user to install it when env.managed is false.

    138 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Build ML Pipeline

    probabl-ai/skills

    Declare the pipeline from data source to predictor as a skrub DataOps graph.

    138 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Evaluate ML Pipeline

    probabl-ai/skills

    Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

    138 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Setup Workspace

    probabl-ai/skills

    Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.

    138 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Audit ML Pipeline

    probabl-ai/skills

    Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

    138 GitHub stars~9.4k tokensUpdated yesterday
    Auto-check passed
  • Explore ML Data

    probabl-ai/skills

    Owns data understanding before any model is designed. An agent skill from probabl-ai/skills.

    138 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed

Questions about Research ML Practice

What does Research ML Practice do?

Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus…. Research ML Practice is an agent skill from probabl-ai/skills. Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus existing EDA.

When should I use Research ML Practice?

Research ML Practice fits situations like: explore-ml-data; build-ml-pipeline load this skill.

How do I install Research ML Practice in Claude Code?

Run `npx skills add probabl-ai/skills --skill research-ml-practice -a claude-code`. Or copy the skill folder (skills/research-ml-practice in probabl-ai/skills) into .claude/skills/research-ml-practice in your project. Claude Code loads it when a task matches its description.

How do I install Research ML Practice in Codex?

Run `npx skills add probabl-ai/skills --skill research-ml-practice -a codex`. Or copy the skill folder (skills/research-ml-practice in probabl-ai/skills) into .agents/skills/research-ml-practice in your project. Codex loads it when a task matches its description.

Can I use Research ML Practice in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add probabl-ai/skills --skill research-ml-practice -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research-ml-practice, .gemini/skills/research-ml-practice, .github/skills/research-ml-practice and .opencode/skills/research-ml-practice in your project.

What does Research ML Practice need to run?

Going by SKILL.md and its folder, Research ML Practice needs the command-line tools its instructions call (git and uv).

Does Research ML Practice access the network?

SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Research ML Practice safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Research ML Practice use?

Research ML Practice is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Research ML Practice use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Research ML Practice?

Skills that share tags, products or a category with Research ML Practice: Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Code Generator (liangdabiao/claude-data-analysis-ultra-main, 290 stars), Data Scientist (davila7/claude-code-templates, 32k stars) and Scientific Toolkit Skill (zLanqing/codex-claude-academic-skills, 4.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Research ML Practice?

probabl-ai (a GitHub organization) maintains it in probabl-ai/skills, which has 138 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 8, 2026.

Source: probabl-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.