Agent skill

Plan ML Experiment

by PKU-YuanGroup in PKU-YuanGroup/OpenAI4S

Plan reproducible machine-learning experiments with leakage-safe random, grouped, or chronological splits; deterministic configuration fingerprints; dataset checksums; seeds, baselines, ablations…

MITAuto-check passedData & Analytics

Install Plan ML Experiment

skills CLI
$ npx skills add PKU-YuanGroup/OpenAI4S --skill plan-ml-experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PKU-YuanGroup/OpenAI4S plan-ml-experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/plan-ml-experiment .claude/skills/plan-ml-experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
plan-ml-experiment
GitHub stars
622
Token cost
~634 tokens
SKILL.md length
238 words
Files
4
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Plan reproducible machine-learning experiments with leakage-safe random, grouped, or chronological splits; deterministic configuration fingerprints; dataset checksums; seeds, baselines, ablations…

  • Works in 6 steps: Write the hypothesis, intervention,… → Choose the unit that must remain… → Reserve the test set for the final… → …
  • Tasks that involve Machine learning
  • SKILL.md covers Planning sequence, Deterministic helpers and Minimum artifact set
  • Runs Python scripts from its folder

What it does

Plan ML Experiment is an agent skill from PKU-YuanGroup/OpenAI4S. Plan reproducible machine-learning experiments with leakage-safe random, grouped, or chronological splits; deterministic configuration fingerprints; dataset checksums; seeds, baselines, ablations, and artifact manifests.

Its SKILL.md is about 630 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `README.md`, `README_zh.md` and `kernel.py`).

It sits in Data & Analytics, covering Machine learning. The repository describes itself as: Open-source AI agent for scientific research. Analyze data in Python/R with Claude, GPT, Gemini, and more. The licence is MIT.

When your agent uses it

  • Tasks that involve Machine learning

Example prompts

  • “/plan-ml-experiment”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Write the hypothesis, intervention, baseline, primary metric, and decision
  2. Choose the unit that must remain independent. Use a grouped split for
  3. Reserve the test set for the final comparison. Fit preprocessing and choose
  4. Fix seeds, environments, input versions, and the exact configuration.
  5. Run the baseline first, then one-factor ablations and the planned model.
  6. Save per-example predictions and a manifest before drawing conclusions.

What it can do on your machine

Read from SKILL.md and the folder at commit 4a72e87. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Plan ML Experiment loads about 634 tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 238 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~634

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PKU-YuanGroup/OpenAI4S at commit 4a72e87, republished under its MIT licence (© PKU-YuanGroup). 238 words, ~634 tokens.

Download SKILL.mdSave it as .claude/skills/plan-ml-experiment/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
plan-ml-experiment
description
Plan reproducible machine-learning experiments with leakage-safe random, grouped, or chronological splits; deterministic configuration fingerprints; dataset checksums; seeds, baselines, ablations, and artifact manifests.
origin
openai4s
category
reproducibility

Plan an ML experiment

Use this skill before training begins. A reproducible experiment is a falsifiable question plus immutable inputs, a leakage-safe evaluation boundary, a declared metric, and enough recorded state to rerun the comparison.

Planning sequence

  1. Write the hypothesis, intervention, baseline, primary metric, and decision rule before inspecting test performance.
  2. Choose the unit that must remain independent. Use a grouped split for patients, molecules/scaffolds, sites, documents, or repeated measures; use a chronological split when deployment predicts the future.
  3. Reserve the test set for the final comparison. Fit preprocessing and choose hyperparameters using training/validation data only.
  4. Fix seeds, environments, input versions, and the exact configuration.
  5. Run the baseline first, then one-factor ablations and the planned model.
  6. Save per-example predictions and a manifest before drawing conclusions.

Deterministic helpers

python
from importlib import import_module

plan = import_module("plan-ml-experiment.kernel")
splits = plan.grouped_split(patient_ids, seed=42)
manifest = plan.experiment_manifest(
    config,
    data_paths=["data/cohort.csv"],
    seeds=[42, 43, 44],
    code_revision="<git commit>",
)

Use random_split(size, ...) only when rows are genuinely independent. grouped_split(groups, ...) keeps each group in exactly one partition. chronological_split(timestamps, ...) orders observations without shuffling. All return original row indices under train, validation, and test.

Minimum artifact set

  • frozen config and config_fingerprint;
  • source artifact/version plus SHA-256 checksums;
  • code revision, runtime/environment versions, and random seeds;
  • split indices or stable sample IDs and the grouping/time policy;
  • baseline, ablation, and final per-example predictions;
  • aggregate metrics with uncertainty and a failure analysis.

Determinism does not prove validity. Hardware kernels may remain nondeterministic, and repeating one biased split does not repair leakage or dataset shift. Report deviations from the plan rather than overwriting it.

© PKU-YuanGroup, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in skills/plan-ml-experiment of PKU-YuanGroup/OpenAI4S.

  • SKILL.md
  • README.md
  • README_zh.md
  • kernel.py

Open the folder on GitHubat commit 4a72e87

Compare with similar skills

Plan ML Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Plan ML Experiment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Plan ML Experiment this skillPKU-YuanGroup/OpenAI4S622—~634Automated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.7k16 repos~3.9kAutomated safety check: PassBSD-3-Clause
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1255 repos~1.4kAutomated safety check: PassMIT
Geomlitalo-goncalves/geoML109—~4.9kAutomated safety check: PassGPL-3.0
QuantMind Training Config Generatorqusong0627/QuantMind1.7k—~1.5kAutomated safety check: PassAGPL-3.0

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 16 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 5 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Geoml

    italo-goncalves/geoML

    Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…

    109 GitHub stars~4.9k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.

    1.7k GitHub stars~1.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Retention Analysis

    liangdabiao/claude-data-analysis-ultra-main

    Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.

    290 GitHub stars~1.3k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from PKU-YuanGroup/OpenAI4S

All 17 skills in this repo
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    622 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Bioprobench

    PKU-YuanGroup/OpenAI4S

    Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

    622 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Reaction Atom Mapping

    PKU-YuanGroup/OpenAI4S

    Map atoms and changed bonds for a complete reaction with RXNMapper.

    622 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Reaction Forward Prediction

    PKU-YuanGroup/OpenAI4S

    Predict ranked products from reactants and reagents with ReactionT5v2-forward; use for outcome prediction or round-trip recovery.

    622 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Reaction Yield Estimation

    PKU-YuanGroup/OpenAI4S

    Estimate yield for a fully specified reactant/reagent/product record with ReactionT5v2-yield.

    622 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Rfdiffusion

    PKU-YuanGroup/OpenAI4S

    Generate de novo protein backbones with RFdiffusion for protein-target binders, hotspot-conditioned interfaces, motif scaffolding, partial diffusion, or symmetric assemblies.

    622 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Questions about Plan ML Experiment

What does Plan ML Experiment do?

Plan reproducible machine-learning experiments with leakage-safe random, grouped, or chronological splits; deterministic configuration fingerprints; dataset checksums; seeds, baselines, ablations…. Plan ML Experiment is an agent skill from PKU-YuanGroup/OpenAI4S. Plan reproducible machine-learning experiments with leakage-safe random, grouped, or chronological splits; deterministic configuration fingerprints; dataset checksums; seeds, baselines, ablations, and artifact manifests.

When should I use Plan ML Experiment?

Plan ML Experiment fits situations like: tasks that involve Machine learning.

How do I install Plan ML Experiment in Claude Code?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill plan-ml-experiment -a claude-code`. Or copy the skill folder (skills/plan-ml-experiment in PKU-YuanGroup/OpenAI4S) into .claude/skills/plan-ml-experiment in your project. Claude Code loads it when a task matches its description.

How do I install Plan ML Experiment in Codex?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill plan-ml-experiment -a codex`. Or copy the skill folder (skills/plan-ml-experiment in PKU-YuanGroup/OpenAI4S) into .agents/skills/plan-ml-experiment in your project. Codex loads it when a task matches its description.

Can I use Plan ML Experiment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PKU-YuanGroup/OpenAI4S --skill plan-ml-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/plan-ml-experiment, .gemini/skills/plan-ml-experiment, .github/skills/plan-ml-experiment and .opencode/skills/plan-ml-experiment in your project.

What does Plan ML Experiment need to run?

Going by SKILL.md and its folder, Plan ML Experiment needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Plan ML Experiment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Plan ML Experiment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Plan ML Experiment use?

Plan ML Experiment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Plan ML Experiment use?

About 634 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Plan ML Experiment?

Skills that share tags, products or a category with Plan ML Experiment: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.7k stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars) and Geoml (italo-goncalves/geoML, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Plan ML Experiment?

PKU-YuanGroup (a GitHub organization) maintains it in PKU-YuanGroup/OpenAI4S, which has 622 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: PKU-YuanGroup/OpenAI4S on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.