Agent skill

ML Experiment Journal

by Leeroo-AI in Leeroo-AI/superml

Keeps a persistent, append-only journal of ML experiment hypotheses and results across sessions, so no hyperparameter or architecture change runs without being logged first.

Apache-2.0Auto-check passedAI & LLM Engineering

Install ML Experiment Journal

skills CLI
$ npx skills add Leeroo-AI/superml --skill ml-experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Leeroo-AI/superml ml-experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-experiment .claude/skills/ml-experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-experiment
GitHub stars
195
Token cost
~1.3k tokens
SKILL.md length
423 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Keeps a persistent, append-only journal of ML experiment hypotheses and results across sessions, so no hyperparameter or architecture change runs without being logged first.

  • Works in 4 steps: Before Any Experiment — Log the Hypothesis → After the Experiment — Log the Result → Before the Next Iteration — Review History → …
  • Starting a new ML experiment and wanting the hypothesis recorded first
  • SKILL.md covers The Iron Law, File Structure, Phases and After This, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill's rule is that no new experiment runs without its hypothesis logged first. Before changing a hyperparameter, dataset, architecture or training recipe, the agent reads experiments/journal.md to see what has already been tried, then appends a planned entry stating what it expects to happen and why. Once results are in, that same entry is updated to completed, with the actual outcome and how it differed from the expectation.

Before proposing the next experiment, the agent rereads the journal and experiments/lessons.md to check whether the exact approach has been tried and to apply any standing rules, and can only propose a new experiment once it can say why it differs from earlier attempts. After every few experiments, or when a pattern shows up, recurring findings are distilled into dated rules in lessons.md, each with its source such as a user correction or an experiment result.

When your agent uses it

  • Starting a new ML experiment and wanting the hypothesis recorded first
  • Logging the result of a finished training run
  • Checking whether an approach has already been tried before repeating it
  • Distilling a pattern across several experiments into a reusable rule

Example prompts

  • “I'm about to try a lower learning rate; log the hypothesis before we run it.”
  • “The run finished; update the journal entry with the actual metrics and how they differ from what we expected.”
  • “Before we try data augmentation, check the journal for whether we've tried something like this.”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Before Any Experiment — Log the Hypothesis
  2. After the Experiment — Log the Result
  3. Before the Next Iteration — Review History
  4. Periodically — Distill Lessons

What it can do on your machine

Read from SKILL.md and the folder at commit a8d980e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Experiment Journal loads about 1.3k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 423 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Leeroo-AI/superml at commit a8d980e, republished under its Apache-2.0 licence (© Leeroo-AI). 423 words, ~1,293 tokens.

Download SKILL.mdSave it as .claude/skills/ml-experiment/SKILL.md (or your agent's skills folder).
name
ml-experiment
description
Use when starting, logging, or reviewing ML experiments — maintains a persistent experiment journal with hypotheses, results, and learnings across sessions

Experiment Journal

Externalize your experimental reasoning. Every ML project is a sequence of hypotheses tested — this skill makes that sequence visible, persistent, and learnable.

The Iron Law

NO NEW EXPERIMENT WITHOUT LOGGING THE HYPOTHESIS FIRST

If you're about to change a hyperparameter, swap a dataset, try a new architecture, or modify a training recipe — write down what you expect to happen and why BEFORE running it. This is how you learn from experiments instead of just running them.

File Structure

Maintain these files in the project root (create if they don't exist):

experiments/
├── journal.md    — Running experiment log (append-only)
└── lessons.md    — Distilled patterns and rules (curated)

Phases

Phase 1: Before Any Experiment — Log the Hypothesis

Before changing anything or running anything new:

  1. Read experiments/journal.md (if it exists) to see what's been tried
  2. Write a new entry:
markdown
### YYYY-MM-DD HH:MM — [Experiment Name]

**Status**: PLANNED

**Hypothesis**: [What you expect to happen and why]
**Change**: [Exactly what's being modified — one variable at a time]
**Config**:
- key_param_1: old_value → new_value
- key_param_2: value (unchanged)
**Expected outcome**: [Specific metric target or qualitative expectation]
**Baseline**: [Current best metric to beat]

Gate: Entry is written before any code runs. No exceptions.

Phase 2: After the Experiment — Log the Result

Once results are in:

  1. Update the journal entry:
markdown
**Status**: COMPLETED
**Actual outcome**: [What actually happened — metrics, behavior]
**Delta**: [How this compared to expectation — better/worse/different than expected]
**Duration**: [Wall time, GPU hours]
**Learning**: [One sentence — what this taught you]
**Next**: [What to try based on this result]

Gate: Result is logged before starting the next experiment.

Phase 3: Before the Next Iteration — Review History

Before proposing or starting the next experiment:

  1. Read experiments/journal.md — scan recent entries
  2. Check: Has this exact approach been tried before? What happened?
  3. Read experiments/lessons.md — are there rules that apply?
  4. Only then propose the next experiment

Gate: You can articulate why this experiment is different from previous attempts.

Phase 4: Periodically — Distill Lessons

After every 3-5 experiments, or when a pattern emerges:

  1. Review recent journal entries for patterns
  2. Add rules to experiments/lessons.md:
markdown
## Lessons

- [YYYY-MM-DD] [Context]: [Lesson]. Source: [user correction / experiment result / KB finding]
  Example: "2024-03-15 QLoRA: alpha/r ratio matters more than absolute rank for 7B models. Source: experiments showed r=32/alpha=64 outperformed r=64/alpha=64"

## Rules (hard-won)

- NEVER [thing that always fails] because [reason]. Learned: [date]
- ALWAYS [thing that always works] when [condition]. Learned: [date]

Gate: Lessons file has been updated before closing out a series of experiments.

Show full SKILL.md (173 more words)Show less

After This

  • Starting a new experiment? Loop back to Phase 1.
  • Need ideas for what to try next? Invoke ml-iterate — it reads your journal and proposes ranked options.
  • Debugging a failed experiment? Invoke ml-debug — include the journal entry as context.
  • Want to verify a config before running? Invoke ml-verify — catch mistakes before wasting GPU time.

Anti-Patterns

MistakeWhy it happensWhat to do instead
Running without logging"I'll just try this quick thing"Even quick experiments get logged — they compound into knowledge
Changing multiple variables"Let me also bump the LR while I'm at it"One variable per experiment. Otherwise you can't attribute the result.
Not recording the baseline"I'll remember what the old score was"Write the baseline metric in the entry. Memory is unreliable.
Skipping the review step"I know what I tried before"Read the journal. You'll find experiments you forgot about.
Never distilling lessons"The journal has everything"A 200-entry journal is noise. Lessons are signal. Distill regularly.

Examples

Starting a fine-tuning experiment:

User: "Let's try QLoRA with rank 64 instead of 32"
Agent: [Reads experiments/journal.md]
Agent: [Writes new entry with hypothesis: "Higher rank captures more task-specific features, expecting +2% accuracy"]
Agent: [Proceeds with implementation]

After getting results:

User: "Training finished, eval accuracy went from 78% to 81%"
Agent: [Updates journal entry with actual outcome, delta (+3% vs expected +2%), learning]
Agent: [Suggests next experiment based on result]

Before next iteration:

User: "What should we try next?"
Agent: [Reads journal — sees rank 64 worked, rank 16 didn't, data augmentation untested]
Agent: [Invokes ml-iterate with full history context]

© Leeroo-AI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ml-experiment of Leeroo-AI/superml.

Open the folder on GitHubat commit a8d980e

Compare with similar skills

ML Experiment Journal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Experiment Journal compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Experiment Journal this skillLeeroo-AI/superml195—~1.3kAutomated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0
Scaffold Examplecomet-ml/comet-examples174—~1kAutomated safety check: PassNone
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
DGX Spark Memory and Thermal Opswshobson/agents40k—~2kAutomated safety check: PassMIT

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Scaffold Example

    comet-ml/comet-examples

    Scaffold a brand-new Comet example in this repo from the canonical template under templates/integration-example/.

    174 GitHub stars~1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub stars~2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from Leeroo-AI/superml

  • ML Experiment Iteration

    Leeroo-AI/superml

    Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.

    195 GitHub stars~4.8k tokensUpdated 6 mo ago
    Auto-check passed
  • ML Training Run Verifier

    Leeroo-AI/superml

    Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.

    195 GitHub stars~3.8k tokensUpdated 6 mo ago
    Auto-check passed
  • ML Failure Debugger

    Leeroo-AI/superml

    Diagnoses failing ML and AI work, such as OOM, NaN, divergence, crashes, slow throughput, wrong outputs and dependency conflicts, with every claim backed by documentation citations.

    195 GitHub stars~9.5k tokensUpdated 6 mo ago
    Auto-check passed
  • ML Implementation Planner

    Leeroo-AI/superml

    Turns build, implement or design requests for ML pipelines into validated implementation plans grounded in a knowledge base or fetched framework documentation.

    195 GitHub stars~11k tokensUpdated 6 mo ago
    Auto-check passed
  • Using Superml

    Leeroo-AI/superml

    A skill your agent uses when starting any conversation involving ML/AI — establishes how to use Leeroopedia KB tools and workflow skills

    195 GitHub stars~5.9k tokensUpdated 6 mo ago
    Auto-check passed
  • ML Research

    Leeroo-AI/superml

    A skill your agent uses when the user wants to understand an ML/AI topic, compare approaches, or survey framework capabilities — "how does X work?", "compare X vs Y"

    195 GitHub stars~7.3k tokensUpdated 6 mo ago
    Auto-check: warnings

Questions about ML Experiment Journal

What does ML Experiment Journal do?

Keeps a persistent, append-only journal of ML experiment hypotheses and results across sessions, so no hyperparameter or architecture change runs without being logged first. The skill's rule is that no new experiment runs without its hypothesis logged first.md to see what has already been tried, then appends a planned entry stating what it expects to happen and why.

When should I use ML Experiment Journal?

ML Experiment Journal fits situations like: starting a new ML experiment and wanting the hypothesis recorded first; logging the result of a finished training run; checking whether an approach has already been tried before repeating it; distilling a pattern across several experiments into a reusable rule.

How do I install ML Experiment Journal in Claude Code?

Run `npx skills add Leeroo-AI/superml --skill ml-experiment -a claude-code`. Or copy the skill folder (skills/ml-experiment in Leeroo-AI/superml) into .claude/skills/ml-experiment in your project. Claude Code loads it when a task matches its description.

How do I install ML Experiment Journal in Codex?

Run `npx skills add Leeroo-AI/superml --skill ml-experiment -a codex`. Or copy the skill folder (skills/ml-experiment in Leeroo-AI/superml) into .agents/skills/ml-experiment in your project. Codex loads it when a task matches its description.

Can I use ML Experiment Journal in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Leeroo-AI/superml --skill ml-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-experiment, .gemini/skills/ml-experiment, .github/skills/ml-experiment and .opencode/skills/ml-experiment in your project.

What does ML Experiment Journal need to run?

SKILL.md names no scripts, command-line tools or credentials: ML Experiment Journal is instructions for the agent only.

Does ML Experiment Journal access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML Experiment Journal safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Experiment Journal use?

ML Experiment Journal is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Experiment Journal use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Experiment Journal?

Skills that share tags, products or a category with ML Experiment Journal: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 792 stars), Scaffold Example (comet-ml/comet-examples, 174 stars) and PyTorch Lightning Training (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Experiment Journal?

Leeroo-AI (a GitHub organization) maintains it in Leeroo-AI/superml, which has 195 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on March 17, 2026.

Source: Leeroo-AI/superml on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.