LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Keeps a persistent, append-only journal of ML experiment hypotheses and results across sessions, so no hyperparameter or architecture change runs without being logged first.
$ npx skills add Leeroo-AI/superml --skill ml-experiment -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Leeroo-AI/superml ml-experiment --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-experiment .claude/skills/ml-experiment && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ml-experiment" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-experiment into .claude/skills/ml-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-experiment", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Leeroo-AI/superml/tree/main/skills/ml-experimentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Leeroo-AI/superml --skill ml-experiment -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Leeroo-AI/superml ml-experiment --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ml-experiment .agents/skills/ml-experiment && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ml-experiment" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-experiment into .agents/skills/ml-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-experiment", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Leeroo-AI/superml --skill ml-experiment -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Leeroo-AI/superml ml-experiment --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ml-experiment .cursor/skills/ml-experiment && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ml-experiment" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-experiment into .cursor/skills/ml-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-experiment", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Leeroo-AI/superml.git --path skills/ml-experiment--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Leeroo-AI/superml --skill ml-experiment -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Leeroo-AI/superml ml-experiment --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ml-experiment .gemini/skills/ml-experiment && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ml-experiment" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-experiment into .gemini/skills/ml-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-experiment", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Leeroo-AI/superml ml-experimentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Leeroo-AI/superml --skill ml-experiment -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ml-experiment .github/skills/ml-experiment && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ml-experiment" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-experiment into .github/skills/ml-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-experiment", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Leeroo-AI/superml --skill ml-experiment -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Leeroo-AI/superml ml-experiment --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Leeroo-AI/superml.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ml-experiment .opencode/skills/ml-experiment && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ml-experiment" agent skill from https://github.com/Leeroo-AI/superml/tree/main/skills/ml-experiment into .opencode/skills/ml-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-experiment", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ml-experimentKeeps a persistent, append-only journal of ML experiment hypotheses and results across sessions, so no hyperparameter or architecture change runs without being logged first.
The skill's rule is that no new experiment runs without its hypothesis logged first. Before changing a hyperparameter, dataset, architecture or training recipe, the agent reads experiments/journal.md to see what has already been tried, then appends a planned entry stating what it expects to happen and why. Once results are in, that same entry is updated to completed, with the actual outcome and how it differed from the expectation.
Before proposing the next experiment, the agent rereads the journal and experiments/lessons.md to check whether the exact approach has been tried and to apply any standing rules, and can only propose a new experiment once it can say why it differs from earlier attempts. After every few experiments, or when a pattern shows up, recurring findings are distilled into dated rules in lessons.md, each with its source such as a user correction or an experiment result.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a8d980e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ML Experiment Journal loads about 1.3k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 423 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Leeroo-AI/superml at commit a8d980e, republished under its Apache-2.0 licence (© Leeroo-AI). 423 words, ~1,293 tokens.
.claude/skills/ml-experiment/SKILL.md (or your agent's skills folder).Externalize your experimental reasoning. Every ML project is a sequence of hypotheses tested — this skill makes that sequence visible, persistent, and learnable.
NO NEW EXPERIMENT WITHOUT LOGGING THE HYPOTHESIS FIRSTIf you're about to change a hyperparameter, swap a dataset, try a new architecture, or modify a training recipe — write down what you expect to happen and why BEFORE running it. This is how you learn from experiments instead of just running them.
Maintain these files in the project root (create if they don't exist):
experiments/
├── journal.md — Running experiment log (append-only)
└── lessons.md — Distilled patterns and rules (curated)Before changing anything or running anything new:
experiments/journal.md (if it exists) to see what's been tried### YYYY-MM-DD HH:MM — [Experiment Name]
**Status**: PLANNED
**Hypothesis**: [What you expect to happen and why]
**Change**: [Exactly what's being modified — one variable at a time]
**Config**:
- key_param_1: old_value → new_value
- key_param_2: value (unchanged)
**Expected outcome**: [Specific metric target or qualitative expectation]
**Baseline**: [Current best metric to beat]Gate: Entry is written before any code runs. No exceptions.
Once results are in:
**Status**: COMPLETED
**Actual outcome**: [What actually happened — metrics, behavior]
**Delta**: [How this compared to expectation — better/worse/different than expected]
**Duration**: [Wall time, GPU hours]
**Learning**: [One sentence — what this taught you]
**Next**: [What to try based on this result]Gate: Result is logged before starting the next experiment.
Before proposing or starting the next experiment:
experiments/journal.md — scan recent entriesexperiments/lessons.md — are there rules that apply?Gate: You can articulate why this experiment is different from previous attempts.
After every 3-5 experiments, or when a pattern emerges:
experiments/lessons.md:## Lessons
- [YYYY-MM-DD] [Context]: [Lesson]. Source: [user correction / experiment result / KB finding]
Example: "2024-03-15 QLoRA: alpha/r ratio matters more than absolute rank for 7B models. Source: experiments showed r=32/alpha=64 outperformed r=64/alpha=64"
## Rules (hard-won)
- NEVER [thing that always fails] because [reason]. Learned: [date]
- ALWAYS [thing that always works] when [condition]. Learned: [date]Gate: Lessons file has been updated before closing out a series of experiments.
| Mistake | Why it happens | What to do instead |
|---|---|---|
| Running without logging | "I'll just try this quick thing" | Even quick experiments get logged — they compound into knowledge |
| Changing multiple variables | "Let me also bump the LR while I'm at it" | One variable per experiment. Otherwise you can't attribute the result. |
| Not recording the baseline | "I'll remember what the old score was" | Write the baseline metric in the entry. Memory is unreliable. |
| Skipping the review step | "I know what I tried before" | Read the journal. You'll find experiments you forgot about. |
| Never distilling lessons | "The journal has everything" | A 200-entry journal is noise. Lessons are signal. Distill regularly. |
Starting a fine-tuning experiment:
User: "Let's try QLoRA with rank 64 instead of 32"
Agent: [Reads experiments/journal.md]
Agent: [Writes new entry with hypothesis: "Higher rank captures more task-specific features, expecting +2% accuracy"]
Agent: [Proceeds with implementation]After getting results:
User: "Training finished, eval accuracy went from 78% to 81%"
Agent: [Updates journal entry with actual outcome, delta (+3% vs expected +2%), learning]
Agent: [Suggests next experiment based on result]Before next iteration:
User: "What should we try next?"
Agent: [Reads journal — sees rank 64 worked, rank 16 didn't, data augmentation untested]
Agent: [Invokes ml-iterate with full history context]© Leeroo-AI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/ml-experiment of Leeroo-AI/superml.
Open the folder on GitHubat commit a8d980e
ML Experiment Journal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ML Experiment Journal this skillLeeroo-AI/superml | 195 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Scaffold Examplecomet-ml/comet-examples | 174 | — | ~1k | Automated safety check: Pass | None | |
| PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | |
| DGX Spark Memory and Thermal Opswshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
comet-ml/comet-examples
Scaffold a brand-new Comet example in this repo from the canonical template under templates/integration-example/.
Orchestra-Research/AI-Research-SKILLs
Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
Leeroo-AI/superml
Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.
Leeroo-AI/superml
Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.
Leeroo-AI/superml
Diagnoses failing ML and AI work, such as OOM, NaN, divergence, crashes, slow throughput, wrong outputs and dependency conflicts, with every claim backed by documentation citations.
Leeroo-AI/superml
Turns build, implement or design requests for ML pipelines into validated implementation plans grounded in a knowledge base or fetched framework documentation.
Leeroo-AI/superml
A skill your agent uses when starting any conversation involving ML/AI — establishes how to use Leeroopedia KB tools and workflow skills
Leeroo-AI/superml
A skill your agent uses when the user wants to understand an ML/AI topic, compare approaches, or survey framework capabilities — "how does X work?", "compare X vs Y"
Categories
Keeps a persistent, append-only journal of ML experiment hypotheses and results across sessions, so no hyperparameter or architecture change runs without being logged first. The skill's rule is that no new experiment runs without its hypothesis logged first.md to see what has already been tried, then appends a planned entry stating what it expects to happen and why.
ML Experiment Journal fits situations like: starting a new ML experiment and wanting the hypothesis recorded first; logging the result of a finished training run; checking whether an approach has already been tried before repeating it; distilling a pattern across several experiments into a reusable rule.
Run `npx skills add Leeroo-AI/superml --skill ml-experiment -a claude-code`. Or copy the skill folder (skills/ml-experiment in Leeroo-AI/superml) into .claude/skills/ml-experiment in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Leeroo-AI/superml --skill ml-experiment -a codex`. Or copy the skill folder (skills/ml-experiment in Leeroo-AI/superml) into .agents/skills/ml-experiment in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Leeroo-AI/superml --skill ml-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-experiment, .gemini/skills/ml-experiment, .github/skills/ml-experiment and .opencode/skills/ml-experiment in your project.
SKILL.md names no scripts, command-line tools or credentials: ML Experiment Journal is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
ML Experiment Journal is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with ML Experiment Journal: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Quality Flywheel (GoogleCloudPlatform/vertex-ai-samples, 792 stars), Scaffold Example (comet-ml/comet-examples, 174 stars) and PyTorch Lightning Training (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Leeroo-AI (a GitHub organization) maintains it in Leeroo-AI/superml, which has 195 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on March 17, 2026.
Source: Leeroo-AI/superml on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.