Agent skill

Experimental Design

by aiming-lab in aiming-lab/AutoResearchClaw

Best practices for designing reproducible ML experiments. An agent skill from aiming-lab/AutoResearchClaw.

MITAuto-check passedResearch & Science

Install Experimental Design

skills CLI
$ npx skills add aiming-lab/AutoResearchClaw --skill experimental-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aiming-lab/AutoResearchClaw experimental-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/researchclaw/skills/builtin/experiment/experimental-design .claude/skills/experimental-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experimental-design
GitHub stars
15k
Token cost
~286 tokens
SKILL.md length
97 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
MIT

At a glance

Best practices for designing reproducible ML experiments. An agent skill from aiming-lab/AutoResearchClaw.

  • Works in 7 steps: ALWAYS include meaningful baselines (not… → Use MULTIPLE random seeds (minimum 3,… → Report mean +/- std across seeds → …
  • Planning ablations
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Controlled experiments

What it does

Experimental Design is an agent skill from aiming-lab/AutoResearchClaw. Best practices for designing reproducible ML experiments. Use when planning ablations, baselines, or controlled experiments.

Its SKILL.md is about 290 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Experimental design. The repository describes itself as: Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞. The licence is MIT.

When your agent uses it

  • Planning ablations
  • Controlled experiments

Example prompts

  • “/experimental-design”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. ALWAYS include meaningful baselines (not just random)
  2. Use MULTIPLE random seeds (minimum 3, ideally 5)
  3. Report mean +/- std across seeds
  4. Design ablations that isolate EACH key component
  5. Control variables: change only ONE thing per comparison
  6. Use standard splits (train/val/test) — never test on training data
  7. Report wall-clock time and memory usage alongside accuracy

What it can do on your machine

Read from SKILL.md and the folder at commit be4ba47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experimental Design loads about 286 tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 97 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~286

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aiming-lab/AutoResearchClaw at commit be4ba47, republished under its MIT licence (© aiming-lab). 97 words, ~286 tokens.

Download SKILL.mdSave it as .claude/skills/experimental-design/SKILL.md (or your agent's skills folder).
name
experimental-design
description
Best practices for designing reproducible ML experiments. Use when planning ablations, baselines, or controlled experiments.
metadata.category
experiment
metadata.trigger-keywords
experiment,ablation,baseline,control,hypothesis,reproducib
metadata.applicable-stages
9,10,12
metadata.priority
2
metadata.version
1.0
metadata.author
researchclaw
metadata.references
Bouthillier et al., Accounting for Variance in ML Benchmarks, MLSys 2021

Experimental Design Best Practice

  1. ALWAYS include meaningful baselines (not just random):
    • At least one classical method baseline
    • At least one recent SOTA method baseline
    • A simple-but-strong baseline (e.g., linear probe, k-NN)
  2. Use MULTIPLE random seeds (minimum 3, ideally 5)
  3. Report mean +/- std across seeds
  4. Design ablations that isolate EACH key component:
    • Remove one component at a time
    • Each ablation must be meaningfully different from baseline
  5. Control variables: change only ONE thing per comparison
  6. Use standard splits (train/val/test) — never test on training data
  7. Report wall-clock time and memory usage alongside accuracy

© aiming-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in researchclaw/skills/builtin/experiment/experimental-design of aiming-lab/AutoResearchClaw.

Open the folder on GitHubat commit be4ba47

Compare with similar skills

Experimental Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experimental Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experimental Design this skillaiming-lab/AutoResearchClaw15k—~286Automated safety check: PassMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k23 repos~5.9kAutomated safety check: NotesMIT
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1287 repos~2.3kAutomated safety check: NotesNone
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.5k—~2.8kAutomated safety check: PassCC-BY-4.0
Research Refine PipelinezjYao36/Auto-Research-Refine1286 repos~1.4kAutomated safety check: NotesNone
Experimental DesignOleafly/Oleafly2064 repos~3.5kAutomated safety check: NotesMIT

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 23 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 7 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.5k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 6 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Experimental Design

    Oleafly/Oleafly

    Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so results are interpretable.

    206 GitHub starsUsed in 4 repos~3.5k tokens
    Research & ScienceAuto-check: notes
  • Research Refine

    zjYao36/Auto-Research-Refine

    Turns a vague research direction into a focused, problem-anchored method plan through up to five review rounds with a second model.

    128 GitHub starsUsed in 7 repos~6.9k tokens
    Research & ScienceAuto-check: notes

More from aiming-lab/AutoResearchClaw

All 34 skills in this repo
  • A-Evolve Agent Improvement

    aiming-lab/AutoResearchClaw

    Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.

    15k GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Auto-check passed
  • Qiskit 2.x Quantum ML Reference

    aiming-lab/AutoResearchClaw

    Reference patterns for writing qiskit 2.x code for variational quantum machine learning: feature maps, VQC training, VQE for chemistry, MPS circuits and noise models.

    15k GitHub stars~4.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Genome-Scale Metabolic Model Builder

    aiming-lab/AutoResearchClaw

    Builds or loads a genome-scale metabolic model in COBRApy, sets its growth medium and objective, and exports it as a validated JSON file for flux analysis.

    15k GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Experimental Design

What does Experimental Design do?

Best practices for designing reproducible ML experiments. An agent skill from aiming-lab/AutoResearchClaw. Experimental Design is an agent skill from aiming-lab/AutoResearchClaw. Best practices for designing reproducible ML experiments.

When should I use Experimental Design?

Experimental Design fits situations like: planning ablations; controlled experiments.

How do I install Experimental Design in Claude Code?

Run `npx skills add aiming-lab/AutoResearchClaw --skill experimental-design -a claude-code`. Or copy the skill folder (researchclaw/skills/builtin/experiment/experimental-design in aiming-lab/AutoResearchClaw) into .claude/skills/experimental-design in your project. Claude Code loads it when a task matches its description.

How do I install Experimental Design in Codex?

Run `npx skills add aiming-lab/AutoResearchClaw --skill experimental-design -a codex`. Or copy the skill folder (researchclaw/skills/builtin/experiment/experimental-design in aiming-lab/AutoResearchClaw) into .agents/skills/experimental-design in your project. Codex loads it when a task matches its description.

Can I use Experimental Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aiming-lab/AutoResearchClaw --skill experimental-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experimental-design, .gemini/skills/experimental-design, .github/skills/experimental-design and .opencode/skills/experimental-design in your project.

What does Experimental Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Experimental Design is instructions for the agent only.

Does Experimental Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experimental Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experimental Design use?

Experimental Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experimental Design use?

About 286 tokens (SKILL.md is roughly 1.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experimental Design?

Skills that share tags, products or a category with Experimental Design: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.5k stars) and Research Refine Pipeline (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experimental Design?

aiming-lab (a GitHub organization) maintains it in aiming-lab/AutoResearchClaw, which has 14,595 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on August 19, 2026.

Source: aiming-lab/AutoResearchClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.