Agent skill

Experiment Craft

by EvoScientist in EvoScientist/EvoSkills

A skill your agent uses when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs…

Apache-2.0Auto-check passedDevelopment

Install Experiment Craft

skills CLI
$ npx skills add EvoScientist/EvoSkills --skill experiment-craft -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvoScientist/EvoSkills experiment-craft --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvoScientist/EvoSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment-craft .claude/skills/experiment-craft && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-craft
GitHub stars
476
Used in
3 other repos
Token cost
~2k tokens
SKILL.md length
919 words
Files
3 (incl. references, assets)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs…

  • Works in 5 steps: Collect Failure Cases → Find a Working Version → Bridge the Gap → …
  • The user wants to debug
  • SKILL.md covers When to Use This Skill, The Debugging Mindset, 5-Step Diagnostic Flow and Counterintuitive Experiment…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment Craft is an agent skill from EvoScientist/EvoSkills. Use this skill when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs, hypotheses, failures, results, and next steps during active research. Apply it to underperforming methods, training that will not converge, regressions after a change, inconsistent results across datasets, aimless experimentation without progress, and questions like 'why doesn't this work?', 'no progress after many attempts', or…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files and assets (for example `assets/experiment-log-template.md` and `references/debugging-methodology.md`).

It sits in Development, covering Hypothesis generation, Debugging and A/B testing. The repository describes itself as: 🧬 Extend EvoScientist with Installable Skill & Knowledge Packs. The licence is Apache-2.0.

When your agent uses it

  • The user wants to debug
  • Systematically iterate on an experiment that already exists
  • They need a structured experiment log for tracking runs
  • Next steps during active research

Example prompts

  • “why doesn”
  • “no progress after many attempts”
  • “how should I investigate this failure?”
  • “/experiment-craft”

Requirements

  • Pre-approved tools (allowed-tools): write_file, edit_file, read_file, think_tool, execute

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Collect Failure Cases
  2. Find a Working Version
  3. Bridge the Gap
  4. Hypothesize and Verify
  5. Propose and Implement a Fix

What it can do on your machine

Read from SKILL.md and the folder at commit 9a9f8cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • write_file
    • edit_file
    • read_file
    • think_tool
    • execute

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Craft loads about 2k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 235 tokens; SKILL.md has 919 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~235
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvoScientist/EvoSkills at commit 9a9f8cf, republished under its Apache-2.0 licence (© EvoScientist). 919 words, ~2,007 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-craft/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
experiment-craft
description
Use this skill when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs, hypotheses, failures, results, and next steps during active research. Apply it to underperforming methods, training that will not converge, regressions after a change, inconsistent results across datasets, aimless experimentation without progress, and questions like 'why doesn't this work?', 'no progress after many attempts', or 'how should I investigate this failure?'. Also use it for setting up practical experiment logging/record-keeping that supports debugging and iteration. Do not use it for designing a brand-new experiment pipeline or full experiment program (use experiment-pipeline), generating research ideas, fixing isolated coding/syntax errors, or writing retrospective summaries into research memory/notes/knowledge bases.
allowed-tools
write_file, edit_file, read_file, think_tool, execute
metadata.author
EvoScientist
metadata.version
1.0.0
metadata.tags
core, experimentation, experiment-design

Experiment Craft

A systematic approach to running, debugging, and iterating on research experiments. The critical skill is not running more experiments — it's understanding WHY experiments fail.

When to Use This Skill

  • User's experiment is not working or producing unexpected results
  • User needs help diagnosing why a method fails on certain data
  • User wants to organize their experiment process with structured logging
  • User asks about debugging research code or iterating on approaches
  • User mentions "experiment debugging", "why doesn't this work", "experiment log", "results are wrong"

This skill is typically loaded from within experiment-pipeline when a stage attempt fails. After debugging, return to the pipeline's stage-gate structure to continue. Can also be used standalone for any experiment debugging.

The Debugging Mindset

Finding WHY experiments fail is the most critical research skill. Not analyzing results leads to two failure modes:

  1. Slow progress: Running random experiments without understanding failure causes
  2. Wasted time: Abandoning good approaches because activation tricks were missed

The goal is not to run more experiments. The goal is to run the RIGHT experiments — ones that isolate causes and test specific hypotheses.

5-Step Diagnostic Flow

When an experiment fails or produces unexpected results, follow these five steps:

Step 1: Collect Failure Cases

Gather concrete examples of bad results. Look at the actual outputs, not just aggregate metrics. What specifically went wrong? Are the failures systematic or random?

Step 2: Find a Working Version

You need a baseline that works. Two ways to find one:

  • Simplify the task: Reduce data complexity, relax the task setting, add more supervision, use easier inputs
  • Remove your changes: Start from the baseline method and remove your algorithmic improvements one by one

If you can't find any working version, simplify further until something works. There is always a simple enough version that works.

Step 3: Bridge the Gap

Starting from the working version, incrementally add complexity until it breaks:

  • Add ONE factor at a time (more complex data, one algorithmic change, one constraint)
  • Find the single factor that causes failure
  • The more atomic the identified cause, the more useful the diagnosis

This step isolates the cause. Without it, you're guessing.

Step 4: Hypothesize and Verify

Based on the isolated cause from Step 3:

  1. List possible explanations for why this factor causes failure
  2. Rank by likelihood (based on your understanding and literature)
  3. Design targeted experiments to verify or eliminate each hypothesis
  4. Confirm the actual cause experimentally — don't rely on intuition alone
Step 5: Propose and Implement a Fix

Based on the confirmed cause:

  • Search for techniques that address this specific cause (use your literature tree from the research-ideation skill)
  • Design a fix that targets the confirmed cause, not the surface symptom
  • Verify the fix works on the original failure cases
  • Check that the fix doesn't break previously working cases

See references/debugging-methodology.md for detailed branching logic and a cause taxonomy.

Show full SKILL.md (443 more words)Show less

Counterintuitive Experiment Rules

Prioritize these rules during experimental work:

  1. Change only one variable at a time: If you change two things and it works, you don't know which one fixed it. If you change two things and it doesn't work, you don't know which one is wrong. Single-variable changes are slower per experiment but faster overall.
  2. Fast iteration requires effective experiments, not more experiments: Blind experimentation makes things worse. One well-designed diagnostic experiment is worth ten random trials.
  3. Some great techniques don't work alone: They need specific activation tricks — learning rate schedules, initialization schemes, data preprocessing steps. Don't discard a technique after one failed attempt. Check related papers for their undisclosed tricks.
  4. Check related papers for their tricks: Papers solving similar technical challenges often have critical implementation details buried in supplementary material or code. These tricks can make the difference between a technique working or failing.
  5. "Once you've ruled out the impossible, whatever remains must be true": Systematic elimination beats intuition. When debugging, explicitly list ALL possible causes, then eliminate them one by one with targeted experiments.

Experiment Logging

Every experiment should be logged with five sections. Use the template at assets/experiment-log-template.md.

SectionWhat to Record
PurposeWhy you're running this experiment; what you expect to learn
SettingData, algorithm changes, hyperparameters — everything needed to reproduce
ResultsQuantitative metrics + qualitative observations + specific good/failure cases
AnalysisDo results match expectations? If not, hypothesized causes ranked by likelihood
Next StepsWhat to do based on the analysis — YOU are the project leader

The "Next Steps" section is the most important. Don't wait for someone to tell you what to do next. Analyze your results and propose the next experiment yourself. This is what distinguishes a researcher from a technician.

Cross-cycle learning: If using experiment-pipeline, your experiment logs feed into evo-memory's ESE (Experiment Strategy Evolution) mechanism. Tag reusable strategies with [Reusable] so ESE can extract them for future cycles.

Return to experiment-pipeline

After completing the 5-step diagnostic flow, return to experiment-pipeline with:

  • Confirmed cause of failure (from Step 4)
  • Proposed fix and its verification status (from Step 5)
  • Updated experiment log entry

Handoff to Paper Writing

When experiments succeed and you have a complete set of results, pass these artifacts to paper-writing:

ArtifactSourceUsed By
Final experiment results (tables and figures)Experiment logsExperiments section
Ablation study resultsDiagnostic experimentsAblation tables
Failure case analysisStep 1 + Step 3Limitations discussion
Key implementation details and tricksSteps 3-5Method section / Supplementary
Baseline comparison resultsStep 2Comparison tables

Reference Navigation

TopicReference FileWhen to Use
Debugging methodologydebugging-methodology.mdDiagnosing why experiments fail
Experiment log templateexperiment-log-template.mdRecording experiment details

© EvoScientist, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references, assets) in skills/experiment-craft of EvoScientist/EvoSkills.

  • SKILL.md
  • assets/experiment-log-template.md
  • references/debugging-methodology.md

Open the folder on GitHubat commit 9a9f8cf

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in EvoScientist/EvoSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Experiment Craft next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Craft compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Craft this skillEvoScientist/EvoSkills4763 repos~2kAutomated safety check: PassApache-2.0
Retroapmantza/pi-lens465—~531Automated safety check: PassMIT
Retrobitsocialnet/5chan135—~865Automated safety check: PassGPL-3.0
Parallel Debuggingwshobson/agents40k—~1.2kAutomated safety check: PassMIT
Debugging ExperimentsPostHog/posthog40k—~5.4kAutomated safety check: PassCustom licence
Debug InvestigatorMathews-Tom/armory328—~3.6kAutomated safety check: PassMIT

Similar skills

  • Retro

    apmantza/pi-lens

    Run the pi-lens retrospective — turn a session, an incident, or a merged bug fix into environment changes (checks, hooks, contract lines, deletions), classified mechanical-vs-judgement, each with…

    465 GitHub stars~531 tokensUpdated today
    DevelopmentAuto-check passed
  • Retro

    bitsocialnet/5chan

    Review a coding session, bug fix, or substantive review correction for evidence-backed ways to prevent recurrence, using existing checks and focused guidance.

    135 GitHub stars~865 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Parallel Debugging

    wshobson/agents

    Debug complex issues using competing hypotheses with parallel investigation, evidence collection, and root cause arbitration.

    40k GitHub stars~1.2k tokensUpdated 4 days ago
    DevelopmentAuto-check passed
  • Debugging Experiments

    PostHog/posthog

    Official

    Debug and support PostHog Experiments (A/B tests) for a customer looking at their own results.

    40k GitHub stars~5.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Debug Investigator

    Mathews-Tom/armory

    Hypothesis-driven debugging with ranked hypotheses, git bisect strategy, instrumentation planning, and minimal reproduction design.

    328 GitHub stars~3.6k tokensUpdated 3 days ago
    DevelopmentAuto-check passed
  • ExecuTorch Knowledge Base

    pytorch/executorch

    Answers ExecuTorch questions from a local wiki on backends, export pitfalls, quantization recipes, runtime errors and SoC compatibility.

    5.1k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from EvoScientist/EvoSkills

All 16 skills in this repo
  • Evomath Tao

    EvoScientist/EvoSkills

    A skill your agent uses whenever the user submits a non-trivial mathematical claim that needs a rigorous proof or audit.

    476 GitHub starsUsed in 2 repos~3.8k tokens
    Auto-check passed
  • Paper Figures

    EvoScientist/EvoSkills

    A skill your agent uses to produce standalone, publication-ready PNG graphics and reproducible matplotlib scripts from tabular data (CSVs or DataFrames).

    476 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Experiment Iterative Coder

    EvoScientist/EvoSkills

    Iterative code refinement through plan → code → evaluate → refine cycles.

    476 GitHub starsUsed in 3 repos~2.5k tokens
    Auto-check passed
  • Paper Planning

    EvoScientist/EvoSkills

    Guides pre-writing planning for academic papers with 4 structured steps: story design (task-challenge-insight-contribution-advantage), experiment planning (comparisons + ablations), figure design…

    476 GitHub starsUsed in 3 repos~2.4k tokens
    Auto-check passed
  • Research Survey

    EvoScientist/EvoSkills

    Generates structured literature survey reports from collected papers using a multi-stage pipeline: outline generation (query-type adaptive) → draft survey → section-by-section expansion → summary…

    476 GitHub starsUsed in 3 repos~2.5k tokens
    Auto-check passed
  • Paper Navigator

    EvoScientist/EvoSkills

    Find and read academic papers (S2 + arXiv). An agent skill from EvoScientist/EvoSkills.

    476 GitHub stars~6.3k tokensUpdated 8 days ago
    Auto-check: notes

Questions about Experiment Craft

What does Experiment Craft do?

A skill your agent uses when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs…. Experiment Craft is an agent skill from EvoScientist/EvoSkills. Use this skill when the user wants to debug, diagnose, or systematically iterate on an experiment that already exists, or when they need a structured experiment log for tracking runs, hypotheses, failures, results, and next steps during active research.

When should I use Experiment Craft?

Experiment Craft fits situations like: the user wants to debug; systematically iterate on an experiment that already exists; they need a structured experiment log for tracking runs; Next steps during active research.

How do I install Experiment Craft in Claude Code?

Run `npx skills add EvoScientist/EvoSkills --skill experiment-craft -a claude-code`. Or copy the skill folder (skills/experiment-craft in EvoScientist/EvoSkills) into .claude/skills/experiment-craft in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Craft in Codex?

Run `npx skills add EvoScientist/EvoSkills --skill experiment-craft -a codex`. Or copy the skill folder (skills/experiment-craft in EvoScientist/EvoSkills) into .agents/skills/experiment-craft in your project. Codex loads it when a task matches its description.

Can I use Experiment Craft in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvoScientist/EvoSkills --skill experiment-craft -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-craft, .gemini/skills/experiment-craft, .github/skills/experiment-craft and .opencode/skills/experiment-craft in your project.

What does Experiment Craft need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Craft is instructions for the agent only. Its frontmatter pre-approves these tools: write_file, edit_file, read_file, think_tool, execute.

Does Experiment Craft access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Craft safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Craft use?

Experiment Craft is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Craft use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Experiment Craft?

Skills that share tags, products or a category with Experiment Craft: Retro (apmantza/pi-lens, 465 stars), Retro (bitsocialnet/5chan, 135 stars), Parallel Debugging (wshobson/agents, 40k stars) and Debugging Experiments (PostHog/posthog, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Craft?

EvoScientist (a GitHub organization) maintains it in EvoScientist/EvoSkills, which has 476 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 30, 2026.

Source: EvoScientist/EvoSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.