Agent skill

Experiment Pipeline

by EvoScientist in EvoScientist/EvoSkills

Guides structured 4-stage experiment execution with attempt budgets and gate conditions: Stage 1 initial implementation (reproduce baseline), Stage 2 hyperparameter tuning, Stage 3 proposed method…

Apache-2.0Auto-check passedDevelopment

Install Experiment Pipeline

skills CLI
$ npx skills add EvoScientist/EvoSkills --skill experiment-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvoScientist/EvoSkills experiment-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvoScientist/EvoSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment-pipeline .claude/skills/experiment-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-pipeline
GitHub stars
478
Used in
3 other repos
Token cost
~4.4k tokens
SKILL.md length
2,216 words
Files
6 (incl. references, assets)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides structured 4-stage experiment execution with attempt budgets and gate conditions: Stage 1 initial implementation (reproduce baseline), Stage 2 hyperparameter tuning, Stage 3 proposed method…

  • Works in 4 steps: Initial Implementation → Hyperparameter Tuning → Proposed Method → …
  • IVE/ESE) and experiment-craft (5-step diagnostic on failure)
  • SKILL.md covers When to Use This Skill, The Pipeline Mindset, Before Starting: Load Prior… and 4-Stage Pipeline Overview, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment Pipeline is an agent skill from EvoScientist/EvoSkills. Guides structured 4-stage experiment execution with attempt budgets and gate conditions: Stage 1 initial implementation (reproduce baseline), Stage 2 hyperparameter tuning, Stage 3 proposed method validation, Stage 4 ablation study. Integrates with evo-memory (load prior strategies, trigger IVE/ESE) and experiment-craft (5-step diagnostic on failure). Use when: user has a planned experiment, needs to reproduce baselines, organize experiment workflow, or systematically validate a method. Do NOT use for debugging a…

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files and assets (for example `assets/pipeline-tracker-template.md`, `assets/stage-log-template.md` and `references/attempt-budget-guide.md`).

It sits in Development. The repository describes itself as: 🧬 Extend EvoScientist with Installable Skill & Knowledge Packs. The licence is Apache-2.0.

When your agent uses it

  • IVE/ESE) and experiment-craft (5-step diagnostic on failure)
  • : user has a planned experiment
  • Needs to reproduce baselines
  • Organize experiment workflow

Example prompts

  • “Use the experiment-pipeline skill to guide structured 4-stage experiment execution with attempt budgets and gate conditions: Stage 1 initial…”
  • “/experiment-pipeline”

Requirements

  • Pre-approved tools (allowed-tools): write_file, edit_file, read_file, think_tool, execute

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Initial Implementation
  2. Hyperparameter Tuning
  3. Proposed Method
  4. Ablation Study

What it can do on your machine

Read from SKILL.md and the folder at commit 9a9f8cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • write_file
    • edit_file
    • read_file
    • think_tool
    • execute

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Pipeline loads about 4.4k tokens when it runs, and up to ~9.5k if it reads all its reference files. Until then it costs about 162 tokens; SKILL.md has 2,216 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~162
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvoScientist/EvoSkills at commit 9a9f8cf, republished under its Apache-2.0 licence (© EvoScientist). 2,216 words, ~4,428 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
experiment-pipeline
description
Guides structured 4-stage experiment execution with attempt budgets and gate conditions: Stage 1 initial implementation (reproduce baseline), Stage 2 hyperparameter tuning, Stage 3 proposed method validation, Stage 4 ablation study. Integrates with evo-memory (load prior strategies, trigger IVE/ESE) and experiment-craft (5-step diagnostic on failure). Use when: user has a planned experiment, needs to reproduce baselines, organize experiment workflow, or systematically validate a method. Do NOT use for debugging a specific experiment failure (use experiment-craft) or designing which experiments to run (use paper-planning).
allowed-tools
write_file, edit_file, read_file, think_tool, execute
metadata.author
EvoScientist
metadata.version
1.0.0
metadata.tags
core, experimentation, experiment-design

Experiment Pipeline

A structured 4-stage framework for executing research experiments from initial implementation through ablation study, with attempt budgets and gate conditions that prevent wasted effort. This follows the Experiment Tree Search design from the EvoScientist paper, where the engineer agent iteratively generates executable code, runs experiments, and records structured execution results at each stage.

When to Use This Skill

  • User has a planned experiment and needs to organize the execution workflow
  • User wants to systematically validate a novel method against baselines
  • User asks about experiment stages, attempt budgets, or when to move on
  • User needs to reproduce baseline results before testing their method
  • User mentions "experiment pipeline", "baseline first", "ablation study", "stage budget", "experiment execution"

The Pipeline Mindset

Experiments fail for two reasons: wrong order and no stopping criteria. Most researchers jump straight to testing their novel method without verifying their baseline setup, then wonder why results don't make sense. Others spend weeks tuning hyperparameters without a budget, hoping the next run will work.

The 4-stage pipeline solves both problems. It enforces a strict order (each stage validates assumptions the next stage depends on) and assigns attempt budgets (forcing systematic thinking over brute-force iteration).

Before Starting: Load Prior Knowledge

If coming from research-ideation, your research proposal (Step 7) provides the experiment plan — datasets, baselines, metrics, and ablation design — that maps directly to Stages 1-4 below.

Before entering the pipeline, load Experimentation Memory (M_E) from prior cycles:

  1. Refer to the evo-memory skill → Read M_E at /memory/experiment-memory.md
  2. Select the top-1 entry (k_E=1) most relevant to the current experiment domain by comparing each entry's Context and Category against the current problem
  3. The selected strategy informs hyperparameter ranges (Stage 2), debugging approaches (Stages 1-3), and training configurations across all stages
  4. If M_E doesn't exist yet (first cycle), skip this step and proceed — your results will seed M_E via ESE after pipeline completion

4-Stage Pipeline Overview

Each stage follows a generate → execute → record → diagnose → revise loop:

StageGoalBudget (N_E^s)Gate Condition
1. Initial ImplementationGet baseline code running and reproduce known results≤20 attemptsMetrics within 2% of reported values (or within reported variance)
2. Hyperparameter TuningOptimize config for your setup≤12 attemptsStable config, variance < 5% across 3 runs
3. Proposed MethodImplement & validate novel method≤12 attemptsOutperforms tuned baseline on primary metric, consistent across 3 runs
4. Ablation StudyProve each component's contribution≤18 attemptsAll claims evidenced with controlled experiments

Each stage saves artifacts to /experiments/stageN_name/.

The Stage Loop

Within every stage, repeat this cycle for each attempt:

  1. Generate: Form a hypothesis or plan for this attempt. What specifically will you try? What do you expect to happen?
  2. Execute: Run the experiment. Record exact configuration, code changes, and runtime.
  3. Record: Log results immediately using the stage log template. Include both metrics and observations.
  4. Diagnose: Compare results to expectations. If they match, assess the gate condition. If they don't, load experiment-craft for the 5-step diagnostic flow.
  5. Revise: Based on diagnosis, either advance to the next stage (gate met) or plan the next attempt (gate not met).

Stage 1: Initial Implementation

Goal: Find or generate executable baseline code and verify it reproduces published results. This stage corresponds to the paper's "initial implementation" — the engineer agent searches for working code, runs it, and records structured execution results.

Why this matters: If you can't get the baseline running and reproducing known results, every subsequent comparison is meaningless. Initial implementation validates your data pipeline, evaluation code, training infrastructure, and understanding of prior work.

Budget: ≤20 attempts (N_E^1=20). Baselines can be tricky — missing details in papers, version mismatches, unreported preprocessing steps. 20 attempts gives enough room to debug without allowing infinite tinkering.

Gate: Primary metrics within 2% of reported values (or within the reported variance if provided).

Process:

  1. Find the original baseline code (official repo, re-implementations, or write from paper description)
  2. Get the code running in your environment — resolve dependencies, fix compatibility issues
  3. Match the exact training configuration from the paper (dataset splits, preprocessing, hyperparameters)
  4. Run and compare metrics. If off by >2%, diagnose the gap
  5. Common pitfalls: different random seeds, different data splits, unreported data augmentation, framework version differences

When to load experiment-craft: If attempts 1-5 all fail significantly (>10% gap), switch to the 5-step diagnostic flow to isolate the cause before burning more attempts.

Output: /experiments/stage1_baseline/ containing results, config, and verified baseline code.

See references/stage-protocols.md for detailed initial implementation checklists.

Stage 2: Hyperparameter Tuning

Goal: Find the optimal hyperparameter configuration for YOUR specific setup.

Why this matters: Published hyperparameters are tuned for the authors' setup. Your hardware, data version, framework version, or subtle implementation differences mean their config may not be optimal for you. Tuning now prevents confounding your novel method's results with suboptimal baselines.

Budget: ≤12 attempts. Hyperparameter tuning has diminishing returns. If 12 structured attempts don't find a stable config, the problem is likely deeper than hyperparameters.

Gate: Stable configuration found — variance < 5% across 3 independent runs with different random seeds.

Process:

  1. Identify the most sensitive hyperparameters (usually: learning rate, batch size, loss weights)
  2. Start with coarse search on the most sensitive parameter
  3. Narrow the range based on results, then move to the next parameter
  4. Validate final config with 3 independent runs

Priority order for tuning: Learning rate → batch size → loss weights → regularization → architecture-specific params. This order reflects typical sensitivity.

When to load experiment-craft: If results are highly unstable (variance > 20%) across runs, there's likely a training instability issue. Use diagnostic flow.

Output: /experiments/stage2_tuning/ containing tuning logs, final config, and stability verification.

See references/attempt-budget-guide.md for budget rationale and adjustment rules.

Stage 3: Proposed Method

Goal: Implement and validate your novel method, demonstrating improvement over the tuned baseline.

Why this matters: This is the core contribution. But because you've verified the baseline (Stage 1) and optimized the config (Stage 2), any improvement you see is genuinely attributable to your method — not to a better-tuned setup or a broken baseline.

Budget: ≤12 attempts. Your method should work within a reasonable number of iterations if the underlying idea is sound. Excessive attempts suggest a fundamental problem, not a tuning issue.

Gate: Outperforms the tuned baseline on the primary metric. The improvement should be consistent across at least 3 runs.

Process:

  1. Implement the core method incrementally — don't add everything at once
  2. Test each component's integration with the baseline pipeline
  3. Run full training and compare against Stage 2 results
  4. If underperforming, isolate which component causes the gap

Integration strategy: Add your method's components one at a time to the working baseline. Each added component should stay within 20% of the baseline's performance — if a single component causes a >20% regression, isolate and debug it before proceeding. Never integrate the full method in one shot.

When to load experiment-craft: When your method underperforms the baseline despite correct implementation. The 5-step diagnostic flow will help distinguish between implementation bugs and fundamental issues.

Critical decision — failure classification: If the method underperforms the baseline after exhausting the attempt budget, hand off to evo-memory for IVE (Idea Validation Evolution) — this is evo-memory's job, not this skill's. IVE triggers under two conditions:

  1. No executable code: Cannot find working code within the budget at any stage.
  2. Worse than baseline: Experiments complete but the method underperforms.

The evo-memory skill will classify the failure as:

  • Implementation failure: Bugs or missing tricks → retryable in a future cycle.
  • Fundamental direction failure: Core idea doesn't work → update ideation memory to prevent retrying.

Output: /experiments/stage3_method/ containing method code, results, comparison with baseline.

Show full SKILL.md (958 more words)Show less

Stage 4: Ablation Study

Goal: Prove that each component of your method contributes meaningfully to the final result.

Why this matters: Reviewers will ask "is component X really necessary?" for every part of your method. Without ablation, you can't answer. More importantly, ablation helps YOU understand why your method works — sometimes components you thought were important aren't, and vice versa.

Budget: ≤18 attempts. Ablation requires multiple controlled experiments — one per component being ablated, plus interaction effects. 18 attempts covers a method with 4-5 components.

Gate: Every claimed contribution is supported by a controlled experiment showing its effect.

Process:

  1. List all components of your method that you claim contribute to performance
  2. Design ablation experiments: remove ONE component at a time, measure the impact
  3. For components that interact, test interaction effects
  4. Verify that no single component's removal improves results (would invalidate the claim)

Three ablation designs:

  • Leave-one-out: Remove each component individually. Shows each component's marginal contribution.
  • Additive: Start from baseline, add components one at a time. Shows incremental gains.
  • Substitution: Replace your component with an alternative approach. Shows your component is better than alternatives, not just better than nothing.

When to load experiment-craft: If ablation results contradict your hypothesis (removing a component improves results), use diagnostic flow to understand why.

Output: /experiments/stage4_ablation/ containing ablation results table, per-component analysis.

See references/stage-protocols.md for detailed ablation design patterns.

Integrating experiment-craft for Diagnosis

When a stage attempt fails, refer to the experiment-craft skill for structured diagnosis:

  1. Follow the experiment-craft diagnostic protocol
  2. Run the 5-step diagnostic flow (observe, hypothesize, test, conclude, prescribe)
  3. The diagnosis does NOT consume your stage budget — it's a free analysis step
  4. The diagnosis output (a prescription) becomes the plan for your next attempt
  5. Return to the pipeline and record the diagnosis in your trajectory log

Trigger points: After any failed attempt in any stage. Especially important:

  • Stage 1: After 5+ failed attempts (>10% gap from reported metrics)
  • Stage 2: When variance > 20% across runs
  • Stage 3: When method consistently underperforms baseline
  • Stage 4: When ablation results contradict your hypothesis

Code Trajectory Logging

Every attempt across all stages should be logged in a structured format that captures not just WHAT you did but WHY and WHAT YOU LEARNED. These logs feed into evo-memory's Experiment Strategy Evolution (ESE) mechanism.

For each attempt, record:

  • Attempt number and stage
  • Hypothesis: What you expected and why
  • Code changes: Summary of what was modified (not a full diff, but the key changes)
  • Result: Metrics and observations
  • Analysis: Whether the hypothesis was confirmed or refuted, and what you learned

See references/code-trajectory-logging.md for the full logging format and how logs feed into evo-memory.

Counterintuitive Pipeline Rules

Prioritize these rules during experiment execution:

  1. Initial implementation is not wasted time: It validates your entire infrastructure — data pipeline, evaluation code, training setup. Skipping it means every subsequent result is built on unverified ground. Most "method doesn't work" bugs are actually baseline setup bugs.

  2. Budget limits prevent rabbit holes: Fixed attempt budgets force you to think systematically. When you know you have 12 attempts, you design each one to maximize information. Without limits, attempt #47 is rarely more informative than attempt #12 — it's just more desperate.

  3. Stage order is non-negotiable: Each stage validates assumptions the next depends on. Skipping Stage 1 means Stage 3 results could be wrong due to a broken baseline. Skipping Stage 2 means Stage 3 improvements might just be better hyperparameters, not a better method. There are no shortcuts.

  4. Ablation is not optional cleanup: It's the primary evidence that your method works for the right reasons. A method that outperforms the baseline but has no ablation is a method you don't understand. Reviewers know this.

  5. Failed attempts are data, not waste: Each failed attempt narrows the search space and reveals something about the problem. Log failures carefully — they feed into evo-memory and prevent future researchers from repeating the same mistakes.

  6. Early termination is a feature: Stopping before budget exhaustion is smart, not lazy. If the gate is clearly unachievable after systematic attempts, escalate to evo-memory IVE rather than burning remaining budget on increasingly random variations.

Handoff to Paper Writing

When all four stages are complete, pass these artifacts to paper-writing:

ArtifactSource StageUsed By
Initial implementation resultsStage 1Comparison tables, setup verification
Optimal hyperparameter configStage 2Reproducibility section
Method vs baseline comparisonStage 3Main results table
Ablation study resultsStage 4Ablation table, contribution claims
Code trajectory logs (all stages)All stagesMethod section details, supplementary
Implementation details and tricksStages 1-3Method section, reproducibility (captured in trajectory log Analysis fields and [Reusable] tags)

Also pass results to evo-memory for evolution updates:

  • If any stage exhausts budget without executable code, OR Stage 3 method underperforms the tuned baseline → trigger IVE (Idea Validation Evolution)
  • If all stages succeeded → trigger ESE (Experiment Strategy Evolution)

Skill Integration

Before Starting (load memory)

Refer to the evo-memory skill to read Experimentation Memory: → Read M_E at /memory/experiment-memory.md

On Failure (within any stage)

Refer to the experiment-craft skill for 5-step diagnostic: → Run diagnosis → Return to pipeline

On IVE Trigger (budget exhausted or method underperforms)

Refer to the evo-memory skill for failure classification: → Run IVE protocol

On Pipeline Success (all 4 stages complete)

Refer to the evo-memory skill for strategy extraction: → Run ESE protocol with trajectory logs

Handoff to Paper Writing

Refer to the paper-writing skill: → Pass all stage artifacts

Reference Navigation

TopicReference FileWhen to Use
Per-stage checklists and patternsstage-protocols.mdDetailed guidance for each stage
Budget rationale and adjustmentattempt-budget-guide.mdWhen budgets feel too tight or too loose
Code trajectory logging formatcode-trajectory-logging.mdRecording attempts for evo-memory
Stage log templatestage-log-template.mdLogging a single stage's progress
Pipeline tracker templatepipeline-tracker-template.mdTracking the full 4-stage pipeline

© EvoScientist, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references, assets) in skills/experiment-pipeline of EvoScientist/EvoSkills.

  • SKILL.md
  • assets/pipeline-tracker-template.md
  • assets/stage-log-template.md
  • references/attempt-budget-guide.md
  • references/code-trajectory-logging.md
  • references/stage-protocols.md

Open the folder on GitHubat commit 9a9f8cf

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in EvoScientist/EvoSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Experiment Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Pipeline this skillEvoScientist/EvoSkills4783 repos~4.4kAutomated safety check: PassApache-2.0
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from EvoScientist/EvoSkills

All 16 skills in this repo
  • Evomath Tao

    EvoScientist/EvoSkills

    A skill your agent uses whenever the user submits a non-trivial mathematical claim that needs a rigorous proof or audit.

    478 GitHub starsUsed in 2 repos~3.8k tokens
    Auto-check passed
  • Paper Figures

    EvoScientist/EvoSkills

    A skill your agent uses to produce standalone, publication-ready PNG graphics and reproducible matplotlib scripts from tabular data (CSVs or DataFrames).

    478 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Experiment Iterative Coder

    EvoScientist/EvoSkills

    Iterative code refinement through plan → code → evaluate → refine cycles.

    478 GitHub starsUsed in 3 repos~2.5k tokens
    Auto-check passed
  • Paper Planning

    EvoScientist/EvoSkills

    Guides pre-writing planning for academic papers with 4 structured steps: story design (task-challenge-insight-contribution-advantage), experiment planning (comparisons + ablations), figure design…

    478 GitHub starsUsed in 3 repos~2.4k tokens
    Auto-check passed
  • Research Survey

    EvoScientist/EvoSkills

    Generates structured literature survey reports from collected papers using a multi-stage pipeline: outline generation (query-type adaptive) → draft survey → section-by-section expansion → summary…

    478 GitHub starsUsed in 3 repos~2.5k tokens
    Auto-check passed
  • Paper Navigator

    EvoScientist/EvoSkills

    Find and read academic papers (S2 + arXiv). An agent skill from EvoScientist/EvoSkills.

    478 GitHub stars~6.3k tokensUpdated 9 days ago
    Auto-check: notes

Categories

Questions about Experiment Pipeline

What does Experiment Pipeline do?

Guides structured 4-stage experiment execution with attempt budgets and gate conditions: Stage 1 initial implementation (reproduce baseline), Stage 2 hyperparameter tuning, Stage 3 proposed method…. Experiment Pipeline is an agent skill from EvoScientist/EvoSkills. Guides structured 4-stage experiment execution with attempt budgets and gate conditions: Stage 1 initial implementation (reproduce baseline), Stage 2 hyperparameter tuning, Stage 3 proposed method validation, Stage 4 ablation study.

When should I use Experiment Pipeline?

Experiment Pipeline fits situations like: IVE/ESE) and experiment-craft (5-step diagnostic on failure); : user has a planned experiment; needs to reproduce baselines; organize experiment workflow.

How do I install Experiment Pipeline in Claude Code?

Run `npx skills add EvoScientist/EvoSkills --skill experiment-pipeline -a claude-code`. Or copy the skill folder (skills/experiment-pipeline in EvoScientist/EvoSkills) into .claude/skills/experiment-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Pipeline in Codex?

Run `npx skills add EvoScientist/EvoSkills --skill experiment-pipeline -a codex`. Or copy the skill folder (skills/experiment-pipeline in EvoScientist/EvoSkills) into .agents/skills/experiment-pipeline in your project. Codex loads it when a task matches its description.

Can I use Experiment Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvoScientist/EvoSkills --skill experiment-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-pipeline, .gemini/skills/experiment-pipeline, .github/skills/experiment-pipeline and .opencode/skills/experiment-pipeline in your project.

What does Experiment Pipeline need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Pipeline is instructions for the agent only. Its frontmatter pre-approves these tools: write_file, edit_file, read_file, think_tool, execute.

Does Experiment Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Pipeline use?

Experiment Pipeline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Pipeline use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.1k tokens, read only when the agent opens those files.

What are the alternatives to Experiment Pipeline?

Skills that share tags, products or a category with Experiment Pipeline: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Pipeline?

EvoScientist (a GitHub organization) maintains it in EvoScientist/EvoSkills, which has 478 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 30, 2026.

Source: EvoScientist/EvoSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.