Agent skill

Experiment Design

by fcakyon in fcakyon/phd-skills

A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

MITAuto-check passedResearch & Science

Install Experiment Design

skills CLI
$ npx skills add fcakyon/phd-skills --skill experiment-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fcakyon/phd-skills experiment-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fcakyon/phd-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/experiment-design .claude/skills/experiment-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-design
GitHub stars
414
Token cost
~987 tokens
SKILL.md length
452 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

  • Works in 7 steps: Understand the Research Question → Single-Variable Isolation → Experiment Matrix → …
  • The user wants to design experiments
  • SKILL.md covers Step 1: Understand the…, Step 2: Single-Variable…, Step 3: Experiment Matrix and Step 4: Resource Estimation, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment Design is an agent skill from fcakyon/phd-skills. Use when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies. Triggers on phrases like "design ablation", "plan experiment", "what experiments should I run", "baseline comparison", or "experiment matrix".

Its SKILL.md is about 990 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Experimental design. The repository describes itself as: PhD Research Skills for Claude Code: paper reproduction, experiment design, paper review, result comparison and more. The licence is MIT.

When your agent uses it

  • The user wants to design experiments
  • Plan ablation studies
  • Structure baselines
  • Create incremental evaluation strategies

Example prompts

  • “design ablation”
  • “plan experiment”
  • “what experiments should I run”
  • “/experiment-design”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Understand the Research Question
  2. Single-Variable Isolation
  3. Experiment Matrix
  4. Resource Estimation
  5. Config Stub Generation
  6. Execution Plan
  7. Analysis Plan

What it can do on your machine

Read from SKILL.md and the folder at commit 67acd61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Design loads about 987 tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 452 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~987

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fcakyon/phd-skills at commit 67acd61, republished under its MIT licence (© fcakyon). 452 words, ~987 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-design/SKILL.md (or your agent's skills folder).
name
experiment-design
description
Use when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies. Triggers on phrases like "design ablation", "plan experiment", "what experiments should I run", "baseline comparison", or "experiment matrix".

Experiment Design Methodology

You are helping a researcher design rigorous experiments. Follow this methodology systematically.

Step 1: Understand the Research Question

Before designing any experiment:

  • Ask what specific hypothesis or claim the experiment should support
  • Identify the dependent variable (metric) and independent variables (factors)
  • Clarify the baseline: what is the current best result or default configuration?

Step 2: Single-Variable Isolation

Every ablation study must change exactly ONE variable at a time. For each factor:

  1. Define the factor — what is being varied (e.g., loss function, learning rate, architecture component)
  2. List levels — all values this factor will take (e.g., CE, focal, VAR)
  3. Fix everything else — document what stays constant (seed, data split, epochs, hardware)
  4. Predict outcome — before running, state what you expect and why

Template for each ablation row:

| Run ID | Factor | Value | Fixed Config | Expected Outcome |
|--------|--------|-------|-------------|-----------------|

Step 3: Experiment Matrix

For multi-factor studies, use a structured matrix:

  1. Full factorial — if factors are few (≤3) and levels are few (≤3 each)
  2. Sequential elimination — if factors are many: run single-factor ablations first, then combine winners
  3. Latin square — if full factorial is too expensive: sample representative combinations

Always calculate total runs before committing:

Total runs = product of all factor levels
GPU hours = total runs × hours_per_run

Step 4: Resource Estimation

For each experiment plan, estimate:

  • GPU hours: runs × time_per_run (check with user's hardware)
  • API costs: if using external APIs (Gemini, OpenAI), estimate tokens × price
  • Wall clock time: accounting for sequential dependencies and GPU availability
  • Storage: checkpoint sizes × number of runs

Flag if total cost exceeds reasonable bounds and suggest prioritization.

Step 5: Config Stub Generation

Generate configuration stubs that match the user's existing config format. Read existing configs first to match:

  • File format (YAML, JSON, TOML)
  • Key naming conventions
  • Directory structure for outputs
  • Logging/tracking integration (wandb, neptune, tensorboard)
Show full SKILL.md (172 more words)Show less

Step 6: Execution Plan

Create a concrete execution plan:

  1. Order runs by dependency (baselines first, then ablations)
  2. Identify which runs can be parallelized across GPUs
  3. Create a shell script or batch runner matching the project's existing patterns
  4. Include checkpointing strategy for long runs

Step 7: Analysis Plan

Before running, define how results will be analyzed:

  • Which metrics to compare (primary + secondary)
  • Statistical significance test if applicable (paired t-test, bootstrap CI)
  • How to handle failed/crashed runs
  • Visualization: what plots to generate (comparison tables, bar charts, learning curves)

Verification Checkpoints

Before finalizing the experiment plan:

  • Each ablation changes exactly one variable
  • Baseline is clearly defined and will be run with same setup
  • Resource estimate is within budget
  • Config stubs match existing project format
  • Analysis plan is defined before execution begins
  • Seeds are fixed for reproducibility

Output Format

Always produce:

  1. Experiment matrix table — all runs with their configurations
  2. Resource estimate — GPU hours, API costs, storage
  3. Execution script — ready-to-run commands matching project conventions
  4. Analysis plan — metrics, comparisons, visualizations

© fcakyon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/experiment-design of fcakyon/phd-skills.

Open the folder on GitHubat commit 67acd61

Compare with similar skills

Experiment Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Design this skillfcakyon/phd-skills414—~987Automated safety check: PassMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k23 repos~5.9kAutomated safety check: NotesMIT
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1287 repos~2.3kAutomated safety check: NotesNone
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.5k—~2.8kAutomated safety check: PassCC-BY-4.0
Research Refine PipelinezjYao36/Auto-Research-Refine1286 repos~1.4kAutomated safety check: NotesNone
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 23 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 7 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.5k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 6 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Experimental Design

    Oleafly/Oleafly

    Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so results are interpretable.

    206 GitHub starsUsed in 4 repos~3.5k tokens
    Research & ScienceAuto-check: notes

More from fcakyon/phd-skills

All 12 skills in this repo
  • Reproduce

    fcakyon/phd-skills

    End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Compare

    fcakyon/phd-skills

    Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

    414 GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Debug

    fcakyon/phd-skills

    Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills.

    414 GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Setup

    fcakyon/phd-skills

    A skill your agent uses when the user wants to set up or troubleshoot a LaTeX environment, choose between biber and bibtex, install packages for a specific venue template, or configure compilation.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check: notes
  • Launch

    fcakyon/phd-skills

    Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

    414 GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Literature Research

    fcakyon/phd-skills

    A skill your agent uses when the user wants to find related work, survey a research area, identify literature gaps, or discover open-source implementations.

    414 GitHub stars~999 tokensUpdated 21 days ago
    Auto-check passed

Questions about Experiment Design

What does Experiment Design do?

A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies. Experiment Design is an agent skill from fcakyon/phd-skills. Use when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

When should I use Experiment Design?

Experiment Design fits situations like: the user wants to design experiments; plan ablation studies; structure baselines; create incremental evaluation strategies.

How do I install Experiment Design in Claude Code?

Run `npx skills add fcakyon/phd-skills --skill experiment-design -a claude-code`. Or copy the skill folder (plugin/skills/experiment-design in fcakyon/phd-skills) into .claude/skills/experiment-design in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Design in Codex?

Run `npx skills add fcakyon/phd-skills --skill experiment-design -a codex`. Or copy the skill folder (plugin/skills/experiment-design in fcakyon/phd-skills) into .agents/skills/experiment-design in your project. Codex loads it when a task matches its description.

Can I use Experiment Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fcakyon/phd-skills --skill experiment-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-design, .gemini/skills/experiment-design, .github/skills/experiment-design and .opencode/skills/experiment-design in your project.

What does Experiment Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Design is instructions for the agent only.

Does Experiment Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Design use?

Experiment Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Design use?

About 987 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment Design?

Skills that share tags, products or a category with Experiment Design: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.5k stars) and Research Refine Pipeline (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Design?

fcakyon (a GitHub user) maintains it in fcakyon/phd-skills, which has 414 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 16, 2026.

Source: fcakyon/phd-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.