Agent skill

Ablation Study Planner

by wanshuiyin in wanshuiyin/Auto-claude-code-research-in-sleep

Plans the ablation studies a paper needs after main results support its claim: Codex designs them like a reviewer, Claude Code checks feasibility and runs them.

MITAuto-check: notesResearch & Science

Install Ablation Study Planner

skills CLI
$ npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wanshuiyin/Auto-claude-code-research-in-sleep ablation-planner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ablation-planner .claude/skills/ablation-planner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ablation-planner
GitHub stars
17k
Token cost
~1.3k tokens
SKILL.md length
321 words
Files
1
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Plans the ablation studies a paper needs after main results support its claim: Codex designs them like a reviewer, Claude Code checks feasibility and runs them.

  • Works in 5 steps: Prepare Context → Codex Designs Ablations → Parse Ablation Plan → …
  • Planning ablations after main results support the paper's claim
  • SKILL.md covers Context: $ARGUMENTS, When to Use, Workflow and Rules
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

This skill plans the ablation studies a paper will need once main results support the claim. It splits the work between two models: Codex designs the ablations from a reviewer's point of view, and Claude Code checks feasibility and does the implementation. It applies when the result-to-claim check returns supported or partial, when you ask for ablation planning, or when the auto review loop flags missing ablations.

The agent first reads the project files for the method and its components, current experiment results from logs or W&B, the claims and the available compute. Codex returns a design that is normalized into a structured plan with component ablations at the highest priority. Claude Code then checks the compute budget, which ablations need code and which are config-only, and what can run in parallel, and it proposes cuts if the budget is tight. It creates configs or scripts, smoke tests each, runs them in the suggested order under descriptive names, logs results in `EXPERIMENT_LOG.md` and updates `findings.md`.

The rules keep roles clear: Codex leads, so Claude Code does not pre-filter the list; every ablation states what it tests and what is expected if the component matters, with no just-try-it runs; config-only ablations come first; and cuts are negotiated with Codex rather than dropped silently.

When your agent uses it

  • Planning ablations after main results support the paper's claim
  • Responding to a reviewer-style request for missing ablations
  • Fitting an ablation list to a limited GPU budget

Example prompts

  • “The main results passed result-to-claim. Plan the ablation studies for the paper.”
  • “The review loop says we lack ablations. Design them and check which need code changes.”
  • “We only have a few GPUs free. Prioritize the ablations and propose cuts.”

Requirements

  • Codex reachable through the MCP tool the skill calls
  • Project files such as the research contract and the experiment log
  • Pre-approved tools (allowed-tools): Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Prepare Context
  2. Codex Designs Ablations
  3. Parse Ablation Plan
  4. CC Reviews Feasibility
  5. Implement and Run

What it can do on your machine

Read from SKILL.md and the folder at commit 26b95cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(*)
    • Read
    • Grep
    • Glob
    • Write
    • Edit
    • mcp__codex__codex
    • mcp__codex__codex-reply

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ablation Study Planner loads about 1.3k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 321 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wanshuiyin/Auto-claude-code-research-in-sleep at commit 26b95cf, republished under its MIT licence (© wanshuiyin). 321 words, ~1,283 tokens.

Download SKILL.mdSave it as .claude/skills/ablation-planner/SKILL.md (or your agent's skills folder).
name
ablation-planner
description
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
allowed-tools
Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply
argument-hint
[method-description-or-claim]

Ablation Planner

Systematically design ablation studies that answer the questions reviewers will ask. Codex leads the design (reviewer perspective), CC reviews feasibility and implements.

Context: $ARGUMENTS

When to Use

  • Main results pass /result-to-claim with claim_supported = yes or partial
  • User explicitly requests ablation planning
  • /auto-review-loop reviewer identifies missing ablations

Workflow

Step 1: Prepare Context

CC reads available project files to build the full picture:

  • Method description and components (from idea-stage/docs/research_contract.md, legacy docs/research_contract.md, or project CLAUDE.md)
  • Current experiment results (from EXPERIMENT_LOG.md, EXPERIMENT_TRACKER.md, or W&B)
  • Confirmed and intended claims (from result-to-claim output or project notes)
  • Available compute resources (from CLAUDE.md server config, if present)
Step 2: Codex Designs Ablations
mcp__codex__codex:
  model: gpt-6-astra
  config: {"model_reasoning_effort": "xhigh"}
  prompt: |
    You are a rigorous ML reviewer planning ablation studies.
    Given this method and results, design ablations that:

    1. Isolate the contribution of each novel component
    2. Answer questions reviewers will definitely ask
    3. Test sensitivity to key hyperparameters
    4. Compare against natural alternative design choices

    Method: [description from project files]
    Components: [list of removable/replaceable components]
    Current results: [key metrics from experiments]
    Claims: [what we claim and current evidence]

    For each ablation, specify:
    - name: what to change (e.g., "remove module X", "replace Y with Z")
    - what_it_tests: the specific question this answers
    - expected_if_component_matters: what we predict if the component is important
    - priority: 1 (must-run) to 5 (nice-to-have)

    Also provide:
    - coverage_assessment: what reviewer questions these ablations answer
    - unnecessary_ablations: experiments that seem useful but won't add insight
    - suggested_order: run order optimized for maximum early information
    - estimated_compute: total GPU-hours estimate
Step 3: Parse Ablation Plan

Normalize Codex response into structured format:

markdown
## Ablation Plan

### Component Ablations (highest priority)
| # | Name | What It Tests | Expected If Matters | Priority |
|---|------|---------------|---------------------|----------|
| 1 | remove module X | contribution of X | performance drops on metric Y | 1 |
| 2 | replace X with simpler Z | value of learned vs fixed | drops, especially on dataset A | 2 |

### Hyperparameter Sensitivity
| # | Parameter | Values to Test | What It Tests | Priority |
|---|-----------|---------------|---------------|----------|
| 3 | lambda | [0.01, 0.1, 1.0] | sensitivity to regularization | 3 |

### Design Choice Comparisons
| # | Name | What It Tests | Priority |
|---|------|---------------|----------|
| 4 | joint vs separate matching | whether joint adds value | 4 |

### Coverage Assessment
[What reviewer questions these ablations answer]

### Unnecessary Ablations
[Experiments that seem useful but won't add insight — skip these]

### Run Order
[Optimized for maximum early information]

### Estimated Compute
[Total GPU-hours]
Step 4: CC Reviews Feasibility

Before running anything, CC checks:

  • Compute budget: can we afford all ablations with available GPUs?
  • Code changes: which ablations need code modifications vs config-only changes?
  • Dependencies: which ablations can run in parallel?
  • Cuts: if budget is tight, propose removing lower-priority ablations and ask Codex to confirm
Step 5: Implement and Run
  1. Create configs/scripts for each ablation (config-only changes first)
  2. Smoke test each ablation before full run
  3. Run in suggested order, using descriptive names (e.g., ablation-no-module-X)
  4. Track results in EXPERIMENT_LOG.md
  5. After all ablations complete → update findings.md with insights

Rules

  • Codex leads the design. CC does not pre-filter or bias the ablation list before Codex sees it. Codex thinks like a reviewer; CC thinks like an engineer.
  • Every ablation must have a clear what_it_tests and expected_if_component_matters. No "just try it" experiments.
  • Config-only ablations take priority over those needing code changes (faster, less error-prone).
  • If total compute exceeds budget, CC proposes cuts and asks Codex to re-prioritize — don't silently drop ablations.
  • Component ablations (remove/replace) take priority over hyperparameter sweeps.
  • Do not generate ablations for components identical to the baseline (no-op ablations).
  • Record all ablation results in EXPERIMENT_LOG.md, including negative results (component removal had no effect = important finding).

© wanshuiyin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ablation-planner of wanshuiyin/Auto-claude-code-research-in-sleep.

Open the folder on GitHubat commit 26b95cf

Compare with similar skills

Ablation Study Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ablation Study Planner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ablation Study Planner this skillwanshuiyin/Auto-claude-code-research-in-sleep17k—~1.3kAutomated safety check: NotesMIT
Light Experiment CodingLight0305/Light-skills640—~2.3kAutomated safety check: PassMIT
Light Research PlanLight0305/Light-skills640—~5.3kAutomated safety check: PassMIT
Scientific Workflow ToolsDrugClaw/DrugClaw126—~712Automated safety check: PassApache-2.0
Bio Experimental Design Multiple TestingGPTomics/bioSkills1.2k1 repos~3.5kAutomated safety check: PassMIT
Scientific Critical Thinkingjaechang-hits/SciAgent-Skills3741 repos~4.7kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Light Experiment Coding

    Light0305/Light-skills

    Builds the code for a frozen research experiment test-first, with leakage controls, seed handling and saved evidence so results can be rerun and audited.

    640 GitHub stars~2.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Light Research Plan

    Light0305/Light-skills

    Light 科研主线第 5 步·研究方案与实验设计:把 idea-critique 放行的 idea 与 data feasibility 拆成能真执行、能写进论文、 能复现的 question/estimand、实验矩阵与预注册包。何时用:idea 已通过审查要落地 / 要设计实验·消融·对比·敏感性· 泛化·鲁棒性 / 写研究方案 PROJECTPLAN / 锁 primary…

    640 GitHub stars~5.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Scientific Workflow Tools

    DrugClaw/DrugClaw

    Research-method workflow guide for hypothesis framing, peer-review style critique, reproducibility planning, study-design checks, and scientific-writing structure.

    126 GitHub stars~712 tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Controls error rates across thousands of simultaneous tests in genomics discovery using false-discovery-rate methods (Benjamini-Hochberg 1995; Benjamini-Yekutieli 2001 for arbitrary dependence…

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Scientific Critical Thinking

    jaechang-hits/SciAgent-Skills

    Evaluating scientific evidence and claims. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.7k tokens
    Research & ScienceAuto-check passed
  • Jop Research Design

    franklee16/academic-research-skills

    A skill your agent uses when defending the research design of a The Journal of Politics (JOP) manuscript — causal identification for quantitative work, experimental and survey-experimental design…

    223 GitHub starsUsed in 1 repo~1.1k tokens
    Research & ScienceAuto-check passed

More from wanshuiyin/Auto-claude-code-research-in-sleep

All 26 skills in this repo
  • Academic Poster Builder

    wanshuiyin/Auto-claude-code-research-in-sleep

    Builds an academic conference poster as a single HTML and CSS file with measurement-based gates, real paper figures and a print-ready PDF rendered through headless Chromium.

    17k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check: notes
  • Proof Run Orchestrator

    wanshuiyin/Auto-claude-code-research-in-sleep

    Runs a mathematical proof project as a stateful pipeline of run directories: a local attempt first, then a manual GPT Pro handoff package, with an optional DeepSeek audit.

    17k GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Render HTML

    wanshuiyin/Auto-claude-code-research-in-sleep

    Render an ARIS Markdown / JSON artifact (IDEAREPORT, AUTOREVIEW, KILLARGUMENT, PAPERPLAN, research-wiki state, etc.) into a single-file HTML view designed for human reading.

    17k GitHub starsUsed in 1 repo~5.4k tokens
    Auto-check: notes
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check: notes
  • Integrity Forensics

    wanshuiyin/Auto-claude-code-research-in-sleep

    Run the Anti-Autoresearch integrity-forensics DETERMINISTIC slice (numeric core + rules-only reporter) against a paper via a SHA-pinned thin launcher, then convert the verdict into a typed policy…

    17k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Interview Cheatsheet

    wanshuiyin/Auto-claude-code-research-in-sleep

    Generate a long-form Chinese interview-prep cheat sheet on a specific ML/LLM topic — formulas with derivations, from-scratch PyTorch code, comparison tables, and 25 高频面试题 (L1 必会 / L2 进阶 / L3 顶级 lab).

    17k GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check: notes

Questions about Ablation Study Planner

What does Ablation Study Planner do?

Plans the ablation studies a paper needs after main results support its claim: Codex designs them like a reviewer, Claude Code checks feasibility and runs them. This skill plans the ablation studies a paper will need once main results support the claim. It splits the work between two models: Codex designs the ablations from a reviewer's point of view, and Claude Code checks feasibility and does the implementation.

When should I use Ablation Study Planner?

Ablation Study Planner fits situations like: planning ablations after main results support the paper's claim; responding to a reviewer-style request for missing ablations; fitting an ablation list to a limited GPU budget.

How do I install Ablation Study Planner in Claude Code?

Run `npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner -a claude-code`. Or copy the skill folder (skills/ablation-planner in wanshuiyin/Auto-claude-code-research-in-sleep) into .claude/skills/ablation-planner in your project. Claude Code loads it when a task matches its description.

How do I install Ablation Study Planner in Codex?

Run `npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner -a codex`. Or copy the skill folder (skills/ablation-planner in wanshuiyin/Auto-claude-code-research-in-sleep) into .agents/skills/ablation-planner in your project. Codex loads it when a task matches its description.

Can I use Ablation Study Planner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill ablation-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ablation-planner, .gemini/skills/ablation-planner, .github/skills/ablation-planner and .opencode/skills/ablation-planner in your project.

What does Ablation Study Planner need to run?

SKILL.md names no scripts, command-line tools or credentials: Ablation Study Planner is instructions for the agent only. Our summary lists: Codex reachable through the MCP tool the skill calls; Project files such as the research contract and the experiment log. Its frontmatter pre-approves these tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply.

Does Ablation Study Planner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ablation Study Planner safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Ablation Study Planner use?

Ablation Study Planner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ablation Study Planner use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ablation Study Planner?

Skills that share tags, products or a category with Ablation Study Planner: Light Experiment Coding (Light0305/Light-skills, 640 stars), Light Research Plan (Light0305/Light-skills, 640 stars), Scientific Workflow Tools (DrugClaw/DrugClaw, 126 stars) and Bio Experimental Design Multiple Testing (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ablation Study Planner?

wanshuiyin (a GitHub user) maintains it in wanshuiyin/Auto-claude-code-research-in-sleep, which has 17,205 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on October 7, 2026.

Source: wanshuiyin/Auto-claude-code-research-in-sleep on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.