Agent skill

Experiment Suite

by ai4s-research in ai4s-research/ai4s-skills

A skill your agent uses when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance)…

MITAuto-check passedResearch & Science

Install Experiment Suite

skills CLI
$ npx skills add ai4s-research/ai4s-skills --skill experiment-suite -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai4s-research/ai4s-skills experiment-suite --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai4s-research/ai4s-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment-suite .claude/skills/experiment-suite && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-suite
GitHub stars
237
Used in
2 other repos
Token cost
~2.5k tokens
SKILL.md length
1,055 words
Files
18 (incl. references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance)…

  • Works in 4 steps: Understand the question and operating mode → Set up the run directory → Build the package (REQUIRED — this is… → …
  • The user has a research question and needs a complete experiment package — design document
  • SKILL.md covers Overview, When to Use, When NOT to Use and Workflow, plus 2 more sections
  • Runs Python scripts from its folder; calls python and python3

What it does

Experiment Suite is an agent skill from ai4s-research/ai4s-skills. Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including reference files (for example `figure_examples/FIGURE_CONTRACT_TEMPLATE.md`, `figure_examples/README.md` and `figure_examples/make_fig_02_horizon_sweep.py`).

It sits in Research & Science, covering Hypothesis generation. It works with Python. The repository describes itself as: Open-source agent skills for AI for Science: topic exploration, literature survey, experiments, paper writing, and integrity audit — driven by any coding agent. The licence is MIT.

When your agent uses it

  • The user has a research question and needs a complete experiment package — design document
  • Results (measured
  • Simulated with honest provenance)
  • Publication-grade figures

Example prompts

  • “/experiment-suite”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Understand the question and operating mode
  2. Set up the run directory
  3. Build the package (REQUIRED — this is the whole job)
  4. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit 744ab20. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Suite loads about 2.5k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 1,055 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~16k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai4s-research/ai4s-skills at commit 744ab20, republished under its MIT licence (© ai4s-research). 1,055 words, ~2,525 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-suite/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
experiment-suite
description
Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report. Single-stage, no Python runtime.

Experiment Suite

Overview

End-to-end experiment package builder. Single stage, full quality from the start. The agent (Claude Code / Cursor / Aider / Codex / …) writes everything directly using its own tools (Write, Bash, WebFetch, …). This skill contains procedure + reference playbooks + figure-example scripts — no Python runtime, no LLM SDK.

The substantive work is decomposed into reference playbooks under references/:

ReferenceTopic
references/00-incremental-execution.mdhow to do this without losing work: batches, persistence, resume — read first
references/01-design-depth.mdwhat a real experiment design contains (motivation → hypothesis → datasets → baselines → metrics → ablations → budget)
references/01a-data-contract.mdruntime dataset binding: source, access route, version, split, and reuse boundary
references/02-code-quality.mdcode-skeleton standards — runnable model.py, data.py, train.py, evaluate.py
references/03-results-protocol.mdresults.json schema; measured / simulated / illustrative provenance
references/04-publication-figures.mdpublication-grade charts, multi-panel layouts, taste rules
references/04a-figure-contract.mdfigure logic before plotting: conclusion, panel map, reviewer risk
references/04b-figure-qa.mdexport bundle, editable text, statistics and image-integrity QA
references/05-report-structure.mdstructured experiment_report.md (problem → design → method → results → analysis → limitations)
references/06-quality-gate.mdself-check before delivery

Also: figure_examples/ — publication-style matplotlib scripts plus a shared style kit the agent can use as starting points.

Read the relevant reference before writing, not after. The full pass does not fit in a single turn — references/00-incremental-execution.md is the only execution mode that completes.

When to Use

  • User wants to "design an experiment" for a research question.
  • User needs runnable code for a specific task (classification / forecasting / detection / …).
  • User wants to compare methods and have a structured report at the end.
  • User needs publication-quality figures of experimental results.

When NOT to Use

  • User only wants a quick code snippet (write code directly).
  • User wants a full paper → paper-writer.
  • User wants a literature survey → literature-survey.

Workflow

Step 1 — Understand the question and operating mode

Confirm with the user:

  • Research question — what we are trying to answer.
  • Task type — classification / regression / forecasting / detection / generation / …
  • Mode
    • measured — user has real data or will run code themselves; provide a path to a measured results.json or run train.py against real data later.
    • simulated (default) — agent generates a plausible-shaped, deterministic results.json as a placeholder. Every figure/table caption must say "simulated".
  • Framework preference — PyTorch (default), JAX, TensorFlow, or sklearn.
  • Compute budget — hours / GPUs available; constrains the code skeleton and hyperparameter plan.

If the user has data and time, push toward measured mode. If not, simulated is acceptable provided disclosures are honest in every artefact.

Step 2 — Set up the run directory
bash
QUESTION="<research_question>"
SLUG=$(python3 -c "import re,hashlib,sys; t=sys.argv[1]; n=re.sub(r'[\\s_]+','-',re.sub(r'[^\\w\\s-]','',t.lower().strip())).strip('-')[:40].rstrip('-'); h=hashlib.sha1(t.encode()).hexdigest()[:8]; print(f'{n}-{h}')" "$QUESTION")
TS=$(date +%Y-%m-%d_%H%M%S)
RUN=output/experiment-suite/$SLUG/$TS

mkdir -p "$RUN/experiment" "$RUN/figures"
ln -sfn "$TS" "output/experiment-suite/$SLUG/latest"

In commands below $RUN = output/experiment-suite/<slug>/latest.

The agent will create five top-level files inside $RUN/:

  • experiment_design.md
  • data_contract.md
  • experiment/{model.py,data.py,train.py,evaluate.py,config.yaml,requirements.txt,README.md}
  • results.json
  • figures/*.pdf plus their make_*.py source and a manifest.json
  • experiment_report.md
Step 3 — Build the package (REQUIRED — this is the whole job)

Open references/00-incremental-execution.md first. Then carry out the six tracks below across many turns, persisting state to $RUN/ after every batch.

3.1 Design — full justification document

Open: references/01-design-depth.md and references/01a-data-contract.md. First write $RUN/data_contract.md as the dataset contract for this run. It must say whether the data are user-supplied, agent-discovered, reused public, controlled, or synthetic fallback. Then write $RUN/experiment_design.md as a real design (≥ 700 words): motivation → hypothesis → datasets → baselines → metrics → ablations → compute budget. Justify every choice.

3.2 Code — actually runnable

Open: references/02-code-quality.md. Fill $RUN/experiment/ with code that an engineer could launch with python train.py --config config.yaml. Real (if minimal) model class, real data loader, real train loop, real eval. The generated data.py and config.yaml are runtime products of this run and should bind to $RUN/data_contract.md, not to a repository-wide hard-coded benchmark. Add a README.md with run instructions.

3.3 Results — honest provenance

Open: references/03-results-protocol.md. Produce $RUN/results.json with a well-formed schema: per-seed entries, per-method per-metric mean & std, ablation block, and a provenance field that names the source.

  • measured mode — the user runs experiment/train.py (or supplies a results JSON) and the agent loads it into $RUN/results.json, setting "simulated": false and "provenance": "loaded from <path>".
  • measured mode, agent-discovered data — the agent may search for and bind an open dataset itself, but the chosen source, split, and access route must first be written into $RUN/data_contract.md; results.json provenance must point back to that binding.
  • simulated mode — the agent writes a deterministic seeded JSON of plausible shape, setting "simulated": true.
Show full SKILL.md (393 more words)Show less
3.4 Figures — 3–6 publication-grade

Open: references/04-publication-figures.md, references/04a-figure-contract.md, references/04b-figure-qa.md, and figure_examples/. Before writing plotting code, define the figure contract in a small working note under $RUN/figures/figure_contract.md:

  • one-sentence conclusion,
  • figure archetype,
  • panel map,
  • evidence hierarchy,
  • statistics needed,
  • reviewer risk.

Then plan and generate at minimum:

  • 1 method-comparison chart (bar or line).
  • 1 ablation breakdown.
  • Optionally training curves, scaling plot, heatmap.

Save each figure into $RUN/figures/<basename>.pdf with its make_*.py source alongside. Prefer saving an editable .svg and print-grade .tiff alongside the PDF when the environment supports it. Append entries to $RUN/figures/manifest.json storing basenames only (never absolute paths) so paper-writer can copy them in directly. Apply the shared publication style (embedded fonts, explicit palette, panel labels, simulated watermark when applicable).

If simulated, watermark the figures or always note "simulated" in their captions in the report.

3.5 Report — structured

Open: references/05-report-structure.md. Write $RUN/experiment_report.md with sections: problem statement → design rationale → method → setup → results → analysis → limitations. Reference figures by filename. This report is the primary deliverable for users who want only the experiment package (no paper-writer follow-up).

3.6 Quality gate

Open: references/06-quality-gate.md. Targets: design ≥ 700 words, code imports cleanly (python -c "import experiment.model" from inside $RUN), results.json passes schema check, ≥ 3 figures, report ≥ 6 sections.

Step 4 — Deliver

Report:

  1. output/experiment-suite/<slug>/latest/experiment_design.md
  2. output/experiment-suite/<slug>/latest/experiment/ — runnable code package.
  3. output/experiment-suite/<slug>/latest/results.json — with provenance.
  4. output/experiment-suite/<slug>/latest/figures/ — publication-grade charts + manifest.json.
  5. output/experiment-suite/<slug>/latest/experiment_report.md — structured report.
  6. Stats per the report format in references/06-quality-gate.md.

Cross-skill data flow (path convention)

The paper-writer skill computing the same slug for the same topic will look here:

  • output/experiment-suite/<slug>/latest/results.json — source of the numbers and the "simulated" flag (drives the disclosure clause in the paper).
  • output/experiment-suite/<slug>/latest/figures/*.pdf (+ manifest.json) — figures to reuse rather than redraw.

Always store basenames in manifest.json. Absolute paths in the manifest break paper-writer's \includegraphics{figures/<basename>}.

Important rules

  • No LLM SDK in this skill. No import anthropic / import openai. The skill is SKILL.md + references + figure examples only.
  • Simulated results must always remain visibly labelled — in results.json ("simulated": true), in figure captions, in the report's top-of-page disclosure, and in any downstream paper's \thanks footnote.
  • Never present simulated results as measured. When in doubt, treat as simulated and disclose.
  • The runnable code is a starting point, not a SOTA reproduction. Be honest about its scope in experiment/README.md.
  • A real experiment package would normally take days of compute for real numbers; the simulated path lets the workflow proceed when that's not possible, with honest disclosure throughout.

© ai4s-research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (references) in skills/experiment-suite of ai4s-research/ai4s-skills.

  • SKILL.md
  • figure_examples/FIGURE_CONTRACT_TEMPLATE.md
  • figure_examples/README.md
  • figure_examples/make_fig_02_horizon_sweep.py
  • figure_examples/make_fig_03_heatmap.py
  • figure_examples/make_fig_04_ablation.py
  • figure_examples/requirements.txt
  • figure_examples/style_kit.py
  • references/00-incremental-execution.md
  • references/01-design-depth.md
  • references/01a-data-contract.md
  • references/02-code-quality.md
  • references/03-results-protocol.md
  • references/04-publication-figures.md
  • references/04a-figure-contract.md
  • references/04b-figure-qa.md
  • references/05-report-structure.md
  • references/06-quality-gate.md

Open the folder on GitHubat commit 744ab20

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in ai4s-research/ai4s-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Experiment Suite next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Suite compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Suite this skillai4s-research/ai4s-skills2372 repos~2.5kAutomated safety check: PassMIT
High Stakes Analytics Decision Lablimingrui679-design/high-stakes-analytics-decision-lab1k—~2.2kAutomated safety check: PassMIT
HypoGeniC Hypothesis GenerationK-Dense-AI/scientific-agent-skills48k1 repos~3.6kAutomated safety check: NotesMIT
News to Research Idea BriefingOpenLAIR/dr-claw1.2k—~1.3kAutomated safety check: NotesCustom licence
Hypothesis Generationspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow84k4 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • High Stakes Analytics Decision Lab

    limingrui679-design/high-stakes-analytics-decision-lab

    Build or review source-backed descriptive, diagnostic, predictive, and prescriptive analysis for consequential decisions.

    1k GitHub stars~2.2k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • HypoGeniC Hypothesis Generation

    K-Dense-AI/scientific-agent-skills

    Plans and audits runs of the HypoGeniC and HypoRefine packages, which propose hypotheses from labeled text datasets, with local checks before any model call.

    48k GitHub starsUsed in 1 repo~3.6k tokens
    Research & ScienceAuto-check: notes
  • Clusters the latest news-feed results by topic and writes a briefing of research idea seeds with citations, plus a structured seeds file, without crawling new sources.

    1.2k GitHub stars~1.3k tokensUpdated 23 days ago
    Research & ScienceAuto-check: notes
  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    47k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed

More from ai4s-research/ai4s-skills

  • Literature Survey

    ai4s-research/ai4s-skills

    A skill your agent uses when the user wants a comprehensive literature survey on a specific research topic.

    237 GitHub starsUsed in 2 repos~2k tokens
    Auto-check passed
  • Integrity Auditor

    ai4s-research/ai4s-skills

    A skill your agent uses when the user wants a paper audited for integrity issues — image misuse, numerical anomalies, logical gaps — and needs a reviewable evidence report.

    237 GitHub starsUsed in 1 repo~5k tokens
    Auto-check passed
  • Paper Writer

    ai4s-research/ai4s-skills

    A skill your agent uses when the user wants a complete, publication-grade research paper on a specific topic — produces 200+ real citations, 4–8 publication-grade figures, and 7 sections of…

    237 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Research Explorer

    ai4s-research/ai4s-skills

    A skill your agent uses when the user has a vague research direction and wants to explore feasible specific topics.

    237 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Ai4s Agent

    ai4s-research/ai4s-skills

    A skill your agent uses when the user wants an end-to-end AI4S research pipeline — broad direction or specific topic in, full research package out (exploration + literature survey + experiment +…

    237 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed
  • Mindmap Render

    ai4s-research/ai4s-skills

    Generate beautiful, high-resolution mindmaps from Markdown unordered lists.

    237 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed

Works with

Questions about Experiment Suite

What does Experiment Suite do?

A skill your agent uses when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance)…. Experiment Suite is an agent skill from ai4s-research/ai4s-skills. Use when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance), publication-grade figures, structured report.

When should I use Experiment Suite?

Experiment Suite fits situations like: the user has a research question and needs a complete experiment package — design document; results (measured; simulated with honest provenance); publication-grade figures.

How do I install Experiment Suite in Claude Code?

Run `npx skills add ai4s-research/ai4s-skills --skill experiment-suite -a claude-code`. Or copy the skill folder (skills/experiment-suite in ai4s-research/ai4s-skills) into .claude/skills/experiment-suite in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Suite in Codex?

Run `npx skills add ai4s-research/ai4s-skills --skill experiment-suite -a codex`. Or copy the skill folder (skills/experiment-suite in ai4s-research/ai4s-skills) into .agents/skills/experiment-suite in your project. Codex loads it when a task matches its description.

Can I use Experiment Suite in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai4s-research/ai4s-skills --skill experiment-suite -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-suite, .gemini/skills/experiment-suite, .github/skills/experiment-suite and .opencode/skills/experiment-suite in your project.

What does Experiment Suite need to run?

Going by SKILL.md and its folder, Experiment Suite needs Python for the scripts in its folder and the command-line tools its instructions call (python and python3). Our summary lists: Python 3.

Does Experiment Suite access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Suite safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Suite use?

Experiment Suite is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Suite use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Experiment Suite?

Skills that share tags, products or a category with Experiment Suite: High Stakes Analytics Decision Lab (limingrui679-design/high-stakes-analytics-decision-lab, 1k stars), HypoGeniC Hypothesis Generation (K-Dense-AI/scientific-agent-skills, 48k stars), News to Research Idea Briefing (OpenLAIR/dr-claw, 1.2k stars) and Hypothesis Generation (spacering-net/codeg, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Suite?

ai4s-research (a GitHub organization) maintains it in ai4s-research/ai4s-skills, which has 237 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on July 28, 2026.

Source: ai4s-research/ai4s-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.