Agent skill

Experiment Designer

by alirezarezvani in alirezarezvani/claude-skills

A skill your agent uses when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

MITAuto-check passedResearch & Science

Install Experiment Designer

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill experiment-designer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills experiment-designer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/product-team/skills/experiment-designer .claude/skills/experiment-designer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-designer
GitHub stars
28k
Used in
1 other repo
Token cost
~783 tokens
SKILL.md length
300 words
Files
4 (incl. scripts, references)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

  • Planning product experiments
  • SKILL.md covers When To Use, Core Workflow, Hypothesis Quality Checklist and Common Experiment Pitfalls, plus 2 more sections
  • Runs Python scripts from its folder; calls python3
  • Writing testable hypotheses

What it does

Experiment Designer is an agent skill from alirezarezvani/claude-skills. Use when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

Its SKILL.md is about 780 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/experiment-playbook.md`, `references/statistics-reference.md` and `scripts/sample_size_calculator.py`).

It sits in Research & Science, covering Experimental design. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • Planning product experiments
  • Writing testable hypotheses
  • Estimating sample size
  • Prioritizing tests

Example prompts

  • “/experiment-designer”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Designer loads about 783 tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 300 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~783
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 300 words, ~783 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-designer/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
experiment-designer
description
Use when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

Experiment Designer

Design, prioritize, and evaluate product experiments with clear hypotheses and defensible decisions.

When To Use

Use this skill for:

  • A/B and multivariate experiment planning
  • Hypothesis writing and success criteria definition
  • Sample size and minimum detectable effect planning
  • Experiment prioritization with ICE scoring
  • Reading statistical output for product decisions

Core Workflow

  1. Write hypothesis in If/Then/Because format
  • If we change [intervention]
  • Then [metric] will change by [expected direction/magnitude]
  • Because [behavioral mechanism]
  1. Define metrics before running test
  • Primary metric: single decision metric
  • Guardrail metrics: quality/risk protection
  • Secondary metrics: diagnostics only
  1. Estimate sample size
  • Baseline conversion or baseline mean
  • Minimum detectable effect (MDE)
  • Significance level (alpha) and power

Use:

bash
python3 scripts/sample_size_calculator.py --baseline-rate 0.12 --mde 0.02 --mde-type absolute
  1. Prioritize experiments with ICE
  • Impact: potential upside
  • Confidence: evidence quality
  • Ease: cost/speed/complexity

ICE Score = (Impact * Confidence * Ease) / 10

  1. Launch with stopping rules
  • Decide fixed sample size or fixed duration in advance
  • Avoid repeated peeking without proper method
  • Monitor guardrails continuously
  1. Interpret results
  • Statistical significance is not business significance
  • Compare point estimate + confidence interval to decision threshold
  • Investigate novelty effects and segment heterogeneity

Hypothesis Quality Checklist

  • Contains explicit intervention and audience
  • Specifies measurable metric change
  • States plausible causal reason
  • Includes expected minimum effect
  • Defines failure condition

Common Experiment Pitfalls

  • Underpowered tests leading to false negatives
  • Running too many simultaneous changes without isolation
  • Changing targeting or implementation mid-test
  • Stopping early on random spikes
  • Ignoring sample ratio mismatch and instrumentation drift
  • Declaring success from p-value without effect-size context

Statistical Interpretation Guardrails

  • p-value < alpha indicates evidence against null, not guaranteed truth.
  • Confidence interval crossing zero/no-effect means uncertain directional claim.
  • Wide intervals imply low precision even when significant.
  • Use practical significance thresholds tied to business impact.

See:

  • references/experiment-playbook.md
  • references/statistics-reference.md

Tooling

scripts/sample_size_calculator.py

Computes required sample size (per variant and total) from:

  • baseline rate
  • MDE (absolute or relative)
  • significance level (alpha)
  • statistical power

Example:

bash
python3 scripts/sample_size_calculator.py \
  --baseline-rate 0.10 \
  --mde 0.015 \
  --mde-type absolute \
  --alpha 0.05 \
  --power 0.8

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in product-team/skills/experiment-designer of alirezarezvani/claude-skills.

  • SKILL.md
  • references/experiment-playbook.md
  • references/statistics-reference.md
  • scripts/sample_size_calculator.py

Open the folder on GitHubat commit 19392f7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in alirezarezvani/claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Experiment Designer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Designer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Designer this skillalirezarezvani/claude-skills28k1 repos~783Automated safety check: PassMIT
Ablation Study Plannerwanshuiyin/Auto-claude-code-research-in-sleep17k—~1.3kAutomated safety check: NotesMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k22 repos~5.9kAutomated safety check: NotesMIT
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.7k—~2.8kAutomated safety check: PassCC-BY-4.0
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1286 repos~2.3kAutomated safety check: NotesNone
Research Refine PipelinezjYao36/Auto-Research-Refine1285 repos~1.4kAutomated safety check: NotesNone

Similar skills

  • Ablation Study Planner

    wanshuiyin/Auto-claude-code-research-in-sleep

    Plans the ablation studies a paper needs after main results support its claim: Codex designs them like a reviewer, Claude Code checks feasibility and runs them.

    17k GitHub stars~1.3k tokensUpdated 2 days ago
    Research & ScienceAuto-check: notes
  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 22 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.7k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 6 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 5 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Questions about Experiment Designer

What does Experiment Designer do?

A skill your agent uses when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor. Experiment Designer is an agent skill from alirezarezvani/claude-skills. Use when planning product experiments, writing testable hypotheses, estimating sample size, prioritizing tests, or interpreting A/B outcomes with practical statistical rigor.

When should I use Experiment Designer?

Experiment Designer fits situations like: planning product experiments; writing testable hypotheses; estimating sample size; prioritizing tests.

How do I install Experiment Designer in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill experiment-designer -a claude-code`. Or copy the skill folder (product-team/skills/experiment-designer in alirezarezvani/claude-skills) into .claude/skills/experiment-designer in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Designer in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill experiment-designer -a codex`. Or copy the skill folder (product-team/skills/experiment-designer in alirezarezvani/claude-skills) into .agents/skills/experiment-designer in your project. Codex loads it when a task matches its description.

Can I use Experiment Designer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill experiment-designer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-designer, .gemini/skills/experiment-designer, .github/skills/experiment-designer and .opencode/skills/experiment-designer in your project.

What does Experiment Designer need to run?

Going by SKILL.md and its folder, Experiment Designer needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Experiment Designer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Designer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Experiment Designer use?

Experiment Designer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Designer use?

About 783 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 870 tokens, read only when the agent opens those files.

What are the alternatives to Experiment Designer?

Skills that share tags, products or a category with Experiment Designer: Ablation Study Planner (wanshuiyin/Auto-claude-code-research-in-sleep, 17k stars), Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.7k stars) and Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Designer?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,891 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.