Agent skill

Measure Experiment Design

by product-on-purpose in product-on-purpose/pm-skills

Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis.

Apache-2.0Auto-check passedResearch & Science

Install Measure Experiment Design

skills CLI
$ npx skills add product-on-purpose/pm-skills --skill measure-experiment-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install product-on-purpose/pm-skills measure-experiment-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/product-on-purpose/pm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/measure-experiment-design .claude/skills/measure-experiment-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
measure-experiment-design
GitHub stars
716
Token cost
~1.1k tokens
SKILL.md length
493 words
Files
7 (incl. references)
Skills in repo
68
Repo updated
First seen
Licence
Apache-2.0

At a glance

Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis.

  • Works in 8 steps: Articulate the Hypothesis → Define the Variants → Choose Primary and Secondary Metrics → …
  • Planning an experiment to validate a product change
  • SKILL.md covers When to Use, When NOT to Use, Instructions and Output Format, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Measure Experiment Design is an agent skill from product-on-purpose/pm-skills. Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis. Use when planning an experiment to validate a product change or test an assumption you have already framed. To articulate the hypothesis itself first, use define-hypothesis.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `HISTORY.md`, `evals/output-scenarios/onboarding-checklist.md` and `evals/output-scenarios/paywall-pricing.md`).

It sits in Research & Science, covering Experimental design, A/B testing and Product metrics. The repository describes itself as: 68 plug-and-play, best-practice product management skills for AI agents: 30 Triple Diamond phase + 11 foundation + 12 utility + 15 tool (Foundation Sprint + Design Sprint). Plus… The licence is Apache-2.0.

When your agent uses it

  • Planning an experiment to validate a product change
  • Test an assumption you have already framed

Example prompts

  • “Use the measure-experiment-design skill to design an A/B test or experiment with variants, success metrics, sample size, and duration for an…”
  • “/measure-experiment-design”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Articulate the Hypothesis
  2. Define the Variants
  3. Choose Primary and Secondary Metrics
  4. Calculate Sample Size
  5. Estimate Duration
  6. Define Targeting and Allocation
  7. Set Success Criteria
  8. Document Risks and Mitigations

What it can do on your machine

Read from SKILL.md and the folder at commit 1cef1a9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Measure Experiment Design loads about 1.1k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 493 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from product-on-purpose/pm-skills at commit 1cef1a9, republished under its Apache-2.0 licence (© product-on-purpose). 493 words, ~1,083 tokens.

Download SKILL.mdSave it as .claude/skills/measure-experiment-design/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
measure-experiment-design
description
Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis. Use when planning an experiment to validate a product change or test an assumption you have already framed. To articulate the hypothesis itself first, use define-hypothesis.
license
Apache-2.0
metadata.phase
measure
metadata.version
2.1.0
metadata.updated
2026-06-10
metadata.category
validation
metadata.frameworks
triple-diamond, lean-startup, design-thinking
metadata.author
product-on-purpose
<!-- PM-Skills | https://github.com/product-on-purpose/pm-skills | Apache 2.0 -->

Experiment Design

An experiment design document defines all parameters needed to run a rigorous A/B test or controlled experiment. It ensures the team aligns on what you're testing, how you'll measure success, and how long to run the test before drawing conclusions. Good experiment design prevents common pitfalls: underpowered tests, unclear success criteria, and decisions based on noise rather than signal.

When to Use

  • Before launching an A/B test to validate a product change
  • When testing a hypothesis that requires quantitative validation
  • After solution design to validate assumptions before full rollout
  • When stakeholders want data-driven evidence for a decision
  • To establish a culture of experimentation and learning

When NOT to Use

  • The hypothesis itself is not yet articulated -> use define-hypothesis first; this skill designs the test for a claim you already have
  • You are analyzing a completed experiment -> use measure-experiment-results
  • You need the event tracking that will measure the experiment -> use measure-instrumentation-spec
  • You are gathering opinions rather than running a controlled test -> use measure-survey-analysis

Instructions

When asked to design an experiment, follow these steps:

  1. Articulate the Hypothesis Write a clear, testable hypothesis in the format: "We believe [change] for [users] will [outcome] as measured by [metric]." One hypothesis per experiment - if you're testing multiple things, run multiple experiments.

  2. Define the Variants Describe the control (current experience) and treatment (new experience) in sufficient detail. Include screenshots, mockups, or precise descriptions so anyone can understand what users will see.

  3. Choose Primary and Secondary Metrics Select one primary metric that will determine success or failure. Add 2-3 secondary metrics to understand the broader impact. Include guardrail metrics to catch unintended negative effects.

  4. Calculate Sample Size Determine how many users you need per variant to detect your minimum detectable effect (MDE) with statistical significance. Specify your significance level (typically 0.05) and power (typically 0.80).

  5. Estimate Duration Based on sample size and available traffic, calculate how long the experiment needs to run. Account for weekly patterns - avoid ending mid-week if behavior varies by day.

  6. Define Targeting and Allocation Specify which users are eligible for the experiment and how traffic is split between variants. Document any exclusions (e.g., employees, specific segments).

  7. Set Success Criteria Define upfront what constitutes a win, a loss, or an inconclusive result. This prevents post-hoc rationalization and moving goalposts.

  8. Document Risks and Mitigations Identify what could go wrong and how you'll detect/address it. Include monitoring plans and rollback criteria.

Show full SKILL.md (88 more words)Show less

Output Format

Use the template in references/TEMPLATE.md to structure the output. A complete design fills every template section: Overview; Hypothesis; Background; Variants; Metrics; Sample Size & Duration; Audience Targeting; Success Criteria; Risks & Mitigations; Implementation Notes; and References.

Quality Checklist

Before finalizing, verify:

  • Hypothesis is falsifiable and specific
  • Only one primary metric is defined
  • Sample size calculation is documented with assumptions
  • Duration accounts for traffic patterns and statistical requirements
  • Success criteria are defined before the experiment starts
  • Guardrail metrics protect against unintended harm

Examples

See references/EXAMPLE.md for a completed example.

© product-on-purpose, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in skills/measure-experiment-design of product-on-purpose/pm-skills.

  • SKILL.md
  • HISTORY.md
  • evals/output-scenarios/onboarding-checklist.md
  • evals/output-scenarios/paywall-pricing.md
  • evals/trigger-fixtures.json
  • references/EXAMPLE.md
  • references/TEMPLATE.md

Open the folder on GitHubat commit 1cef1a9

Compare with similar skills

Measure Experiment Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Measure Experiment Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Measure Experiment Design this skillproduct-on-purpose/pm-skills716—~1.1kAutomated safety check: PassApache-2.0
Craft Experiment Designamplitude/builder-skills159—~522Automated safety check: PassNone
Data Scientistmagnus919/hermes-profiles289—~3.3kAutomated safety check: PassMIT
Ab Testingericrisco/rsc-harness190—~2.4kAutomated safety check: PassMIT
Data Scientistmagnus919/agent-skills119—~4.1kAutomated safety check: PassMIT
Hypothesis Testerjeremylongshore/tons-of-skills-marketplace2.8k—~2.6kAutomated safety check: PassMIT

Similar skills

  • Craft Experiment Design

    amplitude/builder-skills

    Write a hypothesis, define success metrics, and plan a holdout strategy.

    159 GitHub stars~522 tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Data Scientist

    magnus919/hermes-profiles

    PhD-level expertise in data science, statistics, and machine learning.

    289 GitHub stars~3.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    190 GitHub stars~2.4k tokensUpdated today
    Marketing & SEOAuto-check passed
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    119 GitHub stars~4.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Hypothesis Tester

    jeremylongshore/tons-of-skills-marketplace

    Structured hypothesis formulation, experiment design, and results interpretation for Product Managers.

    2.8k GitHub stars~2.6k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Experiment Designer

    mohitagw15856/pm-claude-skills

    Design statistically rigorous A/B tests and interpret experiment results.

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from product-on-purpose/pm-skills

All 68 skills in this repo
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    716 GitHub stars~966 tokensUpdated 3 days ago
    Auto-check passed
  • Define Jtbd Canvas

    product-on-purpose/pm-skills

    Creates a Jobs to be Done canvas capturing the functional, emotional, and social dimensions of a customer job.

    716 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Define Opportunity Tree

    product-on-purpose/pm-skills

    Creates an opportunity solution tree connecting a desired outcome to customer opportunities and candidate solutions, preventing solution-first jumps in continuous discovery.

    716 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Define Problem Statement

    product-on-purpose/pm-skills

    Creates a clear problem framing document with user impact, business context, and success criteria.

    716 GitHub stars~932 tokensUpdated 3 days ago
    Auto-check passed
  • Deliver Acceptance Criteria

    product-on-purpose/pm-skills

    Generates structured Given/When/Then acceptance criteria for a user story or feature slice, covering the happy path, key failure scenarios, and non-functional expectations in testable form.

    716 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Deliver Launch Checklist

    product-on-purpose/pm-skills

    Creates a cross-functional pre-launch checklist covering engineering, design, marketing, support, legal, and operations readiness, with owners, dates, and go/no-go criteria so nothing is missed…

    716 GitHub stars~970 tokensUpdated 3 days ago
    Auto-check passed

Questions about Measure Experiment Design

What does Measure Experiment Design do?

Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis. Measure Experiment Design is an agent skill from product-on-purpose/pm-skills. Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis.

When should I use Measure Experiment Design?

Measure Experiment Design fits situations like: planning an experiment to validate a product change; test an assumption you have already framed.

How do I install Measure Experiment Design in Claude Code?

Run `npx skills add product-on-purpose/pm-skills --skill measure-experiment-design -a claude-code`. Or copy the skill folder (skills/measure-experiment-design in product-on-purpose/pm-skills) into .claude/skills/measure-experiment-design in your project. Claude Code loads it when a task matches its description.

How do I install Measure Experiment Design in Codex?

Run `npx skills add product-on-purpose/pm-skills --skill measure-experiment-design -a codex`. Or copy the skill folder (skills/measure-experiment-design in product-on-purpose/pm-skills) into .agents/skills/measure-experiment-design in your project. Codex loads it when a task matches its description.

Can I use Measure Experiment Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add product-on-purpose/pm-skills --skill measure-experiment-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/measure-experiment-design, .gemini/skills/measure-experiment-design, .github/skills/measure-experiment-design and .opencode/skills/measure-experiment-design in your project.

What does Measure Experiment Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Measure Experiment Design is instructions for the agent only.

Does Measure Experiment Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Measure Experiment Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Measure Experiment Design use?

Measure Experiment Design is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Measure Experiment Design use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Measure Experiment Design?

Skills that share tags, products or a category with Measure Experiment Design: Craft Experiment Design (amplitude/builder-skills, 159 stars), Data Scientist (magnus919/hermes-profiles, 289 stars), Ab Testing (ericrisco/rsc-harness, 190 stars) and Data Scientist (magnus919/agent-skills, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Measure Experiment Design?

product-on-purpose (a GitHub organization) maintains it in product-on-purpose/pm-skills, which has 716 GitHub stars. The repository holds 68 skills in this directory. The repository was last updated on October 8, 2026.

Source: product-on-purpose/pm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.