Agent skill

Measure Experiment Results

by product-on-purpose in product-on-purpose/pm-skills

Documents the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations.

Apache-2.0Auto-check passedMarketing & SEO

Install Measure Experiment Results

skills CLI
$ npx skills add product-on-purpose/pm-skills --skill measure-experiment-results -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install product-on-purpose/pm-skills measure-experiment-results --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/product-on-purpose/pm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/measure-experiment-results .claude/skills/measure-experiment-results && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
measure-experiment-results
GitHub stars
716
Token cost
~989 tokens
SKILL.md length
434 words
Files
5 (incl. references)
Skills in repo
68
Repo updated
First seen
Licence
Apache-2.0

At a glance

Documents the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations.

  • Works in 8 steps: Summarize the Experiment → Restate the Hypothesis → Present Primary Results → …
  • Tasks that involve A/B testing
  • SKILL.md covers When to Use, When NOT to Use, Instructions and Output Format, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Measure Experiment Results is an agent skill from product-on-purpose/pm-skills. Documents the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations. Use after experiments conclude to communicate findings, inform decisions, and build organizational knowledge.

Its SKILL.md is about 990 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `HISTORY.md`, `evals/trigger-fixtures.json` and `references/EXAMPLE.md`).

It sits in Marketing & SEO, covering A/B testing and Statistics. The repository describes itself as: 68 plug-and-play, best-practice product management skills for AI agents: 30 Triple Diamond phase + 11 foundation + 12 utility + 15 tool (Foundation Sprint + Design Sprint). Plus… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve A/B testing
  • Tasks that involve Statistics

Example prompts

  • “Use the measure-experiment-results skill to document the results of a completed experiment or A/B test with statistical analysis, learnings, and…”
  • “/measure-experiment-results”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Summarize the Experiment
  2. Restate the Hypothesis
  3. Present Primary Results
  4. Analyze Secondary Metrics
  5. Segment the Data
  6. Extract Learnings
  7. Make a Recommendation
  8. Define Next Steps

What it can do on your machine

Read from SKILL.md and the folder at commit 1cef1a9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Measure Experiment Results loads about 989 tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 434 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~989
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from product-on-purpose/pm-skills at commit 1cef1a9, republished under its Apache-2.0 licence (© product-on-purpose). 434 words, ~989 tokens.

Download SKILL.mdSave it as .claude/skills/measure-experiment-results/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
measure-experiment-results
description
Documents the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations. Use after experiments conclude to communicate findings, inform decisions, and build organizational knowledge.
license
Apache-2.0
metadata.phase
measure
metadata.version
2.1.0
metadata.updated
2026-06-10
metadata.category
reflection
metadata.frameworks
triple-diamond, lean-startup, design-thinking
metadata.author
product-on-purpose
<!-- PM-Skills | https://github.com/product-on-purpose/pm-skills | Apache 2.0 -->

Experiment Results

An experiment results document captures what happened when you tested a hypothesis, including statistical outcomes, segment analysis, learnings, and clear recommendations. Good results documentation turns individual experiments into organizational knowledge that improves future decision-making.

When to Use

  • After an A/B test or experiment reaches statistical significance
  • When an experiment is ended early (for any reason)
  • To communicate findings to stakeholders who weren't involved
  • During decision-making about whether to ship, iterate, or kill a feature
  • To build a repository of learnings that inform future experiments

When NOT to Use

  • The experiment is not designed or run yet -> use measure-experiment-design
  • The results demand a direction decision -> use iterate-pivot-decision; this skill reports the evidence, that one decides
  • You want the transferable learning banked for the organization -> follow up with iterate-lessons-log
  • Your data is survey responses, not a controlled experiment -> use measure-survey-analysis

Instructions

When asked to document experiment results, follow these steps:

  1. Summarize the Experiment Provide context: what was tested, when it ran, how much traffic it received. Link to the original experiment design document if one exists.

  2. Restate the Hypothesis Remind readers what you believed would happen and why. This frames the results interpretation.

  3. Present Primary Results Show the primary metric outcome clearly: what were the values for control and treatment? Include statistical significance (p-value), confidence intervals, and sample sizes. Be honest about whether results are conclusive.

  4. Analyze Secondary Metrics Present guardrail metrics that ensure you didn't cause unintended harm. Note any secondary metrics that moved unexpectedly.both positive and negative.

  5. Segment the Data Look for differential effects across user segments (platform, tenure, plan type, etc.). Sometimes overall results mask important segment-level insights.

  6. Extract Learnings What did you learn beyond the numbers? Include surprising findings, questions raised, and implications for the product hypothesis. Negative results are valuable learnings.

  7. Make a Recommendation Be clear: should we ship, iterate, or kill? Support the recommendation with the evidence. If the decision is nuanced, explain the trade-offs.

  8. Define Next Steps Specify what happens now.engineering work to ship, follow-up experiments, metrics to continue monitoring, or documentation to update.

Show full SKILL.md (85 more words)Show less

Output Format

Use the template in references/TEMPLATE.md to structure the output. A complete readout fills every template section: Summary; Hypothesis Recap; Results; Segment Analysis; Visualization; Learnings; Recommendation; Next Steps; and Appendix.

Quality Checklist

Before finalizing, verify:

  • Statistical methods and significance are clearly stated
  • Confidence intervals are included (not just p-values)
  • Segment analysis checked for differential effects
  • Secondary/guardrail metrics are reported
  • Learnings go beyond just the numbers
  • Recommendation is clear and actionable
  • Negative or inconclusive results are reported honestly

Examples

See references/EXAMPLE.md for a completed example.

© product-on-purpose, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/measure-experiment-results of product-on-purpose/pm-skills.

  • SKILL.md
  • HISTORY.md
  • evals/trigger-fixtures.json
  • references/EXAMPLE.md
  • references/TEMPLATE.md

Open the folder on GitHubat commit 1cef1a9

Compare with similar skills

Measure Experiment Results next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Measure Experiment Results compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Measure Experiment Results this skillproduct-on-purpose/pm-skills716—~989Automated safety check: PassApache-2.0
Ab Test Analysisnimrodfisher/data-analytics-skills470—~708Automated safety check: PassMIT
Mkt Experimentevolution-foundation/evo-nexus545—~1.2kAutomated safety check: PassCustom licence
A/B Test Analysisphuryn/pm-skills27k—~893Automated safety check: PassMIT
Experimentationandreaskelm/pm-brain234—~2.3kAutomated safety check: PassCustom licence
Experimentai-analyst-lab/ai-analyst304—~2kAutomated safety check: PassMIT

Similar skills

  • Ab Test Analysis

    nimrodfisher/data-analytics-skills

    Rigorous A/B test statistical analysis. An agent skill from nimrodfisher/data-analytics-skills.

    470 GitHub stars~708 tokensUpdated 16 days ago
    Marketing & SEOAuto-check passed
  • Mkt Experiment

    evolution-foundation/evo-nexus

    Autonomous growth experimentation framework. An agent skill from evolution-foundation/evo-nexus.

    545 GitHub stars~1.2k tokensUpdated 5 mo ago
    Marketing & SEOAuto-check passed
  • A/B Test Analysis

    phuryn/pm-skills

    Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

    27k GitHub stars~893 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Experimentation

    andreaskelm/pm-brain

    Design and run product experiments at a practical PM level — A/B tests, hypothesis tests, rollouts, feature flags, and reading results without pretending to be a statistician.

    234 GitHub stars~2.3k tokensUpdated 2 days ago
    Marketing & SEOAuto-check passed
  • Experiment

    ai-analyst-lab/ai-analyst

    The analysis and lifecycle owner for experiments. An agent skill from ai-analyst-lab/ai-analyst.

    304 GitHub stars~2k tokensUpdated 10 days ago
    Marketing & SEOAuto-check passed
  • Statistical Analyst

    alirezarezvani/claude-skills

    Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

    28k GitHub stars~2.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from product-on-purpose/pm-skills

All 68 skills in this repo
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    716 GitHub stars~966 tokensUpdated 3 days ago
    Auto-check passed
  • Define Jtbd Canvas

    product-on-purpose/pm-skills

    Creates a Jobs to be Done canvas capturing the functional, emotional, and social dimensions of a customer job.

    716 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Define Opportunity Tree

    product-on-purpose/pm-skills

    Creates an opportunity solution tree connecting a desired outcome to customer opportunities and candidate solutions, preventing solution-first jumps in continuous discovery.

    716 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Define Problem Statement

    product-on-purpose/pm-skills

    Creates a clear problem framing document with user impact, business context, and success criteria.

    716 GitHub stars~932 tokensUpdated 3 days ago
    Auto-check passed
  • Deliver Acceptance Criteria

    product-on-purpose/pm-skills

    Generates structured Given/When/Then acceptance criteria for a user story or feature slice, covering the happy path, key failure scenarios, and non-functional expectations in testable form.

    716 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Deliver Launch Checklist

    product-on-purpose/pm-skills

    Creates a cross-functional pre-launch checklist covering engineering, design, marketing, support, legal, and operations readiness, with owners, dates, and go/no-go criteria so nothing is missed…

    716 GitHub stars~970 tokensUpdated 3 days ago
    Auto-check passed

Questions about Measure Experiment Results

What does Measure Experiment Results do?

Documents the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations. Measure Experiment Results is an agent skill from product-on-purpose/pm-skills. Documents the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations.

When should I use Measure Experiment Results?

Measure Experiment Results fits situations like: tasks that involve A/B testing; tasks that involve Statistics.

How do I install Measure Experiment Results in Claude Code?

Run `npx skills add product-on-purpose/pm-skills --skill measure-experiment-results -a claude-code`. Or copy the skill folder (skills/measure-experiment-results in product-on-purpose/pm-skills) into .claude/skills/measure-experiment-results in your project. Claude Code loads it when a task matches its description.

How do I install Measure Experiment Results in Codex?

Run `npx skills add product-on-purpose/pm-skills --skill measure-experiment-results -a codex`. Or copy the skill folder (skills/measure-experiment-results in product-on-purpose/pm-skills) into .agents/skills/measure-experiment-results in your project. Codex loads it when a task matches its description.

Can I use Measure Experiment Results in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add product-on-purpose/pm-skills --skill measure-experiment-results -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/measure-experiment-results, .gemini/skills/measure-experiment-results, .github/skills/measure-experiment-results and .opencode/skills/measure-experiment-results in your project.

What does Measure Experiment Results need to run?

SKILL.md names no scripts, command-line tools or credentials: Measure Experiment Results is instructions for the agent only.

Does Measure Experiment Results access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Measure Experiment Results safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Measure Experiment Results use?

Measure Experiment Results is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Measure Experiment Results use?

About 989 tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to Measure Experiment Results?

Skills that share tags, products or a category with Measure Experiment Results: Ab Test Analysis (nimrodfisher/data-analytics-skills, 470 stars), Mkt Experiment (evolution-foundation/evo-nexus, 545 stars), A/B Test Analysis (phuryn/pm-skills, 27k stars) and Experimentation (andreaskelm/pm-brain, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Measure Experiment Results?

product-on-purpose (a GitHub organization) maintains it in product-on-purpose/pm-skills, which has 716 GitHub stars. The repository holds 68 skills in this directory. The repository was last updated on October 8, 2026.

Source: product-on-purpose/pm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.