Agent skill

A/B Test Analysis

by phuryn in phuryn/pm-skills

Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

MITAuto-check passedData & Analytics

Install A/B Test Analysis

skills CLI
$ npx skills add phuryn/pm-skills --skill ab-test-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install phuryn/pm-skills ab-test-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/phuryn/pm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pm-data-analytics/skills/ab-test-analysis .claude/skills/ab-test-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-analysis
GitHub stars
27k
Token cost
~893 tokens
SKILL.md length
338 words
Files
1
Skills in repo
62
Repo updated
First seen
Licence
MIT

At a glance

Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

  • Works in 6 steps: Understand the experiment → Validate the test setup → Calculate statistical significance → …
  • Checking whether a split test has reached significance
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Deciding if a winning variant is safe to roll out given guardrail metrics

What it does

The agent first asks about the experiment: the hypothesis, what the variant changed, the primary metric and any guardrail metrics, how long it ran and the traffic split. It then checks that the test is trustworthy by looking at sample size against the expected effect, whether power falls below 80%, run length against business cycles, sample ratio mismatch and novelty effects.

For the numbers, it computes conversion rates for control and variant, relative lift, a p-value from a two-tailed z-test or chi-squared test, and a 95% confidence interval, and it separates statistical from practical significance. Given raw CSV, Excel or analytics exports, it writes and runs a Python script for the calculations. A table maps each outcome, such as a win with guardrail concerns or a flat result, to ship, investigate, extend, stop or revert, and the agent finishes with a written summary of the test.

When your agent uses it

  • Checking whether a split test has reached significance
  • Deciding if a winning variant is safe to roll out given guardrail metrics
  • Interpreting exported experiment data from an analytics tool
  • Judging whether a test was underpowered or needs more time

Example prompts

  • “Analyze the checkout button test in ./data/checkout-ab.csv and tell me if we should ship variant B.”
  • “Our pricing page test ran for ten days with a 50/50 split. Is the lift real or do we extend it?”
  • “Check this experiment for sample ratio mismatch before I report the result.”

Requirements

  • Python, for the scripts it writes when raw data is provided

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Understand the experiment
  2. Validate the test setup
  3. Calculate statistical significance
  4. Check guardrail metrics
  5. Interpret results
  6. Provide the analysis summary

What it can do on your machine

Read from SKILL.md and the folder at commit 8607e3b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • productcompass.pm

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

A/B Test Analysis loads about 893 tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 338 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~893

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from phuryn/pm-skills at commit 8607e3b, republished under its MIT licence (© phuryn). 338 words, ~893 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-analysis/SKILL.md (or your agent's skills folder).
name
ab-test-analysis
description
Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant.

A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

Context

You are analyzing A/B test results for $ARGUMENTS.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

Instructions
  1. Understand the experiment:

    • What was the hypothesis?
    • What was changed (the variant)?
    • What is the primary metric? Any guardrail metrics?
    • How long did the test run?
    • What is the traffic split?
  2. Validate the test setup:

    • Sample size: Is the sample large enough for the expected effect size?
      • Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
      • Flag if the test is underpowered (<80% power)
    • Duration: Did the test run for at least 1-2 full business cycles?
    • Randomization: Any evidence of sample ratio mismatch (SRM)?
    • Novelty/primacy effects: Was there enough time to wash out initial behavior changes?
  3. Calculate statistical significance:

    • Conversion rate for control and variant
    • Relative lift: (variant - control) / control × 100
    • p-value: Using a two-tailed z-test or chi-squared test
    • Confidence interval: 95% CI for the difference
    • Statistical significance: Is p < 0.05?
    • Practical significance: Is the lift meaningful for the business?

    If the user provides raw data, generate and run a Python script to calculate these.

  4. Check guardrail metrics:

    • Did any guardrail metrics (revenue, engagement, page load time) degrade?
    • A winning primary metric with degraded guardrails may not be a true win
  5. Interpret results:

    OutcomeRecommendation
    Significant positive lift, no guardrail issuesShip it — roll out to 100%
    Significant positive lift, guardrail concernsInvestigate — understand trade-offs before shipping
    Not significant, positive trendExtend the test — need more data or larger effect
    Not significant, flatStop the test — no meaningful difference detected
    Significant negative liftDon't ship — revert to control, analyze why
  6. Provide the analysis summary:

    ## A/B Test Results: [Test Name]
    
    **Hypothesis**: [What we expected]
    **Duration**: [X days] | **Sample**: [N control / M variant]
    
    | Metric | Control | Variant | Lift | p-value | Significant? |
    |---|---|---|---|---|---|
    | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
    | [Guardrail] | ... | ... | ... | ... | ... |
    
    **Recommendation**: [Ship / Extend / Stop / Investigate]
    **Reasoning**: [Why]
    **Next steps**: [What to do]
Show full SKILL.md (37 more words)Show less

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.


Further Reading

© phuryn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in pm-data-analytics/skills/ab-test-analysis of phuryn/pm-skills.

Open the folder on GitHubat commit 8607e3b

Compare with similar skills

A/B Test Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

A/B Test Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
A/B Test Analysis this skillphuryn/pm-skills27k—~893Automated safety check: PassMIT
Data And Funnel Analyticsmanojbajaj95/claude-gtm-plugin104—~2.9kAutomated safety check: PassMIT
Meridian MMM Model Buildinggoogle/meridian1.6k—~2.5kAutomated safety check: PassApache-2.0
Analytics Trackingfreekmurze/dotfiles1k12 repos~2kAutomated safety check: PassNone
Statistical Analystalirezarezvani/claude-skills28k1 repos~2.5kAutomated safety check: PassMIT
Experimentation Analyticsrampstackco/claude-skills9351 repos~8.9kAutomated safety check: PassMIT

Similar skills

  • Data And Funnel Analytics

    manojbajaj95/claude-gtm-plugin

    Analytics tracking, interpretation, funnel analysis, product metrics, and ROI measurement.

    104 GitHub stars~2.9k tokensUpdated 19 days ago
    Data & AnalyticsAuto-check passed
  • Official

    Takes a user through building a Meridian marketing mix model, from loading CSV data and mapping columns to running EDA, fitting and saving the model.

    1.6k GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Analytics Tracking

    freekmurze/dotfiles

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    1k GitHub starsUsed in 12 repos~2k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Analyst

    alirezarezvani/claude-skills

    Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Experimentation Analytics

    rampstackco/claude-skills

    How to read experiment results without fooling yourself. An agent skill from rampstackco/claude-skills.

    935 GitHub starsUsed in 1 repo~8.9k tokens
    Data & AnalyticsAuto-check passed
  • Meta Results Forest Plot Analyzer

    aipoch/medical-research-skills

    Analyzes forest plots for meta-analysis, generating detailed descriptions and formatting figure legends in Chinese or English.

    2k GitHub stars~1.9k tokensUpdated 20 days ago
    Data & AnalyticsAuto-check passed

More from phuryn/pm-skills

All 62 skills in this repo
  • Reviews a diff by anchoring on agreements between two sides of a boundary, forcing a concrete violating execution, and refuting each finding before reporting it.

    27k GitHub stars~3.6k tokensUpdated 22 days ago
    Auto-check passed
  • Team OKR Brainstorm

    phuryn/pm-skills

    Drafts three alternative sets of team OKRs, each with an inspiring objective and measurable key results, tied to the company strategy you provide.

    27k GitHub stars~1.1k tokensUpdated 22 days ago
    Auto-check passed
  • Analyzes uploaded cohort data to compute retention curves and feature adoption trends, builds heatmaps and charts, and suggests qualitative follow-up research.

    27k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed
  • Dummy Dataset Generator

    phuryn/pm-skills

    Generates realistic test datasets with custom columns, row counts and business constraints, output as CSV, JSON, SQL inserts or a runnable Python script.

    27k GitHub stars~983 tokensUpdated 22 days ago
    Auto-check passed
  • Grammar and Flow Checker

    phuryn/pm-skills

    Reviews a draft for grammar, logic and flow problems and returns located, prioritized fix suggestions without rewriting the whole text.

    27k GitHub stars~2.4k tokensUpdated 22 days ago
    Auto-check passed
  • Builds a structured customer interview script with opening, warm-up, jobs-to-be-done exploration and wrap-up sections, following Mom Test rules against leading questions.

    27k GitHub stars~1.2k tokensUpdated 22 days ago
    Auto-check passed

Works with

Questions about A/B Test Analysis

What does A/B Test Analysis do?

Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop. The agent first asks about the experiment: the hypothesis, what the variant changed, the primary metric and any guardrail metrics, how long it ran and the traffic split. It then checks that the test is trustworthy by looking at sample size against the expected effect, whether power falls below 80%, run length against business cycles, sample ratio mismatch and novelty effects.

When should I use A/B Test Analysis?

A/B Test Analysis fits situations like: checking whether a split test has reached significance; deciding if a winning variant is safe to roll out given guardrail metrics; interpreting exported experiment data from an analytics tool; judging whether a test was underpowered or needs more time.

How do I install A/B Test Analysis in Claude Code?

Run `npx skills add phuryn/pm-skills --skill ab-test-analysis -a claude-code`. Or copy the skill folder (pm-data-analytics/skills/ab-test-analysis in phuryn/pm-skills) into .claude/skills/ab-test-analysis in your project. Claude Code loads it when a task matches its description.

How do I install A/B Test Analysis in Codex?

Run `npx skills add phuryn/pm-skills --skill ab-test-analysis -a codex`. Or copy the skill folder (pm-data-analytics/skills/ab-test-analysis in phuryn/pm-skills) into .agents/skills/ab-test-analysis in your project. Codex loads it when a task matches its description.

Can I use A/B Test Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add phuryn/pm-skills --skill ab-test-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-analysis, .gemini/skills/ab-test-analysis, .github/skills/ab-test-analysis and .opencode/skills/ab-test-analysis in your project.

What does A/B Test Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: A/B Test Analysis is instructions for the agent only. Our summary lists: Python, for the scripts it writes when raw data is provided.

Does A/B Test Analysis access the network?

SKILL.md names 1 domain. As links in the text: productcompass.pm. This is read from the text; nothing was executed.

Is A/B Test Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does A/B Test Analysis use?

A/B Test Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does A/B Test Analysis use?

About 893 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to A/B Test Analysis?

Skills that share tags, products or a category with A/B Test Analysis: Data And Funnel Analytics (manojbajaj95/claude-gtm-plugin, 104 stars), Meridian MMM Model Building (google/meridian, 1.6k stars), Analytics Tracking (freekmurze/dotfiles, 1k stars) and Statistical Analyst (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains A/B Test Analysis?

phuryn (a GitHub user) maintains it in phuryn/pm-skills, which has 26,809 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on September 14, 2026.

Source: phuryn/pm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.