Agent skill

Ab Test Planner

by mohitagw15856 in mohitagw15856/pm-claude-skills

Design statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments.

MITAuto-check passedMarketing & SEO

Install Ab Test Planner

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill ab-test-planner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills ab-test-planner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-test-planner .claude/skills/ab-test-planner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-planner
GitHub stars
1.4k
Token cost
~2k tokens
SKILL.md length
997 words
Files
4 (incl. references)
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Design statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments.

  • Asked to set up an experiment
  • SKILL.md covers Required Inputs, Experiment Design Checklist, Hypothesis Template and Sample Size Calculator Logic, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Design an A/B test

What it does

Ab Test Planner is an agent skill from mohitagw15856/pm-claude-skills. Design statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments. Use when asked to set up an experiment, design an A/B test, calculate sample size, or interpret test results. Produces a complete test plan with hypothesis, variant definitions, sample size, duration estimate, guardrail metrics, and a results interpretation guide.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/test-validity-traps.md`, `references/worked-example.md` and `templates/test-plan.md`).

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to set up an experiment
  • Design an A/B test
  • Calculate sample size
  • Interpret test results

Example prompts

  • “/ab-test-planner”

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Test Planner loads about 2k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 997 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 997 words, ~1,972 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-planner/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ab-test-planner
description
Design statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments. Use when asked to set up an experiment, design an A/B test, calculate sample size, or interpret test results. Produces a complete test plan with hypothesis, variant definitions, sample size, duration estimate, guardrail metrics, and a results interpretation guide.

A/B Test Planner Skill

Design experiments that produce trustworthy results — not just directional signals. Every test output includes hypothesis, success metrics, sample size, duration, and a results interpretation guide.

Required Inputs

Ask the user for these if not provided:

  • What is being tested (feature, UI change, copy, pricing, onboarding step)
  • Hypothesis (or ask to help formulate one)
  • Primary metric (conversion rate, click-through, completion rate, etc.)
  • Baseline rate and minimum detectable effect (MDE)
  • Daily eligible users (to calculate duration)

Experiment Design Checklist

Before running any test, confirm:

  • Clear hypothesis with predicted direction
  • Single primary metric (plus up to 2 guardrail metrics)
  • Minimum detectable effect (MDE) defined
  • Sample size calculated
  • Test duration estimated
  • Segment isolated (no overlap with other running tests)
  • Rollback plan defined

Hypothesis Template

"We believe that [change] will cause [primary metric] to [increase/decrease] by [X%] for [user segment], because [rationale based on data or insight]."

Never run a test without a directional hypothesis. "Let's just see what happens" is not a hypothesis.

Sample Size Calculator Logic

Use this formula (provide the output, not the formula, to the user):

  • Baseline conversion rate: Current rate of primary metric
  • MDE: Smallest change worth detecting (recommend 10–20% relative lift for most features)
  • Statistical power: 80% (standard)
  • Significance level: 95% (p < 0.05)

For common scenarios, provide pre-calculated estimates:

Baseline RateMDE (Relative)Required Sample per Variant
5%20%~19,000
10%15%~14,000
20%10%~15,000
40%10%~9,500
60%5%~42,000

Always warn: "These are estimates. Use a tool like Evan Miller's calculator or Statsig for precision."

Test Duration Guidance

Minimum: 2 full weeks (to capture weekly seasonality) Maximum: 4 weeks (novelty effect distorts results beyond this)

Duration = Required sample ÷ (Daily traffic × % exposed)

Flag if traffic is too low to reach significance in under 8 weeks — recommend a different approach (e.g., holdout test, qualitative research).

Output Format

A/B Test Plan — [Test Name] — [Date]

Hypothesis:

[Filled hypothesis template]

Variants:

  • Control (A): [Current experience]
  • Treatment (B): [Changed experience — be specific]

Primary Metric: [Metric name + how measured] Guardrail Metrics: [Metrics that must not degrade]

Target Segment: [Who sees the test — % of traffic, user type] Traffic Split: [50/50 recommended unless ramp-up needed]

Sample Size Required: ~[N] users per variant Estimated Duration: [X] weeks (based on [Y] daily eligible users) Significance Threshold: 95% confidence, 80% power

Exclusions: [Any user segments to exclude and why]

Rollback Trigger: If [guardrail metric] degrades by [X%], stop the test immediately.

Results Interpretation Guide:

  • ✅ Ship if: Treatment shows [X%]+ lift on primary metric at 95% confidence AND guardrail metrics are stable
  • 🔄 Iterate if: Direction is positive but not significant — consider extending or redesigning
  • ❌ Reject if: No lift or negative direction at significance
  • ⚠️ Inconclusive: Do not ship. Do not call it a win.

Guidelines

  • Always recommend against peeking at results before the test reaches planned sample size — explain p-hacking risk
  • If user wants to test multiple variants, explain the multiple comparisons problem and recommend a Bonferroni correction or a Bayesian approach
  • If traffic is very low (<1,000 users/day), recommend qualitative alternatives: moderated testing, 5-second tests, or user interviews
  • Never approve a test with no guardrail metrics — always protect revenue, retention, or core engagement

Anti-Patterns

  • Do not run a test without a directional hypothesis — "let's see what happens" produces uninterpretable results
  • Do not declare a winner before reaching the pre-planned sample size — peeking at results inflates false positive rates
  • Do not test multiple independent changes in a single variant — you won't know which change caused the result
  • Do not use engagement metrics (clicks, time-on-page) as the primary metric when the goal is revenue or retention — proxy metrics mislead
  • Do not ignore guardrail metrics — a conversion lift that causes a support ticket spike is not a win
Show full SKILL.md (387 more words)Show less

Deeper Materials

This skill ships with support files — use them when they are available:

  • references/test-validity-traps.md — The Validity Traps That Quietly Invalidate A/B Tests. Apply it while producing the output; it carries the calibration and judgment calls the method summary above compresses.
  • templates/test-plan.md — a fill-in version of the deliverable with the quality gates inline. Offer it when the user wants to work the document themselves rather than have it generated.

Scoring Rubric (0–40)

Score any output of this skill before handing it over; 32+ is ship-quality.

Dimension0510
Statistical rigourNo sample size, or a number with no stated baseline/MDE behind itSample size present but MDE is guessed or copied from the lookup table without checking the actual baseline; power/significance unstatedSample size derived from the stated baseline and MDE at 80% power / 95% confidence, duration checked against real daily traffic and the 2–4 week window, and the low-traffic escape hatch invoked if it doesn't fit
Hypothesis discipline"Let's see what happens" — no direction, no magnitude, or multiple changes bundled into one variantDirectional hypothesis but missing magnitude, segment, or the evidence-based because; variant purity not confirmedFull template filled (change, metric, direction, magnitude, segment, rationale citing data), and the treatment isolates exactly one change with excluded ideas named as follow-up tests
Guardrails & rollbackNo guardrail metrics, or a rollback line with no thresholdGuardrails named but denominators/definitions ambiguous; rollback trigger vague ("if things look bad")1–2 guardrails protecting revenue or core engagement with pre-agreed definitions, concrete rollback thresholds, and the peeking-vs-harm-monitoring distinction handled explicitly
Decision readinessNo interpretation guide; results will be argued about after the factShip/iterate/reject listed but thresholds fuzzy; inconclusive outcome missing or treated as a soft winAll four outcomes (ship / iterate / reject / inconclusive) mapped to pre-committed thresholds, including what an inconclusive result costs and what each outcome changes next

Quality Checks

  • Hypothesis is directional (predicts a specific direction and magnitude, not "let's see")
  • Primary metric is singular (guardrail metrics are secondary)
  • Sample size is calculated from actual MDE and baseline (not guessed)
  • Test duration accounts for weekly seasonality (minimum 2 weeks)
  • Guardrail metrics are defined (at least one to protect revenue or core engagement)
  • Rollback trigger is specified with a concrete threshold

Example Trigger Phrases

  • "Set up an experiment."
  • "Design an A/B test."
  • "Calculate sample size."
  • "Interpret test results."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/ab-test-planner of mohitagw15856/pm-claude-skills.

  • SKILL.md
  • references/test-validity-traps.md
  • references/worked-example.md
  • templates/test-plan.md

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Ab Test Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Test Planner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Test Planner this skillmohitagw15856/pm-claude-skills1.4k—~2kAutomated safety check: PassMIT
Ab Test Analyzeririnabuht12-oss/marketing-skills4.1k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills716—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Send Experiment Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~4.1kAutomated safety check: PassApache-2.0
Content Experimentation Best Practicessanity-io/agent-toolkit1881 repos~469Automated safety check: PassMIT

Similar skills

  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    4.1k GitHub stars~1.4k tokensUpdated 16 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    716 GitHub stars~966 tokensUpdated 2 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Send Experiment Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an email A/B test", "set up a multivariate subject/CTA test", "run a send-time test", "build a hold-out group", or "is this email result…

    2.9k GitHub starsUsed in 2 repos~4.1k tokens
    Marketing & SEOAuto-check passed
  • Official

    Content experimentation and A/B testing guidance covering experiment design, hypotheses, metrics, sample size, statistical foundations, CMS-managed variants, and common analysis pitfalls.

    188 GitHub starsUsed in 1 repo~469 tokens
    Marketing & SEOAuto-check passed
  • Ads Test

    AgriciDaniel/claude-ads

    Design and evaluate paid-ad experiments with hypotheses, randomization units, sample-size and duration assumptions, guardrails, platform experiment tools, analysis, and decision rules.

    9.9k GitHub stars~312 tokensUpdated 3 days ago
    Marketing & SEOAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Ab Test Planner

What does Ab Test Planner do?

Design statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments. Ab Test Planner is an agent skill from mohitagw15856/pm-claude-skills. Design statistically rigorous A/B tests for product features, UI changes, onboarding flows, and pricing experiments.

When should I use Ab Test Planner?

Ab Test Planner fits situations like: asked to set up an experiment; design an A/B test; calculate sample size; interpret test results.

How do I install Ab Test Planner in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill ab-test-planner -a claude-code`. Or copy the skill folder (skills/ab-test-planner in mohitagw15856/pm-claude-skills) into .claude/skills/ab-test-planner in your project. Claude Code loads it when a task matches its description.

How do I install Ab Test Planner in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill ab-test-planner -a codex`. Or copy the skill folder (skills/ab-test-planner in mohitagw15856/pm-claude-skills) into .agents/skills/ab-test-planner in your project. Codex loads it when a task matches its description.

Can I use Ab Test Planner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill ab-test-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-planner, .gemini/skills/ab-test-planner, .github/skills/ab-test-planner and .opencode/skills/ab-test-planner in your project.

What does Ab Test Planner need to run?

SKILL.md names no scripts, command-line tools or credentials: Ab Test Planner is instructions for the agent only.

Does Ab Test Planner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Test Planner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ab Test Planner use?

Ab Test Planner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Test Planner use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Ab Test Planner?

Skills that share tags, products or a category with Ab Test Planner: Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4.1k stars), Define Hypothesis (product-on-purpose/pm-skills, 716 stars), A B Test Design (Owl-Listener/designer-skills, 2.9k stars) and Send Experiment Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Test Planner?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.