Plan an A/B test by script: sample size per variant, days to run, stopping rules.

MITAuto-check passedMarketing & SEO

Install Ab Test Plan

skills CLI
$ npx skills add indranilbanerjee/digital-marketing-pro --skill ab-test-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install indranilbanerjee/digital-marketing-pro ab-test-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/indranilbanerjee/digital-marketing-pro.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-test-plan .claude/skills/ab-test-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-plan
GitHub stars
862
Used in
1 other repo
Token cost
~1.9k tokens
SKILL.md length
888 words
Files
1
Skills in repo
162
Repo updated
First seen
Licence
MIT

At a glance

Plan an A/B test by script: sample size per variant, days to run, stopping rules.

  • Works in 11 steps: Load brand context: Read… → Check campaign history: Run python… → Run sample size calculator: Execute the… → …
  • Tasks that involve A/B testing
  • SKILL.md covers Purpose, Input Required, Process and Output, plus 1 more section
  • Calls python

What it does

Ab Test Plan is an agent skill from indranilbanerjee/digital-marketing-pro. Plan an A/B test by script: sample size per variant, days to run, stopping rules. "how many visitors per variant"

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: An open-source AI marketing operating system for strategy, SEO, AEO/GEO, paid media, content, CRM, and analytics - grounded in brand context, human approval, and verifiable… The licence is MIT.

When your agent uses it

  • Tasks that involve A/B testing
  • Tasks that involve Experimental design

Example prompts

  • “how many visitors per variant”
  • “/ab-test-plan”

Requirements

  • Python 3

Workflow steps

11 steps, taken from the first numbered list in SKILL.md.

  1. Load brand context: Read ~/.claude-marketing/brands/_active-brand.json for the active slug, then load…
  2. Check campaign history: Run python "${CLAUDE_PLUGIN_ROOT}/scripts/campaign-tracker.py" --brand {slug} --action list-campaigns to review…
  3. Run sample size calculator: Execute the calculator with the baseline rate, MDE, MDE type, significance, and power. The --mde-type flag…
  4. Build hypothesis statement: Structure the hypothesis in the format: "If [specific change], then [primary metric] will [direction and…
  5. Design test variants: Define the control (current experience) and one or more treatment variants. Specify exactly what changes in each…
  6. Define primary and secondary metrics: Identify the primary success metric (the one that determines the winner) and secondary metrics to…
  7. Calculate test duration: Based on sample size requirements and daily traffic, estimate the number of days needed. Ensure the duration…
  8. Create monitoring plan: Define interim checkpoints for technical QA (not statistical peeking), sample ratio mismatch (SRM) detection, and…
  9. Define stopping rules and decision criteria: Specify when to call the test (sample size reached + significance threshold met), when to…
  10. Assess traffic feasibility: Verify that the daily traffic can reach the required sample size within a reasonable timeframe (under 8…
  11. Document pre-registration: Record the test plan before launch -- hypothesis, metrics, sample size, duration, and decision criteria -- to…

What it can do on your machine

Read from SKILL.md and the folder at commit 9e949f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Test Plan loads about 1.9k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 888 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from indranilbanerjee/digital-marketing-pro at commit 9e949f3, republished under its MIT licence (© indranilbanerjee). 888 words, ~1,888 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-plan/SKILL.md (or your agent's skills folder).
name
ab-test-plan
description
Plan an A/B test by script: sample size per variant, days to run, stopping rules. "how many visitors per variant"
argument-hint
[element-to-test]

/digital-marketing-pro:ab-test-plan

Script location. If your host does not set ${CLAUDE_PLUGIN_ROOT}, the scripts are in this plugin's scripts/ folder, next to skills/.

Purpose

Dedicated A/B test planning with a structured hypothesis framework, statistical sample size calculation, variant design, and monitoring plan. Produces a complete experiment specification with statistical rigor and clear decision criteria.

Input Required

The user must provide (or will be prompted for):

  • Element to test: The specific page, component, or experience being tested (landing page headline, CTA button, pricing page layout, email subject line, checkout flow, form design, etc.)
  • Current conversion rate: Baseline conversion rate for the metric being tested (or best estimate)
  • Desired minimum detectable effect (MDE): The smallest improvement worth detecting. MDE is ABSOLUTE by default — expressed in the same units as the baseline (baseline 5.0% and you want to catch a +1.0 percentage-point lift, i.e. 5.0% → 6.0% ⇒ --mde 0.01 --mde-type absolute). To express it as a relative lift instead (a 10% relative improvement on a 5% baseline = 5.5% ⇒ --mde 0.10 --mde-type relative), pass --mde-type relative. This distinction is the single most common sample-size error: the same "10%" read as absolute vs. relative changes the required sample size by roughly two orders of magnitude (~200×) at a 5% baseline. Always confirm which the user means.
  • Daily traffic or impressions: Average daily visitors or impressions to the test page or element
  • Significance level: Desired confidence level, default 95% (alpha = 0.05)
  • Statistical power: Desired power, default 80% (beta = 0.20)
  • Number of variants: How many variants to test (default 1 treatment + 1 control; more for multivariate)
  • Business context: What prompted the test idea (analytics data, user feedback, competitive analysis, heuristic audit, stakeholder request)

Process

  1. Load brand context: Read ~/.claude-marketing/brands/_active-brand.json for the active slug, then load ~/.claude-marketing/brands/{slug}/profile.json. Apply voice, compliance, industry context. Check guidelines/_manifest.json for restrictions, messaging, channel styles, voice-and-tone rules, and templates. If a template matching this command exists in ~/.claude-marketing/brands/{slug}/templates/, apply its format. If no brand exists, prompt for /digital-marketing-pro:brand-setup or proceed with defaults.
  2. Check campaign history: Run python "${CLAUDE_PLUGIN_ROOT}/scripts/campaign-tracker.py" --brand {slug} --action list-campaigns to review past test results and avoid re-testing already-validated hypotheses.
  3. Run sample size calculator: Execute the calculator with the baseline rate, MDE, MDE type, significance, and power. The --mde-type flag defaults to absolute — always confirm with the user which interpretation they mean before computing (the two differ by roughly two orders of magnitude, ~200×, at a 5% baseline):
    bash
    # Absolute MDE — detect a 1.0 percentage-point lift on a 5% baseline (5.0% → 6.0%)
    python "${CLAUDE_PLUGIN_ROOT}/scripts/sample-size-calculator.py" --baseline-rate 0.05 --mde 0.01 --mde-type absolute --significance 0.95 --power 0.80
    
    # Relative MDE — detect a 10% relative lift on a 5% baseline (5.0% → 5.5%)
    python "${CLAUDE_PLUGIN_ROOT}/scripts/sample-size-calculator.py" --baseline-rate 0.05 --mde 0.10 --mde-type relative --significance 0.95 --power 0.80
    This determines the required sample size per variant. Later, when the test has run, evaluate the result with python "${CLAUDE_PLUGIN_ROOT}/scripts/significance-tester.py" --control-visitors {n} --control-conversions {n} --variant-visitors {n} --variant-conversions {n} --confidence 0.95.
  4. Build hypothesis statement: Structure the hypothesis in the format: "If [specific change], then [primary metric] will [direction and magnitude] because [rationale grounded in data, user research, or established UX principle]."
  5. Design test variants: Define the control (current experience) and one or more treatment variants. Specify exactly what changes in each variant -- copy, layout, color, imagery, flow, or functionality. For multivariate tests, define the variable matrix and interaction effects to watch.
  6. Define primary and secondary metrics: Identify the primary success metric (the one that determines the winner) and secondary metrics to monitor for unintended effects (e.g., testing CTA click rate as primary, but watching bounce rate, time on page, and downstream conversion as secondary guardrails).
  7. Calculate test duration: Based on sample size requirements and daily traffic, estimate the number of days needed. Ensure the duration spans at least one full business cycle (7 days minimum) to account for day-of-week variation. Flag if duration exceeds 8 weeks (validity risk).
  8. Create monitoring plan: Define interim checkpoints for technical QA (not statistical peeking), sample ratio mismatch (SRM) detection, and guardrail metric alerts that would trigger early test stoppage for data quality or user experience reasons.
  9. Define stopping rules and decision criteria: Specify when to call the test (sample size reached + significance threshold met), when to stop early (guardrail violations, SRM detected, implementation bugs), and the protocol for inconclusive results (extend, redesign, or implement based on directional signal).
  10. Assess traffic feasibility: Verify that the daily traffic can reach the required sample size within a reasonable timeframe (under 8 weeks). If traffic is insufficient, recommend reducing the number of variants, increasing the MDE, or using qualitative methods instead.
  11. Document pre-registration: Record the test plan before launch -- hypothesis, metrics, sample size, duration, and decision criteria -- to prevent post-hoc rationalization and ensure scientific rigor.
Show full SKILL.md (158 more words)Show less

Output

A structured A/B test plan containing:

  • Hypothesis statement in If/Then/Because format with supporting evidence or rationale
  • Control and variant descriptions with specific, implementable change details
  • Required sample size per variant and total sample size
  • Estimated test duration in days based on traffic volume and required sample size
  • Primary metric and secondary metric definitions with measurement methods
  • Guardrail metrics that trigger early stoppage if degraded
  • Monitoring dashboard specification with interim checkpoint schedule
  • Statistical analysis plan (frequentist or Bayesian, one-tailed or two-tailed, correction for multiple comparisons)
  • Stopping rules for early termination (guardrail violations, SRM detection, critical bugs)
  • Go/no-go decision criteria with clear thresholds for winner declaration
  • Post-test action plan for winning, losing, and inconclusive scenarios
  • Traffic feasibility assessment with low-traffic alternative recommendations if applicable
  • Test documentation template for recording results and learnings in the campaign tracker

Agents Used

  • cro-specialist -- Hypothesis design, variant specification, sample size calculation, statistical analysis planning, monitoring framework, stopping rules, traffic feasibility assessment, and experiment documentation

© indranilbanerjee, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ab-test-plan of indranilbanerjee/digital-marketing-pro.

Open the folder on GitHubat commit 9e949f3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in indranilbanerjee/digital-marketing-pro, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Ab Test Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Test Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Test Plan this skillindranilbanerjee/digital-marketing-pro8621 repos~1.9kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills4.1k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills716—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Ab Testingericrisco/rsc-harness180—~2.4kAutomated safety check: PassMIT

Similar skills

  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    4.1k GitHub stars~1.4k tokensUpdated 16 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    716 GitHub stars~966 tokensUpdated 2 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    180 GitHub stars~2.4k tokensUpdated today
    Marketing & SEOAuto-check passed
  • Ab Test Stats

    guia-matthieu/clawfu-skills

    Calculate A/B test statistical significance. An agent skill from guia-matthieu/clawfu-skills.

    150 GitHub stars~1k tokensUpdated 9 days ago
    Marketing & SEOAuto-check passed

More from indranilbanerjee/digital-marketing-pro

All 162 skills in this repo
  • Import Template

    indranilbanerjee/digital-marketing-pro

    Import a deliverable template as a reusable placeholder template per brand.

    862 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed
  • Aeo Audit

    indranilbanerjee/digital-marketing-pro

    Run a one-time AEO audit of six AI answer engines, scored per surface.

    862 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Agent Readiness Audit

    indranilbanerjee/digital-marketing-pro

    Audit agent readiness by script: AI-crawler rules, product schema, no-JS HTML, feeds.

    862 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Backlink Gap

    indranilbanerjee/digital-marketing-pro

    Find backlink gap domains linking to competitors, not you, scored by script.

    862 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • C2pa Metadata

    indranilbanerjee/digital-marketing-pro

    Embed C2PA provenance in AI-generated images, video or PDF by script.

    862 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Audit

    indranilbanerjee/digital-marketing-pro

    Audit all campaigns running for a brand across channels, with 4-tier triage.

    862 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes

Categories

Questions about Ab Test Plan

What does Ab Test Plan do?

Plan an A/B test by script: sample size per variant, days to run, stopping rules. Ab Test Plan is an agent skill from indranilbanerjee/digital-marketing-pro. Plan an A/B test by script: sample size per variant, days to run, stopping rules.

When should I use Ab Test Plan?

Ab Test Plan fits situations like: tasks that involve A/B testing; tasks that involve Experimental design.

How do I install Ab Test Plan in Claude Code?

Run `npx skills add indranilbanerjee/digital-marketing-pro --skill ab-test-plan -a claude-code`. Or copy the skill folder (skills/ab-test-plan in indranilbanerjee/digital-marketing-pro) into .claude/skills/ab-test-plan in your project. Claude Code loads it when a task matches its description.

How do I install Ab Test Plan in Codex?

Run `npx skills add indranilbanerjee/digital-marketing-pro --skill ab-test-plan -a codex`. Or copy the skill folder (skills/ab-test-plan in indranilbanerjee/digital-marketing-pro) into .agents/skills/ab-test-plan in your project. Codex loads it when a task matches its description.

Can I use Ab Test Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add indranilbanerjee/digital-marketing-pro --skill ab-test-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-plan, .gemini/skills/ab-test-plan, .github/skills/ab-test-plan and .opencode/skills/ab-test-plan in your project.

What does Ab Test Plan need to run?

Going by SKILL.md and its folder, Ab Test Plan needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Ab Test Plan access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Test Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ab Test Plan use?

Ab Test Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Test Plan use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ab Test Plan?

Skills that share tags, products or a category with Ab Test Plan: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4.1k stars), Define Hypothesis (product-on-purpose/pm-skills, 716 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Test Plan?

indranilbanerjee (a GitHub user) maintains it in indranilbanerjee/digital-marketing-pro, which has 862 GitHub stars. The repository holds 162 skills in this directory. The repository was last updated on October 9, 2026.

Source: indranilbanerjee/digital-marketing-pro on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.