Agent skill

Ab Test Setup

by borghei in borghei/Claude-Skills

Design and analyze A/B tests: sample size, test duration, and statistical significance for conversion experiments.

MITAuto-check passedMarketing & SEO

Install Ab Test Setup

skills CLI
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills ab-test-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/marketing/ab-test-setup .claude/skills/ab-test-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-setup
GitHub stars
881
Token cost
~1.2k tokens
SKILL.md length
395 words
Files
6 (incl. scripts, references)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

Design and analyze A/B tests: sample size, test duration, and statistical significance for conversion experiments.

  • Works in 5 steps: Define hypothesis and success metric → Run sample_size_calculator.py with… → Create test configuration JSON (see… → …
  • Setting up an A/B test
  • SKILL.md covers Overview, Clarify First, Quick Start and Tools Overview, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Ab Test Setup is an agent skill from borghei/Claude-Skills. Design and analyze A/B tests: sample size, test duration, and statistical significance for conversion experiments. Use when setting up an A/B test, calculating sample size, designing an experiment, or analyzing results.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/ab-testing-guide.md`, `scripts/results_analyzer.py` and `scripts/sample_size_calculator.py`).

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Setting up an A/B test
  • Calculating sample size
  • Designing an experiment
  • Analyzing results

Example prompts

  • “/ab-test-setup”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Define hypothesis and success metric
  2. Run sample_size_calculator.py with baseline conversion and minimum detectable effect
  3. Create test configuration JSON (see Common Patterns)
  4. Run test_designer.py to generate complete test plan
  5. Share plan with stakeholders for alignment before launch

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Test Setup loads about 1.2k tokens when it runs, and up to ~2.6k if it reads all its reference files. Until then it costs about 58 tokens; SKILL.md has 395 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 395 words, ~1,198 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-setup/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
ab-test-setup
description
Design and analyze A/B tests: sample size, test duration, and statistical significance for conversion experiments. Use when setting up an A/B test, calculating sample size, designing an experiment, or analyzing results.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
marketing
metadata.domain
experimentation
metadata.updated
2026-04-02
metadata.tags
ab-testing, experimentation, statistics, sample-size, conversion-rate

A/B Test Setup Skill

Overview

Production-ready A/B testing toolkit for calculating sample sizes, designing rigorous test plans, and analyzing results with statistical significance testing. Designed for growth teams, product managers, and marketers who need to make data-driven decisions from controlled experiments.

Clarify First

Before designing the test, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Hypothesis + primary metric — what change you expect and the single metric that judges it (drives test plan + analysis)
  • Baseline conversion rate — the current rate the metric sits at today (drives sample size calculation)
  • Minimum detectable effect (MDE) — smallest lift worth detecting (drives required samples + duration)
  • Daily traffic available — eligible visitors per day per variant (determines how long the test must run)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

bash
# Calculate required sample sizes for a test
python scripts/sample_size_calculator.py --baseline 0.05 --mde 0.10 --power 0.80

# Design a complete A/B test plan
python scripts/test_designer.py test_config.json

# Analyze A/B test results
python scripts/results_analyzer.py results.json

Tools Overview

ToolPurposeInputOutput
sample_size_calculator.pySample size calculationBaseline rate, MDE, powerRequired samples + duration
test_designer.pyTest plan designJSON test configComplete test plan document
results_analyzer.pyResults analysisJSON with test resultsStatistical analysis + recommendation

Workflows

Workflow 1: New A/B Test Setup
  1. Define hypothesis and success metric
  2. Run sample_size_calculator.py with baseline conversion and minimum detectable effect
  3. Create test configuration JSON (see Common Patterns)
  4. Run test_designer.py to generate complete test plan
  5. Share plan with stakeholders for alignment before launch
Show full SKILL.md (157 more words)Show less
Workflow 2: Test Results Analysis
  1. Collect test results into JSON format
  2. Run results_analyzer.py to get statistical significance
  3. Review confidence interval, p-value, and effect size
  4. Check for segment-level effects if overall result is inconclusive
  5. Make ship/no-ship decision based on analysis
Workflow 3: Experimentation Program Review
  1. Compile results from multiple past tests
  2. Run results_analyzer.py --batch on all results
  3. Review win rate, average effect size, and velocity
  4. Identify patterns in winning vs losing tests
  5. Optimize test pipeline based on learnings

Reference Documentation

See references/ab-testing-guide.md for comprehensive methodology covering:

  • Statistical foundations (z-tests, confidence intervals)
  • Sample size theory and trade-offs
  • Common experimentation pitfalls
  • Multi-variant and sequential testing
  • Bayesian vs frequentist approaches

Common Patterns

Pattern: Test Configuration JSON
json
{
  "test_name": "Homepage CTA Button Color",
  "hypothesis": "Changing the CTA button from blue to green will increase click-through rate",
  "metric_primary": "cta_click_rate",
  "metric_secondary": ["signup_rate", "bounce_rate"],
  "baseline_rate": 0.045,
  "minimum_detectable_effect": 0.10,
  "significance_level": 0.05,
  "power": 0.80,
  "variants": [
    {"name": "control", "description": "Current blue CTA button"},
    {"name": "treatment", "description": "Green CTA button"}
  ],
  "daily_traffic": 5000,
  "allocation": {"control": 0.50, "treatment": 0.50}
}
Pattern: Test Results JSON
json
{
  "test_name": "Homepage CTA Button Color",
  "variants": {
    "control": {"visitors": 12500, "conversions": 563},
    "treatment": {"visitors": 12500, "conversions": 625}
  },
  "metric": "cta_click_rate",
  "significance_level": 0.05
}
Quick Reference: Common Effect Sizes
ContextSmall EffectMedium EffectLarge Effect
Conversion Rate2-5% relative5-15% relative> 15% relative
Revenue per User1-3%3-8%> 8%
Engagement Rate3-5%5-10%> 10%

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in marketing/ab-test-setup of borghei/Claude-Skills.

  • SKILL.md
  • examples/test_results.csv
  • references/ab-testing-guide.md
  • scripts/results_analyzer.py
  • scripts/sample_size_calculator.py
  • scripts/test_designer.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Ab Test Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Test Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Test Setup this skillborghei/Claude-Skills881—~1.2kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills3.9k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills715—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Ab Test Planindranilbanerjee/digital-marketing-pro8551 repos~2kAutomated safety check: PassMIT

Similar skills

  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    715 GitHub stars~966 tokensUpdated 4 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Ab Test Plan

    indranilbanerjee/digital-marketing-pro

    Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…

    855 GitHub starsUsed in 1 repo~2k tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated today
    Marketing & SEOAuto-check passed

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Ab Test Setup

What does Ab Test Setup do?

Design and analyze A/B tests: sample size, test duration, and statistical significance for conversion experiments. Ab Test Setup is an agent skill from borghei/Claude-Skills. Design and analyze A/B tests: sample size, test duration, and statistical significance for conversion experiments.

When should I use Ab Test Setup?

Ab Test Setup fits situations like: setting up an A/B test; calculating sample size; designing an experiment; analyzing results.

How do I install Ab Test Setup in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a claude-code`. Or copy the skill folder (marketing/ab-test-setup in borghei/Claude-Skills) into .claude/skills/ab-test-setup in your project. Claude Code loads it when a task matches its description.

How do I install Ab Test Setup in Codex?

Run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a codex`. Or copy the skill folder (marketing/ab-test-setup in borghei/Claude-Skills) into .agents/skills/ab-test-setup in your project. Codex loads it when a task matches its description.

Can I use Ab Test Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-setup, .gemini/skills/ab-test-setup, .github/skills/ab-test-setup and .opencode/skills/ab-test-setup in your project.

What does Ab Test Setup need to run?

Going by SKILL.md and its folder, Ab Test Setup needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Ab Test Setup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Test Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ab Test Setup use?

Ab Test Setup is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Test Setup use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Ab Test Setup?

Skills that share tags, products or a category with Ab Test Setup: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.9k stars), Define Hypothesis (product-on-purpose/pm-skills, 715 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Test Setup?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.