Agent skill

Ab Test Setup

by OpenClaudia in OpenClaudia/openclaudia-skills

Design, plan, and analyze A/B tests with statistical rigor. An agent skill from OpenClaudia/openclaudia-skills.

MITAuto-check passedMarketing & SEO

Install Ab Test Setup

skills CLI
$ npx skills add OpenClaudia/openclaudia-skills --skill ab-test-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenClaudia/openclaudia-skills ab-test-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenClaudia/openclaudia-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-test-setup .claude/skills/ab-test-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-setup
GitHub stars
708
Token cost
~1.7k tokens
SKILL.md length
618 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
MIT

At a glance

Design, plan, and analyze A/B tests with statistical rigor. An agent skill from OpenClaudia/openclaudia-skills.

  • Works in 8 steps: Gather Test Context → Hypothesis Framework → Sample Size and Duration → …
  • The user asks about A/B testing
  • SKILL.md covers Step 1: Gather Test Context, Step 2: Hypothesis Framework, Step 3: Sample Size and Duration and Step 4: Test Types, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ab Test Setup is an agent skill from OpenClaudia/openclaudia-skills. Design, plan, and analyze A/B tests with statistical rigor. Use when the user asks about A/B testing, split testing, experiment design, statistical significance, sample size calculation, test duration, multivariate testing, or conversion experiments. Trigger phrases include "A/B test", "split test", "experiment", "statistical significance", "sample size", "test duration", "which version wins", "conversion experiment", "hypothesis test", "variant testing".

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: 77 open-source marketing skills for Claude Code, Codex, and other AI coding agents. SEO, content, email, ads, analytics, and growth. The licence is MIT.

When your agent uses it

  • The user asks about A/B testing
  • Experiment design
  • Statistical significance
  • Sample size calculation

Example prompts

  • “A/B test”
  • “split test”
  • “experiment”
  • “/ab-test-setup”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Gather Test Context
  2. Hypothesis Framework
  3. Sample Size and Duration
  4. Test Types
  5. Test Design by Element
  6. Running the Test
  7. Common Pitfalls
  8. Test Prioritization (ICE Scoring)

What it can do on your machine

Read from SKILL.md and the folder at commit 28bf209. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Test Setup loads about 1.7k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 618 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenClaudia/openclaudia-skills at commit 28bf209, republished under its MIT licence (© OpenClaudia). 618 words, ~1,684 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-setup/SKILL.md (or your agent's skills folder).
name
ab-test-setup
description
Design, plan, and analyze A/B tests with statistical rigor. Use when the user asks about A/B testing, split testing, experiment design, statistical significance, sample size calculation, test duration, multivariate testing, or conversion experiments. Trigger phrases include "A/B test", "split test", "experiment", "statistical significance", "sample size", "test duration", "which version wins", "conversion experiment", "hypothesis test", "variant testing".

A/B Test Design and Analysis

You are an expert in experimentation and A/B testing. When the user asks you to design a test, calculate sample sizes, analyze results, or plan an experimentation roadmap, follow this framework.

Step 1: Gather Test Context

Establish: page/feature being tested, current conversion rate, monthly traffic, primary metric, secondary metrics, guardrail metrics, duration constraints, testing platform (Optimizely, VWO, custom).

Step 2: Hypothesis Framework

Hypothesis Template
OBSERVATION: [What we noticed in data/research/feedback]
HYPOTHESIS: If we [specific change], then [metric] will [change] by [amount],
            because [behavioral/psychological reasoning].
CONTROL (A): [Current state]
VARIANT (B): [Proposed change]
PRIMARY METRIC: [Single metric that determines winner]
GUARDRAILS: [Metrics that must not degrade]
Hypothesis Categories
  • Clarity: "Users don't understand what we offer" -- test headline, value prop
  • Motivation: "Users aren't motivated to act" -- test social proof, urgency, benefits
  • Friction: "Process is too difficult" -- test form length, step count, layout
  • Trust: "Users don't trust us" -- test testimonials, guarantees, badges
  • Relevance: "Content doesn't match intent" -- test personalization, segmentation

Step 3: Sample Size and Duration

Sample Size Formula
n = (Z_alpha/2 + Z_beta)^2 * (p1*(1-p1) + p2*(1-p2)) / (p2 - p1)^2
Where: Z_alpha/2 = 1.96 (95%), Z_beta = 0.84 (80% power), p2 = p1 * (1 + MDE)
Quick Reference (per variant, 95% significance, 80% power)
Baseline CR10% MDE15% MDE20% MDE25% MDE
2%385,040173,47098,74063,850
3%253,670114,30065,08042,110
5%148,64067,04038,20024,730
10%70,42031,78018,12011,740
15%44,31020,01011,4207,400
20%31,31014,1408,0705,230

Duration = (Sample size per variant x Number of variants) / Daily traffic. Minimum 7 days, maximum 8 weeks.

If duration exceeds 8 weeks: increase MDE, reduce variants, test a higher-traffic page, use a micro-conversion metric, or accept lower power.

Step 4: Test Types

TypeWhatWhenCaution
A/BTwo versions, 50/50 splitOne specific change, sufficient trafficMinimum 7 days
A/B/nControl + 2-4 variantsMultiple approaches to same elementNeeds proportionally more traffic
MVTMultiple element combinationsHigh traffic (100K+/month)Combinations multiply fast
BanditDynamic traffic allocationHigh opportunity costHarder to reach significance
Pre/PostBefore vs. after (no split)Cannot split trafficWeakest causal evidence

Step 5: Test Design by Element

Headline Tests

Test: value prop angle, specificity, social proof integration, question vs. statement, length. Measure: conversion rate, bounce rate, scroll depth.

CTA Tests

Test: button copy (action vs. benefit), color (contrast), size, placement, surrounding copy. Measure: click-through rate, conversion rate.

Layout Tests

Test: single vs. two column, long vs. short form, section order, video vs. static hero, with vs. without nav. Measure: conversion rate, scroll depth. Guardrail: page load time.

Show full SKILL.md (259 more words)Show less
Pricing Tests

Test: price point, billing display, tier count, feature allocation, default plan, anchoring, decoy pricing. Measure: revenue per visitor (not just CR). Guardrail: support tickets, refund rate.

Copy Tests

Test: tone, length, format (paragraphs vs. bullets), emotional angle, proof type. Measure: conversion rate, read depth.

Step 6: Running the Test

Pre-Launch Checklist
  • Hypothesis documented with primary metric defined
  • Sample size calculated, traffic sufficient
  • QA on both variants across devices and browsers
  • Tracking verified -- conversions fire correctly for both variants
  • No other tests on same page/funnel
  • Traffic allocation set (50/50)
  • Exclusion criteria defined (bots, internal IPs)
  • Stakeholders aligned on decision criteria before launch
During the Test
  • Do not peek for first 3-5 days (early results are misleading)
  • Do not stop early unless guardrail metrics violated
  • Monitor for technical issues and tracking accuracy
  • Watch for sample ratio mismatch (SRM): >1% deviation means setup problem
  • Do not add variants mid-test
Post-Test Analysis
TEST RESULTS
============
Test: [name] | Duration: [days] | Sample: [n] | Split: [%/%]
SRM Check: [Pass/Fail]

| Variant | Visitors | Conversions | CR | vs Control | p-value | Significant? |
|---------|----------|-------------|-----|------------|---------|--------------|
| Control | X,XXX | XXX | X.XX% | -- | -- | -- |
| Var B | X,XXX | XXX | X.XX% | +X.X% | 0.XXX | Yes/No |

DECISION: [Implement / Keep Control / Iterate]
REASONING: [Data-based rationale]
NEXT TEST: [What to test next]

Step 7: Common Pitfalls

  1. Peeking: Checking daily inflates false positives to 25-30%. Commit to sample size upfront.
  2. Underpowered tests: "No result" often means "not enough data."
  3. Too many variables: Isolate one variable per test.
  4. Ignoring segments: Overall flat, but mobile wins / desktop loses. Always segment.
  5. Novelty effect: Run 2+ weeks to account for novelty wearing off.
  6. Multiple comparisons: One primary metric. Bonferroni correction for extras.
  7. Practical significance: A significant 0.1% lift may not be worth implementing.

Step 8: Test Prioritization (ICE Scoring)

Impact (1-10): How much will this move the metric?
Confidence (1-10): How likely to produce a result?
Ease (1-10): How easy to implement?
ICE Score = (Impact + Confidence + Ease) / 3
Roadmap Template
EXPERIMENTATION ROADMAP
Quarter: [Q] | Page: [target] | Traffic: [volume] | Current CR: [X%]

| Priority | Test | ICE | Duration | Status |
|----------|------|-----|----------|--------|
| 1 | ... | 8.3 | 14 days | Ready |
| 2 | ... | 7.7 | 21 days | Ready |
| 3 | ... | 7.0 | 14 days | Idea |

Run tests sequentially on the same page to avoid interaction effects. Provide a backlog ranked by ICE score.

© OpenClaudia, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ab-test-setup of OpenClaudia/openclaudia-skills.

Open the folder on GitHubat commit 28bf209

Compare with similar skills

Ab Test Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Test Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Test Setup this skillOpenClaudia/openclaudia-skills708—~1.7kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills3.8k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills713—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Ab Test Planindranilbanerjee/digital-marketing-pro8541 repos~2kAutomated safety check: PassMIT

Similar skills

  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.8k GitHub stars~1.4k tokensUpdated 13 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    713 GitHub stars~966 tokensUpdated 2 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Ab Test Plan

    indranilbanerjee/digital-marketing-pro

    Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…

    854 GitHub starsUsed in 1 repo~2k tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Marketing & SEOAuto-check passed

More from OpenClaudia/openclaudia-skills

All 74 skills in this repo
  • Competitor Traffic Report

    OpenClaudia/openclaudia-skills

    Build an interactive competitive-traffic report for any company and its rivals — monthly visits (SimilarWeb), organic search traffic and Domain Rating (Ahrefs) — as one self-contained HTML page with…

    708 GitHub stars~1.2k tokensUpdated 19 days ago
    Auto-check passed
  • Geo Difficulty

    OpenClaudia/openclaudia-skills

    Score how hard a keyword is to rank for in the AI-search era — page-level URL Rating of real competitors (not just domain DR), Ahrefs keyword difficulty, and whether a given site already ranks or is…

    708 GitHub stars~694 tokensUpdated 19 days ago
    Auto-check passed
  • Gsc Portfolio Audit

    OpenClaudia/openclaudia-skills

    Audit EVERY Google Search Console property at once — rank all sites by clicks and impressions with period-over-period deltas, then diff keywords per site to surface what is newly ranking, rising…

    708 GitHub stars~1k tokensUpdated 19 days ago
    Auto-check passed
  • Similarweb Traffic

    OpenClaudia/openclaudia-skills

    Fetch website traffic estimates (monthly visits, traffic sources, top countries, keywords, engagement, ranks) for any domain from SimilarWeb.

    708 GitHub stars~882 tokensUpdated 19 days ago
    Auto-check passed
  • Ahrefs Python

    OpenClaudia/openclaudia-skills

    Manages Ahrefs API usage in Python using ahrefs-python library.

    708 GitHub stars~1.9k tokensUpdated 19 days ago
    Auto-check passed
  • Affiliate Marketing

    OpenClaudia/openclaudia-skills

    Build and manage an affiliate marketing program. An agent skill from OpenClaudia/openclaudia-skills.

    708 GitHub stars~2.2k tokensUpdated 19 days ago
    Auto-check passed

Categories

Questions about Ab Test Setup

What does Ab Test Setup do?

Design, plan, and analyze A/B tests with statistical rigor. An agent skill from OpenClaudia/openclaudia-skills. Ab Test Setup is an agent skill from OpenClaudia/openclaudia-skills. Design, plan, and analyze A/B tests with statistical rigor.

When should I use Ab Test Setup?

Ab Test Setup fits situations like: the user asks about A/B testing; experiment design; statistical significance; sample size calculation.

How do I install Ab Test Setup in Claude Code?

Run `npx skills add OpenClaudia/openclaudia-skills --skill ab-test-setup -a claude-code`. Or copy the skill folder (skills/ab-test-setup in OpenClaudia/openclaudia-skills) into .claude/skills/ab-test-setup in your project. Claude Code loads it when a task matches its description.

How do I install Ab Test Setup in Codex?

Run `npx skills add OpenClaudia/openclaudia-skills --skill ab-test-setup -a codex`. Or copy the skill folder (skills/ab-test-setup in OpenClaudia/openclaudia-skills) into .agents/skills/ab-test-setup in your project. Codex loads it when a task matches its description.

Can I use Ab Test Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenClaudia/openclaudia-skills --skill ab-test-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-setup, .gemini/skills/ab-test-setup, .github/skills/ab-test-setup and .opencode/skills/ab-test-setup in your project.

What does Ab Test Setup need to run?

SKILL.md names no scripts, command-line tools or credentials: Ab Test Setup is instructions for the agent only.

Does Ab Test Setup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Test Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ab Test Setup use?

Ab Test Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Test Setup use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ab Test Setup?

Skills that share tags, products or a category with Ab Test Setup: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.8k stars), Define Hypothesis (product-on-purpose/pm-skills, 713 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Test Setup?

OpenClaudia (a GitHub organization) maintains it in OpenClaudia/openclaudia-skills, which has 708 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on September 18, 2026.

Source: OpenClaudia/openclaudia-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.