Agent skill

Ab Testing Framework

by thatrebeccarae in thatrebeccarae/claude-marketing

A/B and multivariate testing methodology. An agent skill from thatrebeccarae/claude-marketing.

MITAuto-check passedMarketing & SEO

Install Ab Testing Framework

skills CLI
$ npx skills add thatrebeccarae/claude-marketing --skill ab-testing-framework -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install thatrebeccarae/claude-marketing ab-testing-framework --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/thatrebeccarae/claude-marketing.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-testing-framework .claude/skills/ab-testing-framework && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-testing-framework
GitHub stars
162
Token cost
~1.3k tokens
SKILL.md length
553 words
Files
4
Skills in repo
55
Repo updated
First seen
Licence
MIT

At a glance

A/B and multivariate testing methodology. An agent skill from thatrebeccarae/claude-marketing.

  • Works in 5 steps: Hypothesis → Sample Size Calculation → Test Execution Rules → …
  • The user asks about A/B testing
  • SKILL.md covers Install, Test Design Process, Common Testing Pitfalls and What to Test (Prioritized by…, plus 1 more section
  • Calls git

What it does

Ab Testing Framework is an agent skill from thatrebeccarae/claude-marketing. A/B and multivariate testing methodology. Design experiments, calculate sample sizes, determine statistical significance, avoid common pitfalls, and interpret results. Platform-agnostic framework applicable to landing pages, emails, ads, pricing, and product features. Use when the user asks about A/B testing, split testing, experiment design, statistical significance, or conversion experiments.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `EXAMPLES.md` and `REFERENCE.md`).

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: A full marketing department for Claude Code. Skill packs for Klaviyo, Shopify, GA4, Looker Studio, paid media, and more. Audit, optimize, and report using natural language. The licence is MIT.

When your agent uses it

  • The user asks about A/B testing
  • Experiment design
  • Statistical significance
  • Conversion experiments

Example prompts

  • “/ab-testing-framework”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Hypothesis
  2. Sample Size Calculation
  3. Test Execution Rules
  4. Statistical Analysis
  5. Decision Framework

What it can do on your machine

Read from SKILL.md and the folder at commit a8a63ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Testing Framework loads about 1.3k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 553 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from thatrebeccarae/claude-marketing at commit a8a63ec, republished under its MIT licence (© thatrebeccarae). 553 words, ~1,342 tokens.

Download SKILL.mdSave it as .claude/skills/ab-testing-framework/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ab-testing-framework
description
A/B and multivariate testing methodology. Design experiments, calculate sample sizes, determine statistical significance, avoid common pitfalls, and interpret results. Platform-agnostic framework applicable to landing pages, emails, ads, pricing, and product features. Use when the user asks about A/B testing, split testing, experiment design, statistical significance, or conversion experiments.
license
MIT
origin
custom
author
Rebecca Rae Barton
author_url
https://github.com/thatrebeccarae
metadata.version
1.0.0
metadata.category
analytics
metadata.domain
experimentation
metadata.updated
2026-03-18
metadata.tested
2026-03-18
metadata.tested_with
Claude Code v2.1

A/B Testing Framework

Design, run, and analyze conversion experiments with statistical rigor.

Install

bash
git clone https://github.com/thatrebeccarae/claude-marketing.git && cp -r claude-marketing/skills/ab-testing-framework ~/.claude/skills/

Test Design Process

Step 1: Hypothesis

Template: If we [change X], then [metric Y] will [increase/decrease] by [Z%] because [reason].

Good hypothesis: "If we change the CTA from Get Started to Start Free Trial, then signup rate will increase by 15% because it reduces uncertainty about cost."

Bad hypothesis: "If we change the button color, conversions will improve." (No reasoning, no expected magnitude.)

Step 2: Sample Size Calculation

To determine how long to run a test:

Required sample per variation = 16 * (p * (1-p)) / (MDE^2)

Where:
  p = baseline conversion rate (as decimal)
  MDE = minimum detectable effect (as decimal)
Baseline Rate10% MDE20% MDE30% MDE
1%253,41463,35428,157
3%82,36920,5929,152
5%48,64012,1605,404
10%23,0405,7602,560
20%10,2402,5601,138

Minimum test duration: 2 full business weeks (to capture day-of-week effects), even if sample size is reached sooner.

Step 3: Test Execution Rules
  1. Random assignment — visitors must be randomly assigned to control/variant
  2. No peeking — do not check results before reaching sample size
  3. No mid-test changes — do not modify variants during the test
  4. Even traffic split — 50/50 for A/B, even splits for multivariate
  5. Single variable — change only one thing per test (unless multivariate)
  6. Full duration — run for the pre-calculated duration, not until significance
Step 4: Statistical Analysis
Frequentist Approach

Z-test for proportions:

Z = (p1 - p2) / sqrt(p_pooled * (1 - p_pooled) * (1/n1 + 1/n2))

Where:
  p1, p2 = conversion rates of control and variant
  p_pooled = (x1 + x2) / (n1 + n2)
  n1, n2 = sample sizes

p-value interpretation:

  • p < 0.05: Statistically significant (95% confidence)
  • p < 0.01: Highly significant (99% confidence)
  • p >= 0.05: Not significant — do not declare a winner
Bayesian Approach

When to use Bayesian:

  • Low traffic (small sample sizes)
  • Need to make decisions faster
  • Want probability of each variant being best (not just "significant or not")

Interpretation: "There is a 94% probability that Variant B is better than Control" vs frequentist "We reject the null hypothesis at 95% confidence."

Step 5: Decision Framework
ResultSignificanceAction
Variant winsp < 0.05Implement variant
Control winsp < 0.05Keep control, learn from failure
No differencep >= 0.05Keep control, test something bigger
Variant winsp = 0.05-0.10Consider traffic — may need more time
Show full SKILL.md (224 more words)Show less

Common Testing Pitfalls

  1. Peeking — checking results early inflates false positive rate from 5% to 26%+
  2. Stopping early — reaching significance != reaching required sample size
  3. Testing too many variants — each variant needs full sample size
  4. Ignoring segments — overall winner may be loser for key segments
  5. Too small an effect — testing for 2% lift needs enormous sample sizes
  6. Not accounting for seasonality — run full weeks, avoid holidays
  7. Multiple metrics — primary metric must be pre-declared; secondary are directional
  8. Survivorship bias — only measuring users who complete, not those who abandon
  9. Simpson paradox — segment-level winners can reverse at aggregate level
  10. Novelty effect — new designs get temporary lift; re-test after 2-4 weeks

What to Test (Prioritized by Impact)

High Impact
  • Value proposition / headline
  • CTA text and placement
  • Pricing and offer structure
  • Form length (fields removed)
  • Page layout (single column vs multi)
  • Social proof presence and placement
Medium Impact
  • Image/video vs static
  • Testimonial format (text vs video)
  • Navigation presence on landing pages
  • Trust badges and security signals
  • Urgency elements (countdown, stock)
Low Impact (Usually Not Worth Testing)
  • Button color (unless extreme contrast issue)
  • Font changes
  • Minor copy tweaks
  • Icon styles
  • Footer content

Integration with Other Skills

  • cro-auditor — CRO audit generates test hypotheses; this skill designs the experiments
  • google-analytics — GA4 for experiment data and segment analysis
  • copywriting-frameworks — Generate variant copy using proven frameworks

© thatrebeccarae, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in skills/ab-testing-framework of thatrebeccarae/claude-marketing.

  • SKILL.md
  • EXAMPLES.md
  • LICENSE
  • REFERENCE.md

Open the folder on GitHubat commit a8a63ec

Compare with similar skills

Ab Testing Framework next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Testing Framework compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Testing Framework this skillthatrebeccarae/claude-marketing162—~1.3kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills3.9k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills715—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Ab Test Planindranilbanerjee/digital-marketing-pro8551 repos~2kAutomated safety check: PassMIT

Similar skills

  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    715 GitHub stars~966 tokensUpdated 4 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Ab Test Plan

    indranilbanerjee/digital-marketing-pro

    Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…

    855 GitHub starsUsed in 1 repo~2k tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated yesterday
    Marketing & SEOAuto-check passed

More from thatrebeccarae/claude-marketing

All 55 skills in this repo
  • Google Analytics

    thatrebeccarae/claude-marketing

    Analyze Google Analytics data, review website performance metrics, identify traffic patterns, and suggest data-driven improvements.

    162 GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check: notes
  • Content Creator

    thatrebeccarae/claude-marketing

    Comprehensive content marketing toolkit with brand voice analysis, SEO optimization scripts, content frameworks, social media strategy, and content calendar planning.

    162 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Klaviyo Analyst

    thatrebeccarae/claude-marketing

    Klaviyo marketing operations and analyst expertise. An agent skill from thatrebeccarae/claude-marketing.

    162 GitHub stars~5k tokensUpdated 4 mo ago
    Auto-check: notes
  • Klaviyo Developer

    thatrebeccarae/claude-marketing

    Klaviyo API and developer integration expertise. An agent skill from thatrebeccarae/claude-marketing.

    162 GitHub stars~4.9k tokensUpdated 4 mo ago
    Auto-check: notes
  • Looker Studio

    thatrebeccarae/claude-marketing

    Looker Studio (formerly Google Data Studio) expertise. An agent skill from thatrebeccarae/claude-marketing.

    162 GitHub stars~3.6k tokensUpdated 4 mo ago
    Auto-check: notes
  • Shopify

    thatrebeccarae/claude-marketing

    Shopify e-commerce platform marketing expertise. An agent skill from thatrebeccarae/claude-marketing.

    162 GitHub stars~2.4k tokensUpdated 4 mo ago
    Auto-check: notes

Categories

Questions about Ab Testing Framework

What does Ab Testing Framework do?

A/B and multivariate testing methodology. An agent skill from thatrebeccarae/claude-marketing. Ab Testing Framework is an agent skill from thatrebeccarae/claude-marketing. A/B and multivariate testing methodology.

When should I use Ab Testing Framework?

Ab Testing Framework fits situations like: the user asks about A/B testing; experiment design; statistical significance; conversion experiments.

How do I install Ab Testing Framework in Claude Code?

Run `npx skills add thatrebeccarae/claude-marketing --skill ab-testing-framework -a claude-code`. Or copy the skill folder (skills/ab-testing-framework in thatrebeccarae/claude-marketing) into .claude/skills/ab-testing-framework in your project. Claude Code loads it when a task matches its description.

How do I install Ab Testing Framework in Codex?

Run `npx skills add thatrebeccarae/claude-marketing --skill ab-testing-framework -a codex`. Or copy the skill folder (skills/ab-testing-framework in thatrebeccarae/claude-marketing) into .agents/skills/ab-testing-framework in your project. Codex loads it when a task matches its description.

Can I use Ab Testing Framework in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add thatrebeccarae/claude-marketing --skill ab-testing-framework -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-testing-framework, .gemini/skills/ab-testing-framework, .github/skills/ab-testing-framework and .opencode/skills/ab-testing-framework in your project.

What does Ab Testing Framework need to run?

Going by SKILL.md and its folder, Ab Testing Framework needs the command-line tools its instructions call (git).

Does Ab Testing Framework access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ab Testing Framework safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ab Testing Framework use?

Ab Testing Framework is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Testing Framework use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ab Testing Framework?

Skills that share tags, products or a category with Ab Testing Framework: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.9k stars), Define Hypothesis (product-on-purpose/pm-skills, 715 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Testing Framework?

thatrebeccarae (a GitHub user) maintains it in thatrebeccarae/claude-marketing, which has 162 GitHub stars. The repository holds 55 skills in this directory. The repository was last updated on May 14, 2026.

Source: thatrebeccarae/claude-marketing on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.