Agent skill

Ab Testing

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

MITAuto-check passedMarketing & SEO

Install Ab Testing

skills CLI
$ npx skills add ericrisco/rsc-harness --skill ab-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness ab-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-testing .claude/skills/ab-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-testing
GitHub stars
167
Token cost
~2.4k tokens
SKILL.md length
1,193 words
Files
6 (incl. scripts, references)
Skills in repo
227
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

  • Works in 5 steps: Hypothesis and metrics → Sample size from MDE, baseline, and power → Run discipline → …
  • Analyzing a controlled experiment — falsifiable hypothesis
  • SKILL.md covers Pre-test checklist — every…, Step 1 — Hypothesis and metrics, Step 2 — Sample size from MDE,… and Step 3 — Run discipline, plus 4 more sections
  • Runs Shell scripts from its folder

What it does

Ab Testing is an agent skill from ericrisco/rsc-harness. Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is analytics), NOT north-star/KPI trees (that is kpi-framework), NOT projecting metrics forward (that is forecasting).

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/pitfalls.md`).

It sits in Marketing & SEO, covering A/B testing, Experimental design and Forecasting and time series. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Analyzing a controlled experiment — falsifiable hypothesis
  • Sample size from an MDE
  • Reading significance/CI/power
  • Rescuing tests that wont go significant

Example prompts

  • “/ab-testing”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Hypothesis and metrics
  2. Sample size from MDE, baseline, and power
  3. Run discipline
  4. Analyze
  5. CUPED variance reduction

What it can do on your machine

Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Testing loads about 2.4k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 1,193 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,193 words, ~2,438 tokens.

Download SKILL.mdSave it as .claude/skills/ab-testing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
ab-testing
description
Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is `analytics`), NOT north-star/KPI trees (that is `kpi-framework`), NOT projecting metrics forward (that is `forecasting`).
tags
ab-testing, experimentation, statistics, cuped, sample-size, hypothesis-testing
recommends
analytics, kpi-framework, forecasting, data-cleaning, python, reporting
origin
risco

A/B testing — design and read a defensible experiment

An experiment without a pre-committed sample size and a single primary metric is not an experiment. It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost entirely before traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four.

Pre-test checklist — every line true before any traffic

Each one is a place experiments die silently.

  • A falsifiable hypothesis — names the change, the direction, and the metric it moves.
  • Exactly ONE primary metric. More than one primary = multiple comparisons = inflated false positives.
  • Guardrail metrics — what you refuse to harm (latency, refunds, unsubscribes) even for a win.
  • The randomization unit = the analysis unit (usually the user). Mixing them is pseudoreplication.
  • An MDE — the smallest lift that would change a decision. Not "any difference."
  • A computed sample size and the duration it implies at your real daily eligible traffic.
  • A fixed stop rule — a date or an n you commit to before launch. No "we'll see how it looks."

Step 1 — Hypothesis and metrics

State a null you can reject. "The new checkout button changes purchase conversion" with H0: conversion equal across arms, H1: it differs. Vague aspirations ("improve the funnel") have no rejection region.

Pick one primary metric and freeze it. Why: every extra primary metric is another coin flip at α, so three "primary" metrics turn a 5% false-positive rate into roughly 14%. Demote the rest to secondary.

Randomize on the same unit you analyze on. If a user sees the variant on every visit, randomize by user, not by session — analyzing 50k sessions from 8k users treats correlated observations as independent and fabricates significance.

text
Bad:  "We think the redesign will improve engagement and revenue and retention."  (no null, 3 primaries, no number)
Good: "H0: 30-day purchase conversion is equal between control and the new one-click button.
       H1: it differs. Primary: purchase conversion. Guardrails: refund rate, p95 checkout latency.
       Randomize by user_id. MDE: +1.5pp absolute on a 12% baseline."

Step 2 — Sample size from MDE, baseline, and power

Defaults: power 0.80, α 0.05 (two-sided). The MDE is yours to choose — it is the smallest effect that would actually change what you do.

Rule: required n scales with ~1/MDE². Why: halving the smallest effect you care to detect roughly quadruples the traffic and time. This is the single most expensive decision in the design, so set the MDE to a business threshold, never to "whatever is small."

For a conversion rate (proportion):

python
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize

p1, p2 = 0.12, 0.135                       # baseline, baseline + MDE (1.5pp)
h = proportion_effectsize(p1, p2)          # Cohen's h (arcsine transform)
n = NormalIndPower().solve_power(effect_size=h, alpha=0.05, power=0.80, ratio=1.0)
print(int(-(-n // 1)))                      # n PER ARM, rounded up

For a continuous metric (revenue per user, time on page) use Welch-style sizing:

python
from statsmodels.stats.power import TTestIndPower

effect = mde_in_units / pooled_std         # Cohen's d
n = TTestIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.80, ratio=1.0)

Then convert n to a calendar plan: days = ceil((n_per_arm * num_arms) / daily_eligible_users). If that is 9 days, run a clean two full weeks anyway — weekday/weekend mix is part of the population, and a 6-day test oversamples whoever shows up Tuesday. Full worked example (12% baseline, +1.5pp MDE, 80% power) plus runnable sizing, n→duration, CUPED θ and SRM snippets: references/sample-size-and-cuped.md.

Step 3 — Run discipline

Fixed horizon is the default. Commit to the n/date from Step 2 and read the result once, at the end.

Do not peek and stop at first significance. Why: checking repeatedly and stopping the moment p < 0.05 inflates the Type-I error far above 5% — with enough looks, a null test crosses 0.05 most of the time. If you genuinely need to stop early, use a sequential / always-valid method (confidence sequences, e.g. Netflix's anytime-valid CIs) that holds Type-I error under continuous monitoring. Sequential is strong for killing losers early and weak for calling winners early — for a confident win, the fixed-horizon read is tighter.

Gate on SRM before you trust anything. Compute a chi-square test on the observed split versus the intended ratio. If p < 0.001 the assignment or logging is broken — a bot filter dropping one arm, a redirect, a caching bug. Fix the instrumentation and rerun; do not "adjust for it."

The peeking Type-I math, sequential/always-valid options, SRM diagnosis, novelty/primacy effects, Simpson's paradox in segments and HARKing all live in references/pitfalls.md.

Step 4 — Analyze

Pick the test by metric type:

Metric typeTest
Binary conversion (proportion)Two-proportion z-test (statsmodels.stats.proportion.proportions_ztest)
Continuous, roughly normal / large nWelch's t-test (scipy.stats.ttest_ind(..., equal_var=False))
Continuous, heavy-tailed / skewed (revenue)Mann-Whitney U, or t-test on a log/winsorized metric

Report lift + confidence interval + p-value together. Never p alone. Why: p < 0.05 with a CI of [+0.1pp, +5pp] is "statistically there, practically a coin toss" — the CI tells you the size, p only tells you it is not exactly zero. Practical significance = compare the CI to your MDE: if the whole interval sits above the MDE, ship; if it straddles the MDE, you detected something too small to matter.

Multiple comparisons. Two regimes:

  • Small set of pre-declared decision metrics → Bonferroni (divide α by the count). Conservative, simple.
  • Large exploratory scan of many metrics/segments → Benjamini-Hochberg (FDR). It keeps far more power than Bonferroni on big scans (in a 20-effect example, ~17 detected vs ~12 under Bonferroni).
Show full SKILL.md (409 more words)Show less

Step 5 — CUPED variance reduction

CUPED (Controlled-experiment Using Pre-Experiment Data) subtracts predictable pre-period noise so the same traffic buys more power — or the same power needs less traffic. The adjusted metric:

text
Y_cuped = Y − θ · (X − E[X])        where  θ = Cov(Y, X) / Var(X)

Estimate θ by regressing the in-experiment metric Y on the pre-experiment covariate X (e.g. each user's spend in the 4 weeks before the test), then analyze Y_cuped with the same test as Step 4.

When it pays: recurring users with a strong pre-period signal. Reported wins — Netflix ~40% variance reduction on engagement, Statsig 50%+ on common metrics → significance in roughly half the time/traffic.

When it does nothing — do not bother: brand-new users (no pre-period data), a covariate uncorrelated with the outcome, or — the cardinal sin — a covariate measured after assignment, which biases the estimate. The covariate MUST be pre-treatment and independent of which arm a user lands in. Runnable θ-via-OLS snippet in references/sample-size-and-cuped.md.

Anti-patterns

BadWhy it is wrongDo instead
Peek daily, stop the day p < 0.05Repeated looks inflate Type-I error far above αFix n/date up front; or a sequential method that holds α
No sample size set before launchYou will stop on noise and call it a winCompute n from MDE/baseline/power in Step 2
Several "primary" metricsEach is a coin flip at α; 3 metrics ≈ 14% false-positiveOne frozen primary; the rest are secondary
Ignore the observed splitAn SRM means assignment/logging is broken; results are garbageChi-square SRM gate before reading anything
Report only the p-valueHides effect size — p < 0.05 can be practically zeroAlways lift + CI + p; compare CI to MDE
CUPED on a post-assignment covariateCovariate correlated with the arm biases θUse only pre-treatment, assignment-independent covariates
Call a winner from an underpowered test"Not significant" then ≠ "no effect"; you lacked powerReach planned n, or report the CI and say "inconclusive, here is the range"
Decide the hypothesis after seeing results (HARKing)Turns the whole analysis into a fishing expeditionPre-register hypothesis + primary metric before launch
Run 6 days because it "looks significant"Oversamples one weekday slice of the populationRun full weeks; honor the fixed horizon

Checkable artifact

When this skill emits a Python sizing/analysis script or an experiment-design doc, run scripts/verify.sh from your project root. It confirms the script executes under python3 and prints a numeric sample size, and that any design doc names a primary metric, an MDE, and power/alpha. It is read-only and soft-passes when no artifact is present (a design-only conversation).

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/ab-testing of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/pitfalls.md
  • references/sample-size-and-cuped.md
  • scripts/verify.sh

Open the folder on GitHubat commit e3d5b33

Compare with similar skills

Ab Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Testing this skillericrisco/rsc-harness167—~2.4kAutomated safety check: PassMIT
Define Hypothesisproduct-on-purpose/pm-skills715—~966Automated safety check: PassApache-2.0
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
App Analyticsappeeky/aso-skills2.2k—~1.6kAutomated safety check: PassMIT
Analytics Metrics Kpinicepkg/ai-workflow285—~2.1kAutomated safety check: PassMIT
Craft Experiment Designamplitude/builder-skills159—~522Automated safety check: PassNone

Similar skills

  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    715 GitHub stars~966 tokensUpdated 3 days ago
    Marketing & SEOAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • App Analytics

    appeeky/aso-skills

    When the user wants to set up, interpret, or improve their app analytics and tracking.

    2.2k GitHub stars~1.6k tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Analytics Metrics Kpi

    nicepkg/ai-workflow

    Master metrics definition, KPI tracking, dashboarding, A/B testing, and data-driven decision making.

    285 GitHub stars~2.1k tokensUpdated 8 mo ago
    Business, Finance & HRAuto-check passed
  • Craft Experiment Design

    amplitude/builder-skills

    Write a hypothesis, define success metrics, and plan a holdout strategy.

    159 GitHub stars~522 tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Power Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…

    174 GitHub stars~2.2k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from ericrisco/rsc-harness

All 227 skills in this repo
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    167 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    167 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    167 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    167 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    167 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Angular

    ericrisco/rsc-harness

    A skill your agent uses when building, refactoring, or debugging Angular (v20/21+): standalone components, signals, zoneless change detection, @if/@for/@defer control flow, inject() DI…

    167 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Ab Testing

What does Ab Testing do?

A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…. Ab Testing is an agent skill from ericrisco/rsc-harness. Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant.

When should I use Ab Testing?

Ab Testing fits situations like: analyzing a controlled experiment — falsifiable hypothesis; sample size from an MDE; reading significance/CI/power; rescuing tests that wont go significant.

How do I install Ab Testing in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill ab-testing -a claude-code`. Or copy the skill folder (skills/ab-testing in ericrisco/rsc-harness) into .claude/skills/ab-testing in your project. Claude Code loads it when a task matches its description.

How do I install Ab Testing in Codex?

Run `npx skills add ericrisco/rsc-harness --skill ab-testing -a codex`. Or copy the skill folder (skills/ab-testing in ericrisco/rsc-harness) into .agents/skills/ab-testing in your project. Codex loads it when a task matches its description.

Can I use Ab Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill ab-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-testing, .gemini/skills/ab-testing, .github/skills/ab-testing and .opencode/skills/ab-testing in your project.

What does Ab Testing need to run?

Going by SKILL.md and its folder, Ab Testing needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Ab Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ab Testing use?

Ab Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Testing use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Ab Testing?

Skills that share tags, products or a category with Ab Testing: Define Hypothesis (product-on-purpose/pm-skills, 715 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), App Analytics (appeeky/aso-skills, 2.2k stars) and Analytics Metrics Kpi (nicepkg/ai-workflow, 285 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Testing?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.