Define Hypothesis
product-on-purpose/pm-skills
Defines a testable hypothesis with clear success metrics and a validation approach.
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
$ npx skills add ericrisco/rsc-harness --skill ab-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness ab-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-testing .claude/skills/ab-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ab-testing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing into .claude/skills/ab-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill ab-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness ab-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ab-testing .agents/skills/ab-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ab-testing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing into .agents/skills/ab-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill ab-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness ab-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ab-testing .cursor/skills/ab-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ab-testing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing into .cursor/skills/ab-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/ab-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill ab-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness ab-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ab-testing .gemini/skills/ab-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ab-testing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing into .gemini/skills/ab-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness ab-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill ab-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ab-testing .github/skills/ab-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ab-testing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing into .github/skills/ab-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill ab-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness ab-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ab-testing .opencode/skills/ab-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ab-testing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing into .opencode/skills/ab-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ab-testingA skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
Ab Testing is an agent skill from ericrisco/rsc-harness. Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is analytics), NOT north-star/KPI trees (that is kpi-framework), NOT projecting metrics forward (that is forecasting).
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/pitfalls.md`).
It sits in Marketing & SEO, covering A/B testing, Experimental design and Forecasting and time series. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ab Testing loads about 2.4k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 1,193 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,193 words, ~2,438 tokens.
.claude/skills/ab-testing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.An experiment without a pre-committed sample size and a single primary metric is not an experiment. It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost entirely before traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four.
Each one is a place experiments die silently.
State a null you can reject. "The new checkout button changes purchase conversion" with H0: conversion equal across arms, H1: it differs. Vague aspirations ("improve the funnel") have no rejection region.
Pick one primary metric and freeze it. Why: every extra primary metric is another coin flip at α, so three "primary" metrics turn a 5% false-positive rate into roughly 14%. Demote the rest to secondary.
Randomize on the same unit you analyze on. If a user sees the variant on every visit, randomize by user, not by session — analyzing 50k sessions from 8k users treats correlated observations as independent and fabricates significance.
Bad: "We think the redesign will improve engagement and revenue and retention." (no null, 3 primaries, no number)
Good: "H0: 30-day purchase conversion is equal between control and the new one-click button.
H1: it differs. Primary: purchase conversion. Guardrails: refund rate, p95 checkout latency.
Randomize by user_id. MDE: +1.5pp absolute on a 12% baseline."Defaults: power 0.80, α 0.05 (two-sided). The MDE is yours to choose — it is the smallest effect that would actually change what you do.
Rule: required n scales with ~1/MDE². Why: halving the smallest effect you care to detect roughly quadruples the traffic and time. This is the single most expensive decision in the design, so set the MDE to a business threshold, never to "whatever is small."
For a conversion rate (proportion):
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize
p1, p2 = 0.12, 0.135 # baseline, baseline + MDE (1.5pp)
h = proportion_effectsize(p1, p2) # Cohen's h (arcsine transform)
n = NormalIndPower().solve_power(effect_size=h, alpha=0.05, power=0.80, ratio=1.0)
print(int(-(-n // 1))) # n PER ARM, rounded upFor a continuous metric (revenue per user, time on page) use Welch-style sizing:
from statsmodels.stats.power import TTestIndPower
effect = mde_in_units / pooled_std # Cohen's d
n = TTestIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.80, ratio=1.0)Then convert n to a calendar plan: days = ceil((n_per_arm * num_arms) / daily_eligible_users). If that
is 9 days, run a clean two full weeks anyway — weekday/weekend mix is part of the population, and a
6-day test oversamples whoever shows up Tuesday. Full worked example (12% baseline, +1.5pp MDE, 80%
power) plus runnable sizing, n→duration, CUPED θ and SRM snippets: references/sample-size-and-cuped.md.
Fixed horizon is the default. Commit to the n/date from Step 2 and read the result once, at the end.
Do not peek and stop at first significance. Why: checking repeatedly and stopping the moment p < 0.05 inflates the Type-I error far above 5% — with enough looks, a null test crosses 0.05 most of the time. If you genuinely need to stop early, use a sequential / always-valid method (confidence sequences, e.g. Netflix's anytime-valid CIs) that holds Type-I error under continuous monitoring. Sequential is strong for killing losers early and weak for calling winners early — for a confident win, the fixed-horizon read is tighter.
Gate on SRM before you trust anything. Compute a chi-square test on the observed split versus the intended ratio. If p < 0.001 the assignment or logging is broken — a bot filter dropping one arm, a redirect, a caching bug. Fix the instrumentation and rerun; do not "adjust for it."
The peeking Type-I math, sequential/always-valid options, SRM diagnosis, novelty/primacy effects,
Simpson's paradox in segments and HARKing all live in references/pitfalls.md.
Pick the test by metric type:
| Metric type | Test |
|---|---|
| Binary conversion (proportion) | Two-proportion z-test (statsmodels.stats.proportion.proportions_ztest) |
| Continuous, roughly normal / large n | Welch's t-test (scipy.stats.ttest_ind(..., equal_var=False)) |
| Continuous, heavy-tailed / skewed (revenue) | Mann-Whitney U, or t-test on a log/winsorized metric |
Report lift + confidence interval + p-value together. Never p alone. Why: p < 0.05 with a CI of [+0.1pp, +5pp] is "statistically there, practically a coin toss" — the CI tells you the size, p only tells you it is not exactly zero. Practical significance = compare the CI to your MDE: if the whole interval sits above the MDE, ship; if it straddles the MDE, you detected something too small to matter.
Multiple comparisons. Two regimes:
CUPED (Controlled-experiment Using Pre-Experiment Data) subtracts predictable pre-period noise so the same traffic buys more power — or the same power needs less traffic. The adjusted metric:
Y_cuped = Y − θ · (X − E[X]) where θ = Cov(Y, X) / Var(X)Estimate θ by regressing the in-experiment metric Y on the pre-experiment covariate X (e.g. each
user's spend in the 4 weeks before the test), then analyze Y_cuped with the same test as Step 4.
When it pays: recurring users with a strong pre-period signal. Reported wins — Netflix ~40% variance reduction on engagement, Statsig 50%+ on common metrics → significance in roughly half the time/traffic.
When it does nothing — do not bother: brand-new users (no pre-period data), a covariate uncorrelated
with the outcome, or — the cardinal sin — a covariate measured after assignment, which biases the
estimate. The covariate MUST be pre-treatment and independent of which arm a user lands in. Runnable
θ-via-OLS snippet in references/sample-size-and-cuped.md.
| Bad | Why it is wrong | Do instead |
|---|---|---|
| Peek daily, stop the day p < 0.05 | Repeated looks inflate Type-I error far above α | Fix n/date up front; or a sequential method that holds α |
| No sample size set before launch | You will stop on noise and call it a win | Compute n from MDE/baseline/power in Step 2 |
| Several "primary" metrics | Each is a coin flip at α; 3 metrics ≈ 14% false-positive | One frozen primary; the rest are secondary |
| Ignore the observed split | An SRM means assignment/logging is broken; results are garbage | Chi-square SRM gate before reading anything |
| Report only the p-value | Hides effect size — p < 0.05 can be practically zero | Always lift + CI + p; compare CI to MDE |
| CUPED on a post-assignment covariate | Covariate correlated with the arm biases θ | Use only pre-treatment, assignment-independent covariates |
| Call a winner from an underpowered test | "Not significant" then ≠ "no effect"; you lacked power | Reach planned n, or report the CI and say "inconclusive, here is the range" |
| Decide the hypothesis after seeing results (HARKing) | Turns the whole analysis into a fishing expedition | Pre-register hypothesis + primary metric before launch |
| Run 6 days because it "looks significant" | Oversamples one weekday slice of the population | Run full weeks; honor the fixed horizon |
When this skill emits a Python sizing/analysis script or an experiment-design doc, run
scripts/verify.sh from your project root. It confirms the script executes under python3 and prints a
numeric sample size, and that any design doc names a primary metric, an MDE, and power/alpha. It is
read-only and soft-passes when no artifact is present (a design-only conversation).
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/ab-testing of ericrisco/rsc-harness.
Open the folder on GitHubat commit e3d5b33
Ab Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ab Testing this skillericrisco/rsc-harness | 167 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Define Hypothesisproduct-on-purpose/pm-skills | 715 | — | ~966 | Automated safety check: Pass | Apache-2.0 | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | |
| App Analyticsappeeky/aso-skills | 2.2k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Analytics Metrics Kpinicepkg/ai-workflow | 285 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Craft Experiment Designamplitude/builder-skills | 159 | — | ~522 | Automated safety check: Pass | None |
product-on-purpose/pm-skills
Defines a testable hypothesis with clear success metrics and a validation approach.
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
appeeky/aso-skills
When the user wants to set up, interpret, or improve their app analytics and tracking.
nicepkg/ai-workflow
Master metrics definition, KPI tracking, dashboarding, A/B testing, and data-driven decision making.
amplitude/builder-skills
Write a hypothesis, define success metrics, and plan a holdout strategy.
gaasher/Agent-Loop-Skills
A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
ericrisco/rsc-harness
A skill your agent uses when building, refactoring, or debugging Angular (v20/21+): standalone components, signals, zoneless change detection, @if/@for/@defer control flow, inject() DI…
Categories
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…. Ab Testing is an agent skill from ericrisco/rsc-harness. Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant.
Ab Testing fits situations like: analyzing a controlled experiment — falsifiable hypothesis; sample size from an MDE; reading significance/CI/power; rescuing tests that wont go significant.
Run `npx skills add ericrisco/rsc-harness --skill ab-testing -a claude-code`. Or copy the skill folder (skills/ab-testing in ericrisco/rsc-harness) into .claude/skills/ab-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill ab-testing -a codex`. Or copy the skill folder (skills/ab-testing in ericrisco/rsc-harness) into .agents/skills/ab-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill ab-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-testing, .gemini/skills/ab-testing, .github/skills/ab-testing and .opencode/skills/ab-testing in your project.
Going by SKILL.md and its folder, Ab Testing needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Ab Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ab Testing: Define Hypothesis (product-on-purpose/pm-skills, 715 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), App Analytics (appeeky/aso-skills, 2.2k stars) and Analytics Metrics Kpi (nicepkg/ai-workflow, 285 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.