Agent skill

Stat Ab Testing

by asgard-ai-platform in asgard-ai-platform/skills

Design and analyze A/B tests with proper statistical methodology including sample size calculation, randomization, frequentist and Bayesian approaches, and sequential testing.

MITAuto-check passedMarketing & SEO

Install Stat Ab Testing

skills CLI
$ npx skills add asgard-ai-platform/skills --skill stat-ab-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install asgard-ai-platform/skills stat-ab-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/asgard-ai-platform/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/stat-ab-testing .claude/skills/stat-ab-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stat-ab-testing
GitHub stars
241
Token cost
~1k tokens
SKILL.md length
332 words
Files
4 (incl. references)
Skills in repo
209
Repo updated
First seen
Licence
MIT

At a glance

Design and analyze A/B tests with proper statistical methodology including sample size calculation, randomization, frequentist and Bayesian approaches, and sequential testing.

  • Works in 7 steps: Hypothesis: What do you expect to happen… → Primary metric: ONE key metric… → Guardrail metrics: Metrics that must NOT… → …
  • The user needs to set up an experiment
  • SKILL.md covers Framework, Output Format, Gotchas and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stat Ab Testing is an agent skill from asgard-ai-platform/skills. Design and analyze A/B tests with proper statistical methodology including sample size calculation, randomization, frequentist and Bayesian approaches, and sequential testing. Use this skill when the user needs to set up an experiment, calculate required sample size, interpret test results, or decide between testing methodologies — even if they say 'should we A/B test this', 'how many users do we need', 'is the test result conclusive', or 'can we stop the test early'.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `examples/sample_scenario.md`, `references/bandits.md` and `references/bayesian-ab.md`).

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: 301 open-source coding agent skills across 22 domains — methodology, judgment & gotchas packaged as Claude Agent Skills for the Asgard AI Platform. The licence is MIT.

When your agent uses it

  • The user needs to set up an experiment
  • Calculate required sample size
  • Interpret test results
  • Decide between testing methodologies — even if they say should we A/B test this

Example prompts

  • “should we A/B test this”
  • “how many users do we need”
  • “is the test result conclusive”
  • “/stat-ab-testing”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Hypothesis: What do you expect to happen and why?
  2. Primary metric: ONE key metric (conversion, revenue, retention)
  3. Guardrail metrics: Metrics that must NOT degrade (page load time, error rate)
  4. Randomization unit: User, session, or device?
  5. Sample size: Calculated from baseline, MDE, α, power
  6. Duration: Account for weekly cycles (minimum 1-2 full weeks)
  7. Stopping rules: Pre-defined — do NOT peek and stop early without correction

What it can do on your machine

Read from SKILL.md and the folder at commit 4e7f4f8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stat Ab Testing loads about 1k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 122 tokens; SKILL.md has 332 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from asgard-ai-platform/skills at commit 4e7f4f8, republished under its MIT licence (© asgard-ai-platform). 332 words, ~1,049 tokens.

Download SKILL.mdSave it as .claude/skills/stat-ab-testing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
stat-ab-testing
description
Design and analyze A/B tests with proper statistical methodology including sample size calculation, randomization, frequentist and Bayesian approaches, and sequential testing. Use this skill when the user needs to set up an experiment, calculate required sample size, interpret test results, or decide between testing methodologies — even if they say 'should we A/B test this', 'how many users do we need', 'is the test result conclusive', or 'can we stop the test early'.
metadata.category
WP-21 設計/資訊/傳播/公衛
metadata.tags
statistics, ab-testing, experimentation

A/B Testing Statistics

Framework

IRON LAW: Calculate Sample Size BEFORE Running the Test

Running a test without knowing the required sample size leads to two
failures: stopping too early (false positives) or running too long (waste).

Required inputs: baseline conversion rate, minimum detectable effect (MDE),
significance level (α), power (1-β). Calculate BEFORE starting.
Sample Size Formula (Proportions)
n per group ≈ (Z_α/2 + Z_β)² × [p₁(1-p₁) + p₂(1-p₂)] / (p₁ - p₂)²

Quick reference (α=0.05, power=0.8):

Baseline RateMDE (relative)N per Group
5%10% (→5.5%)~58,000
5%20% (→6.0%)~15,000
10%10% (→11%)~15,000
10%20% (→12%)~4,000
Testing Approaches
ApproachHow It WorksBest When
Frequentist (fixed-horizon)Set sample size, run to completion, then analyzeStandard practice, well-understood
BayesianUpdate beliefs with data, compute probability of improvementWant probability statements ("90% chance B is better")
Sequential testingCheck results at intervals with adjusted thresholdsNeed to stop early if clear winner, or limit downside risk
Experiment Design Checklist
  1. Hypothesis: What do you expect to happen and why?
  2. Primary metric: ONE key metric (conversion, revenue, retention)
  3. Guardrail metrics: Metrics that must NOT degrade (page load time, error rate)
  4. Randomization unit: User, session, or device?
  5. Sample size: Calculated from baseline, MDE, α, power
  6. Duration: Account for weekly cycles (minimum 1-2 full weeks)
  7. Stopping rules: Pre-defined — do NOT peek and stop early without correction
Analysis Steps
  1. Check randomization balance (are groups comparable on pre-treatment metrics?)
  2. Calculate observed difference and confidence interval
  3. Run significance test (z-test for proportions, t-test for continuous)
  4. Check guardrail metrics
  5. Interpret with practical significance in mind

Output Format

markdown
# A/B Test Design: {Experiment Name}

## Hypothesis
- H₀: {no difference}
- H₁: {expected improvement}
- Primary metric: {metric}
- MDE: {X% relative}

## Sample Size
- Baseline rate: {X%}
- Required N per group: {N}
- Estimated duration: {days/weeks}

## Results (post-test)
| Metric | Control | Treatment | Diff | CI (95%) | p-value |
|--------|---------|-----------|------|----------|---------|
| {primary} | X% | X% | +X% | [X, X] | {value} |

## Decision
{Ship / Don't ship / Extend test} — {rationale}

Gotchas

  • Peeking inflates false positives: Checking results daily and stopping when p < 0.05 can produce a 30%+ false positive rate. Use sequential testing methods if you need to peek.
  • Novelty effect: New features may show a lift that fades as users get used to them. Run tests long enough (2+ weeks) to stabilize.
  • Simpson's paradox: An overall positive result can be negative in every subgroup (or vice versa). Segment by key dimensions.
  • Network effects / interference: If treatment users interact with control users (social features, marketplace), independence is violated. Use cluster randomization.
  • Statistical significance threshold is arbitrary: α=0.05 is convention, not truth. For high-stakes decisions (pricing, major UX changes), consider α=0.01.

References

  • For Bayesian A/B testing methodology, see references/bayesian-ab.md
  • For multi-armed bandit approach, see references/bandits.md

© asgard-ai-platform, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in stat-ab-testing of asgard-ai-platform/skills.

  • SKILL.md
  • examples/sample_scenario.md
  • references/bandits.md
  • references/bayesian-ab.md

Open the folder on GitHubat commit 4e7f4f8

Compare with similar skills

Stat Ab Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stat Ab Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stat Ab Testing this skillasgard-ai-platform/skills241—~1kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills3.9k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills715—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Ab Test Planindranilbanerjee/digital-marketing-pro8551 repos~2kAutomated safety check: PassMIT

Similar skills

  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    715 GitHub stars~966 tokensUpdated 4 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Ab Test Plan

    indranilbanerjee/digital-marketing-pro

    Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…

    855 GitHub starsUsed in 1 repo~2k tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated today
    Marketing & SEOAuto-check passed

More from asgard-ai-platform/skills

All 209 skills in this repo
  • Algo Sc Eoq

    asgard-ai-platform/skills

    Calculate Economic Order Quantity to minimize total inventory cost (ordering + holding).

    241 GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Algo Ecom Bm25

    asgard-ai-platform/skills

    Implement BM25 ranking function for e-commerce product search relevance scoring.

    241 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Mfg Cpk

    asgard-ai-platform/skills

    Calculate Cpk process capability index to assess whether a process meets specification requirements.

    241 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Price Elasticity

    asgard-ai-platform/skills

    Calculate price elasticity of demand to quantify how price changes affect sales volume.

    241 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Bayesian

    asgard-ai-platform/skills

    Apply Bayesian averaging to rank items by combining observed ratings with prior expectations.

    241 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Elo

    asgard-ai-platform/skills

    Implement Elo rating system to rank items or players from pairwise comparison outcomes.

    241 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Stat Ab Testing

What does Stat Ab Testing do?

Design and analyze A/B tests with proper statistical methodology including sample size calculation, randomization, frequentist and Bayesian approaches, and sequential testing. Stat Ab Testing is an agent skill from asgard-ai-platform/skills. Design and analyze A/B tests with proper statistical methodology including sample size calculation, randomization, frequentist and Bayesian approaches, and sequential testing.

When should I use Stat Ab Testing?

Stat Ab Testing fits situations like: the user needs to set up an experiment; calculate required sample size; interpret test results; decide between testing methodologies — even if they say should we A/B test this.

How do I install Stat Ab Testing in Claude Code?

Run `npx skills add asgard-ai-platform/skills --skill stat-ab-testing -a claude-code`. Or copy the skill folder (stat-ab-testing in asgard-ai-platform/skills) into .claude/skills/stat-ab-testing in your project. Claude Code loads it when a task matches its description.

How do I install Stat Ab Testing in Codex?

Run `npx skills add asgard-ai-platform/skills --skill stat-ab-testing -a codex`. Or copy the skill folder (stat-ab-testing in asgard-ai-platform/skills) into .agents/skills/stat-ab-testing in your project. Codex loads it when a task matches its description.

Can I use Stat Ab Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add asgard-ai-platform/skills --skill stat-ab-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stat-ab-testing, .gemini/skills/stat-ab-testing, .github/skills/stat-ab-testing and .opencode/skills/stat-ab-testing in your project.

What does Stat Ab Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Stat Ab Testing is instructions for the agent only.

Does Stat Ab Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Stat Ab Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stat Ab Testing use?

Stat Ab Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stat Ab Testing use?

About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.

What are the alternatives to Stat Ab Testing?

Skills that share tags, products or a category with Stat Ab Testing: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.9k stars), Define Hypothesis (product-on-purpose/pm-skills, 715 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stat Ab Testing?

asgard-ai-platform (a GitHub organization) maintains it in asgard-ai-platform/skills, which has 241 GitHub stars. The repository holds 209 skills in this directory. The repository was last updated on June 6, 2026.

Source: asgard-ai-platform/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.