Agent skill

Ab Test Analysis

by nimrodfisher in nimrodfisher/data-analytics-skills

Rigorous A/B test statistical analysis. An agent skill from nimrodfisher/data-analytics-skills.

MITAuto-check passedMarketing & SEO

Install Ab Test Analysis

skills CLI
$ npx skills add nimrodfisher/data-analytics-skills --skill ab-test-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nimrodfisher/data-analytics-skills ab-test-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nimrodfisher/data-analytics-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/03-data-analysis-investigation/ab-test-analysis .claude/skills/ab-test-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-analysis
GitHub stars
468
Token cost
~708 tokens
SKILL.md length
337 words
Files
4 (incl. scripts, references, assets)
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Rigorous A/B test statistical analysis. An agent skill from nimrodfisher/data-analytics-skills.

  • Works in 6 steps: Confirm test design — verify the… → Check for sample ratio mismatch (SRM) —… → Calculate per-variant metrics — compute… → …
  • Analyzing experiment results
  • Runs Python scripts from its folder
  • Calculating statistical significance

What it does

Ab Test Analysis is an agent skill from nimrodfisher/data-analytics-skills. Rigorous A/B test statistical analysis. Use when analyzing experiment results, calculating statistical significance, checking for sample ratio mismatch, or validating test design before launch.

Its SKILL.md is about 710 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts, reference files and assets (for example `assets/ab_test_report_template.md`, `references/ab_test_design_guide.md` and `scripts/ab_test_analyzer.py`).

It sits in Marketing & SEO, covering A/B testing and Statistics. The repository describes itself as: A comprehensive list of Claude & Codex skills for a wide range of data analytics tasks. The licence is MIT.

When your agent uses it

  • Analyzing experiment results
  • Calculating statistical significance
  • Checking for sample ratio mismatch
  • Validating test design before launch

Example prompts

  • “/ab-test-analysis”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Confirm test design — verify the hypothesis, the control and treatment definitions, the randomisation unit (user/session/device), the…
  2. Check for sample ratio mismatch (SRM) — run a chi-square test on the actual vs. expected split. If SRM is detected, stop and investigate…
  3. Calculate per-variant metrics — compute the rate (or mean) and 95% confidence interval for the primary metric in each variant. Document…
  4. Run the significance test — execute a two-proportion z-test (for rates) or Welch's t-test (for means). Record z-score, p-value, and 95% CI…
  5. Check guardrail metrics — run the same significance test for each guardrail metric. A significant degradation on any guardrail is a…
  6. Produce the recommendation — synthesise SRM result, power, significance, and guardrail checks into a clear ship / no-ship / extend…

What it can do on your machine

Read from SKILL.md and the folder at commit 9449d36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Test Analysis loads about 708 tokens when it runs, and up to ~1.5k if it reads all its reference files. Until then it costs about 53 tokens; SKILL.md has 337 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~708
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from nimrodfisher/data-analytics-skills at commit 9449d36, republished under its MIT licence (© nimrodfisher). 337 words, ~708 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-analysis/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ab-test-analysis
description
Rigorous A/B test statistical analysis. Use when analyzing experiment results, calculating statistical significance, checking for sample ratio mismatch, or validating test design before launch.

A/B Test Analysis

When to use

  • An experiment has finished and the team needs a ship / no-ship recommendation
  • Results look directionally positive but the team is unsure if they're statistically significant
  • A test has been running for weeks without a clear winner and someone needs to decide whether to continue
  • A new experiment needs sample-size planning before launch
  • Results are disputed and need a rigorous, documented analysis

Process

  1. Confirm test design — verify the hypothesis, the control and treatment definitions, the randomisation unit (user/session/device), the primary metric, any guardrail metrics, and the target split ratio.
  2. Check for sample ratio mismatch (SRM) — run a chi-square test on the actual vs. expected split. If SRM is detected, stop and investigate the randomisation pipeline before interpreting results. Use scripts/ab_test_analyzer.py --check-srm.
  3. Calculate per-variant metrics — compute the rate (or mean) and 95% confidence interval for the primary metric in each variant. Document absolute and relative difference.
  4. Run the significance test — execute a two-proportion z-test (for rates) or Welch's t-test (for means). Record z-score, p-value, and 95% CI for the effect. Use references/statistical_tests_reference.md if unsure which test applies.
  5. Check guardrail metrics — run the same significance test for each guardrail metric. A significant degradation on any guardrail is a blocker regardless of primary metric results.
  6. Produce the recommendation — synthesise SRM result, power, significance, and guardrail checks into a clear ship / no-ship / extend decision. Quantify the expected business impact if shipped. Record in assets/ab_test_report_template.md.

Inputs the skill needs

  • Test plan or hypothesis document (variant definitions, randomisation unit, primary metric)
  • Data with at minimum: user_id, variant assignment, primary metric outcome
  • Optional: guardrail metric values per user, daily aggregate data for temporal validity checks
  • Target split ratio (e.g., 50/50)
  • Minimum detectable effect or business threshold for "worth shipping"

Output

  • scripts/ab_test_analyzer.py — runs SRM check, significance test, power analysis, and guardrail checks from a CSV or summary stats input
  • references/statistical_tests_reference.md — which test to use and when
  • references/ab_test_design_guide.md — SRM causes, power planning, peeking and multiple testing
  • assets/ab_test_report_template.md — structured report: design, results, checks, recommendation, expected impact

© nimrodfisher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references, assets) in 03-data-analysis-investigation/ab-test-analysis of nimrodfisher/data-analytics-skills.

  • SKILL.md
  • assets/ab_test_report_template.md
  • references/ab_test_design_guide.md
  • scripts/ab_test_analyzer.py

Open the folder on GitHubat commit 9449d36

Compare with similar skills

Ab Test Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Test Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Test Analysis this skillnimrodfisher/data-analytics-skills468—~708Automated safety check: PassMIT
Measure Experiment Resultsproduct-on-purpose/pm-skills715—~989Automated safety check: PassApache-2.0
Mkt Experimentevolution-foundation/evo-nexus545—~1.2kAutomated safety check: PassCustom licence
A/B Test Analysisphuryn/pm-skills27k—~893Automated safety check: PassMIT
Experimentai-analyst-lab/ai-analyst304—~2kAutomated safety check: PassMIT
Statistical Analystalirezarezvani/claude-skills28k—~2.5kAutomated safety check: PassMIT

Similar skills

  • Measure Experiment Results

    product-on-purpose/pm-skills

    Documents the results of a completed experiment or A/B test with statistical analysis, learnings, and recommendations.

    715 GitHub stars~989 tokensUpdated today
    Marketing & SEOAuto-check passed
  • Mkt Experiment

    evolution-foundation/evo-nexus

    Autonomous growth experimentation framework. An agent skill from evolution-foundation/evo-nexus.

    545 GitHub stars~1.2k tokensUpdated 4 mo ago
    Marketing & SEOAuto-check passed
  • A/B Test Analysis

    phuryn/pm-skills

    Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

    27k GitHub stars~893 tokensUpdated 24 days ago
    Data & AnalyticsAuto-check passed
  • Experiment

    ai-analyst-lab/ai-analyst

    The analysis and lifecycle owner for experiments. An agent skill from ai-analyst-lab/ai-analyst.

    304 GitHub stars~2k tokensUpdated 8 days ago
    Marketing & SEOAuto-check passed
  • Statistical Analyst

    alirezarezvani/claude-skills

    Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

    28k GitHub stars~2.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Experimentation Analytics

    rampstackco/claude-skills

    How to read experiment results without fooling yourself. An agent skill from rampstackco/claude-skills.

    941 GitHub stars~8.9k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed

More from nimrodfisher/data-analytics-skills

All 31 skills in this repo
  • Analysis Assumptions Log

    nimrodfisher/data-analytics-skills

    Track and document analytical assumptions and decisions. An agent skill from nimrodfisher/data-analytics-skills.

    468 GitHub stars~578 tokensUpdated 14 days ago
    Auto-check passed
  • Analysis QA Checklist

    nimrodfisher/data-analytics-skills

    Pre-delivery quality assurance for analysis work. An agent skill from nimrodfisher/data-analytics-skills.

    468 GitHub stars~470 tokensUpdated 14 days ago
    Auto-check passed
  • Business Metrics Calculator

    nimrodfisher/data-analytics-skills

    Standard business metric calculation with industry benchmarks.

    468 GitHub stars~668 tokensUpdated 14 days ago
    Auto-check passed
  • Cohort Analysis

    nimrodfisher/data-analytics-skills

    Time-based cohort analysis with retention and behaviour tracking.

    468 GitHub stars~660 tokensUpdated 14 days ago
    Auto-check passed
  • Context Packager

    nimrodfisher/data-analytics-skills

    Efficiently package context for AI-assisted analysis. An agent skill from nimrodfisher/data-analytics-skills.

    468 GitHub stars~500 tokensUpdated 14 days ago
    Auto-check passed
  • Data Catalog Entry

    nimrodfisher/data-analytics-skills

    Create standardized metadata for data assets. An agent skill from nimrodfisher/data-analytics-skills.

    468 GitHub stars~616 tokensUpdated 14 days ago
    Auto-check passed

Questions about Ab Test Analysis

What does Ab Test Analysis do?

Rigorous A/B test statistical analysis. An agent skill from nimrodfisher/data-analytics-skills. Ab Test Analysis is an agent skill from nimrodfisher/data-analytics-skills. Rigorous A/B test statistical analysis.

When should I use Ab Test Analysis?

Ab Test Analysis fits situations like: analyzing experiment results; calculating statistical significance; checking for sample ratio mismatch; validating test design before launch.

How do I install Ab Test Analysis in Claude Code?

Run `npx skills add nimrodfisher/data-analytics-skills --skill ab-test-analysis -a claude-code`. Or copy the skill folder (03-data-analysis-investigation/ab-test-analysis in nimrodfisher/data-analytics-skills) into .claude/skills/ab-test-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Ab Test Analysis in Codex?

Run `npx skills add nimrodfisher/data-analytics-skills --skill ab-test-analysis -a codex`. Or copy the skill folder (03-data-analysis-investigation/ab-test-analysis in nimrodfisher/data-analytics-skills) into .agents/skills/ab-test-analysis in your project. Codex loads it when a task matches its description.

Can I use Ab Test Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nimrodfisher/data-analytics-skills --skill ab-test-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-analysis, .gemini/skills/ab-test-analysis, .github/skills/ab-test-analysis and .opencode/skills/ab-test-analysis in your project.

What does Ab Test Analysis need to run?

Going by SKILL.md and its folder, Ab Test Analysis needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Ab Test Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Test Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ab Test Analysis use?

Ab Test Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Test Analysis use?

About 708 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 822 tokens, read only when the agent opens those files.

What are the alternatives to Ab Test Analysis?

Skills that share tags, products or a category with Ab Test Analysis: Measure Experiment Results (product-on-purpose/pm-skills, 715 stars), Mkt Experiment (evolution-foundation/evo-nexus, 545 stars), A/B Test Analysis (phuryn/pm-skills, 27k stars) and Experiment (ai-analyst-lab/ai-analyst, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Test Analysis?

nimrodfisher (a GitHub user) maintains it in nimrodfisher/data-analytics-skills, which has 468 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on September 25, 2026.

Source: nimrodfisher/data-analytics-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.