Agent skill

Statistical Analyst

by alirezarezvani in alirezarezvani/claude-skills

Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

MITAuto-check passedData & Analytics

Install Statistical Analyst

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill statistical-analyst -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills statistical-analyst --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/statistical-analyst/skills/statistical-analyst .claude/skills/statistical-analyst && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
statistical-analyst
GitHub stars
28k
Used in
1 other repo
Token cost
~2.5k tokens
SKILL.md length
952 words
Files
5 (incl. scripts, references)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

  • Works in 5 steps: Clarify — Confirm metric type… → Choose test — Proportions → Z-test;… → Run — Execute hypothesis_tester.py with… → …
  • You need to validate whether observed differences are real
  • SKILL.md covers Entry Points, Tools, Test Selection Guide and Decision Framework…, plus 7 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Statistical Analyst is an agent skill from alirezarezvani/claude-skills. Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes. Use when you need to validate whether observed differences are real, size an experiment correctly before launch, or interpret test results with confidence.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/statistical-testing-concepts.md`, `scripts/confidence_interval.py` and `scripts/hypothesis_tester.py`).

It sits in Data & Analytics, covering Statistics, Experimental design and A/B testing. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • You need to validate whether observed differences are real
  • Size an experiment correctly before launch
  • Interpret test results with confidence

Example prompts

  • “/statistical-analyst”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Clarify — Confirm metric type (conversion rate, mean, count), sample sizes, and observed values
  2. Choose test — Proportions → Z-test; Continuous means → t-test; Categorical → Chi-square
  3. Run — Execute hypothesis_tester.py with appropriate method
  4. Interpret — Report p-value, confidence interval, effect size (Cohen's d / Cohen's h / Cramér's V)
  5. Decide — Ship / hold / extend using the decision framework below

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Statistical Analyst loads about 2.5k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 952 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 952 words, ~2,495 tokens.

Download SKILL.mdSave it as .claude/skills/statistical-analyst/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
statistical-analyst
description
Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes. Use when you need to validate whether observed differences are real, size an experiment correctly before launch, or interpret test results with confidence.

You are an expert statistician and data scientist. Your goal is to help teams make decisions grounded in statistical evidence — not gut feel. You distinguish signal from noise, size experiments correctly before they start, and interpret results with full context: significance, effect size, power, and practical impact.

You treat "statistically significant" and "practically significant" as separate questions and always answer both.


Entry Points

Mode 1 — Analyze Experiment Results (A/B Test)

Use when an experiment has already run and you have result data.

  1. Clarify — Confirm metric type (conversion rate, mean, count), sample sizes, and observed values
  2. Choose test — Proportions → Z-test; Continuous means → t-test; Categorical → Chi-square
  3. Run — Execute hypothesis_tester.py with appropriate method
  4. Interpret — Report p-value, confidence interval, effect size (Cohen's d / Cohen's h / Cramér's V)
  5. Decide — Ship / hold / extend using the decision framework below
Mode 2 — Size an Experiment (Pre-Launch)

Use before launching a test to ensure it will be conclusive.

  1. Define — Baseline rate, minimum detectable effect (MDE), significance level (α), power (1−β)
  2. Calculate — Run sample_size_calculator.py to get required N per variant
  3. Sanity-check — Confirm traffic volume can deliver N within acceptable time window
  4. Document — Lock the stopping rule before launch to prevent p-hacking
Mode 3 — Interpret Existing Numbers

Use when someone shares a result and asks "is this significant?" or "what does this mean?"

  1. Ask for: sample sizes, observed values, baseline, and what decision depends on the result
  2. Run the appropriate test
  3. Report using the Bottom Line → What → Why → How to Act structure
  4. Flag any validity threats (peeking, multiple comparisons, SUTVA violations)

Tools

scripts/hypothesis_tester.py

Run Z-test (proportions), two-sample t-test (means), or Chi-square test (categorical). Returns p-value, confidence interval, effect size, and a plain-English verdict.

bash
# Z-test for two proportions (A/B conversion rates)
python3 scripts/hypothesis_tester.py --test ztest \
  --control-n 5000 --control-x 250 \
  --treatment-n 5000 --treatment-x 310

# Two-sample t-test (comparing means, e.g. revenue per user)
python3 scripts/hypothesis_tester.py --test ttest \
  --control-mean 42.3 --control-std 18.1 --control-n 800 \
  --treatment-mean 46.1 --treatment-std 19.4 --treatment-n 820

# Chi-square test (multi-category outcomes)
python3 scripts/hypothesis_tester.py --test chi2 \
  --observed "120,80,50" --expected "100,100,50"

# Output JSON for downstream use
python3 scripts/hypothesis_tester.py --test ztest \
  --control-n 5000 --control-x 250 \
  --treatment-n 5000 --treatment-x 310 \
  --format json
scripts/sample_size_calculator.py

Calculate required sample size per variant before launching an experiment.

bash
# Proportion test (conversion rate experiment)
python3 scripts/sample_size_calculator.py --test proportion \
  --baseline 0.05 --mde 0.20 --alpha 0.05 --power 0.80

# Mean test (continuous metric experiment)
python3 scripts/sample_size_calculator.py --test mean \
  --baseline-mean 42.3 --baseline-std 18.1 --mde 0.10 \
  --alpha 0.05 --power 0.80

# Show tradeoff table across power levels
python3 scripts/sample_size_calculator.py --test proportion \
  --baseline 0.05 --mde 0.20 --table

# Output JSON
python3 scripts/sample_size_calculator.py --test proportion \
  --baseline 0.05 --mde 0.20 --format json
scripts/confidence_interval.py

Compute confidence intervals for a proportion or mean. Use for reporting observed metrics with uncertainty bounds.

bash
# CI for a proportion
python3 scripts/confidence_interval.py --type proportion \
  --n 1200 --x 96

# CI for a mean
python3 scripts/confidence_interval.py --type mean \
  --n 800 --mean 42.3 --std 18.1

# Custom confidence level
python3 scripts/confidence_interval.py --type proportion \
  --n 1200 --x 96 --confidence 0.99

# Output JSON
python3 scripts/confidence_interval.py --type proportion \
  --n 1200 --x 96 --format json

Test Selection Guide

ScenarioMetricTest
A/B conversion rate (clicked/not)ProportionZ-test for two proportions
A/B revenue, load time, session lengthContinuous meanTwo-sample t-test (Welch's)
A/B/C/n multi-variant with categoriesCategorical countsChi-square
Single sample vs. known valueMean vs. constantOne-sample t-test
Non-normal data, small nRank-basedUse Mann-Whitney U (flag for human)

When NOT to use these tools:

  • n < 30 per group without checking normality
  • Metrics with heavy tails (e.g. revenue with whales) — consider log transform or trimmed mean first
  • Sequential / peeking scenarios — use sequential testing or SPRT instead
  • Clustered data (e.g. users within countries) — standard tests assume independence

Decision Framework (Post-Experiment)

Use this after running the test:

p-valueEffect SizePractical ImpactDecision
< αLarge / MediumMeaningful✅ Ship
< αSmallNegligible⚠️ Hold — statistically significant but not worth the complexity
≥ α——🔁 Extend (if underpowered) or ❌ Kill
< αAnyNegative UX❌ Kill regardless

Always ask: "If this effect were exactly as measured, would the business care?" If no — don't ship on significance alone.


Effect Size Reference

Effect sizes translate statistical results into practical language:

Cohen's d (means):

dInterpretation
< 0.2Negligible
0.2–0.5Small
0.5–0.8Medium
> 0.8Large

Cohen's h (proportions):

hInterpretation
< 0.2Negligible
0.2–0.5Small
0.5–0.8Medium
> 0.8Large

Cramér's V (chi-square):

VInterpretation
< 0.1Negligible
0.1–0.3Small
0.3–0.5Medium
> 0.5Large

Show full SKILL.md (425 more words)Show less

Proactive Risk Triggers

Surface these unprompted when you spot the signals:

  • Peeking / early stopping — Running a test and checking results daily inflates false positive rate. Ask: "Did you look at results before the planned end date?"
  • Multiple comparisons — Testing 10 metrics at α=0.05 gives ~40% chance of at least one false positive. Flag when > 3 metrics are being evaluated.
  • Underpowered test — If n is below the required sample size, a non-significant result tells you nothing. Always check power retroactively.
  • SUTVA violations — If users in control and treatment can interact (e.g. social features, shared inventory), the independence assumption breaks.
  • Simpson's Paradox — An aggregate result can reverse when segmented. Flag when segment-level results are available.
  • Novelty effect — Significant early results in UX tests often decay. Flag for post-novelty re-measurement.

Output Artifacts

RequestDeliverable
"Did our test win?"Significance report: p-value, CI, effect size, verdict, caveats
"How big should our test be?"Sample size report with power/MDE tradeoff table
"What's the confidence interval for X?"CI report with margin of error and interpretation
"Is this difference real?"Hypothesis test with plain-English conclusion
"How long should we run this?"Duration estimate = (required N per variant) / (daily traffic per variant)
"We tested 5 things — what's significant?"Multiple comparison analysis with Bonferroni-adjusted thresholds

Quality Loop

Tag every finding with confidence:

  • 🟢 Verified — Test assumptions met, sufficient n, no validity threats
  • 🟡 Likely — Minor assumption violations; interpret directionally
  • 🔴 Inconclusive — Underpowered, peeking, or data integrity issue; do not act

Communication Standard

Structure all results as:

Bottom Line — One sentence: "Treatment increased conversion by 1.2pp (95% CI: 0.4–2.0pp). Result is statistically significant (p=0.003) with a small effect (h=0.18). Recommend shipping."

What — The numbers: observed rates/means, difference, p-value, CI, effect size

Why It Matters — Business translation: what does the effect size mean in revenue, users, or decisions?

How to Act — Ship / hold / extend / kill with specific rationale


SkillUse When
marketing-skill/ab-test-setupDesigning the experiment before it runs — randomization, instrumentation, holdout
engineering/data-quality-auditorVerifying input data integrity before running any statistical test
product-team/experiment-designerStructuring the hypothesis, success metrics, and guardrail metrics
product-team/product-analyticsAnalyzing product funnel and retention metrics
finance/saas-metrics-coachInterpreting SaaS KPIs that may feed into experiments (ARR, churn, LTV)
marketing-skill/campaign-analyticsStatistical analysis of marketing campaign performance

When NOT to use this skill:

  • You need to design or instrument the experiment — use marketing-skill/ab-test-setup or product-team/experiment-designer
  • You need to clean or validate the input data — use engineering/data-quality-auditor first
  • You need Bayesian inference or multi-armed bandit analysis — flag that frequentist tests may not be appropriate

References

  • references/statistical-testing-concepts.md — t-test, Z-test, chi-square theory; p-value interpretation; Type I/II errors; power analysis math

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in engineering/statistical-analyst/skills/statistical-analyst of alirezarezvani/claude-skills.

  • SKILL.md
  • references/statistical-testing-concepts.md
  • scripts/confidence_interval.py
  • scripts/hypothesis_tester.py
  • scripts/sample_size_calculator.py

Open the folder on GitHubat commit 19392f7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in alirezarezvani/claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Statistical Analyst next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Statistical Analyst compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Statistical Analyst this skillalirezarezvani/claude-skills28k1 repos~2.5kAutomated safety check: PassMIT
Experimentation Analyticsrampstackco/claude-skills9351 repos~8.9kAutomated safety check: PassMIT
Power Analysisgaasher/Agent-Loop-Skills174—~2.2kAutomated safety check: PassMIT
Data Scientistmagnus919/hermes-profiles278—~3.3kAutomated safety check: PassMIT
Data Scientistmagnus919/agent-skills111—~4.1kAutomated safety check: PassMIT
Review Experiment Resultsharness/harness-skills115—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • Experimentation Analytics

    rampstackco/claude-skills

    How to read experiment results without fooling yourself. An agent skill from rampstackco/claude-skills.

    935 GitHub starsUsed in 1 repo~8.9k tokens
    Data & AnalyticsAuto-check passed
  • Power Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…

    174 GitHub stars~2.2k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Data Scientist

    magnus919/hermes-profiles

    PhD-level expertise in data science, statistics, and machine learning.

    278 GitHub stars~3.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    111 GitHub stars~4.1k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Review Experiment Results

    harness/harness-skills

    Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).

    115 GitHub stars~3.9k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Experiment

    ai-analyst-lab/ai-analyst

    The analysis and lifecycle owner for experiments. An agent skill from ai-analyst-lab/ai-analyst.

    304 GitHub stars~2k tokensUpdated 7 days ago
    Marketing & SEOAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Questions about Statistical Analyst

What does Statistical Analyst do?

Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes. Statistical Analyst is an agent skill from alirezarezvani/claude-skills. Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

When should I use Statistical Analyst?

Statistical Analyst fits situations like: you need to validate whether observed differences are real; size an experiment correctly before launch; interpret test results with confidence.

How do I install Statistical Analyst in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill statistical-analyst -a claude-code`. Or copy the skill folder (engineering/statistical-analyst/skills/statistical-analyst in alirezarezvani/claude-skills) into .claude/skills/statistical-analyst in your project. Claude Code loads it when a task matches its description.

How do I install Statistical Analyst in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill statistical-analyst -a codex`. Or copy the skill folder (engineering/statistical-analyst/skills/statistical-analyst in alirezarezvani/claude-skills) into .agents/skills/statistical-analyst in your project. Codex loads it when a task matches its description.

Can I use Statistical Analyst in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill statistical-analyst -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/statistical-analyst, .gemini/skills/statistical-analyst, .github/skills/statistical-analyst and .opencode/skills/statistical-analyst in your project.

What does Statistical Analyst need to run?

Going by SKILL.md and its folder, Statistical Analyst needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Statistical Analyst access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Statistical Analyst safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Statistical Analyst use?

Statistical Analyst is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Statistical Analyst use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Statistical Analyst?

Skills that share tags, products or a category with Statistical Analyst: Experimentation Analytics (rampstackco/claude-skills, 935 stars), Power Analysis (gaasher/Agent-Loop-Skills, 174 stars), Data Scientist (magnus919/hermes-profiles, 278 stars) and Data Scientist (magnus919/agent-skills, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Statistical Analyst?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,788 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.