Agent skill

Statistical Analyst

by borghei in borghei/Claude-Skills

Applied statistics for business and product questions — test selection, assumption checks, power planning, effect sizes with intervals, multiplicity correction.

MITAuto-check passedData & Analytics

Install Statistical Analyst

skills CLI
$ npx skills add borghei/Claude-Skills --skill statistical-analyst -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills statistical-analyst --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-analytics/statistical-analyst .claude/skills/statistical-analyst && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
statistical-analyst
GitHub stars
881
Token cost
~3.5k tokens
SKILL.md length
1,798 words
Files
13 (incl. scripts, references, assets)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

Applied statistics for business and product questions — test selection, assumption checks, power planning, effect sizes with intervals, multiplicity correction.

  • Works in 5 steps: Write down the business question and the… → Fill in the outcome type, group… → Run the selector. It returns one test,… → …
  • Interpreting an experiment
  • SKILL.md covers When to use this skill, Inputs the skill expects, Clarify First and Workflows, plus 3 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Statistical Analyst is an agent skill from borghei/Claude-Skills. Applied statistics for business and product questions — test selection, assumption checks, power planning, effect sizes with intervals, multiplicity correction. Use when interpreting an experiment, sizing a study, or vetting a claim.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts, reference files and assets (for example `assets/experiment_design_template.md`, `assets/sample_contingency.json` and `assets/sample_experiment.json`).

It sits in Data & Analytics, covering Statistics and Data analysis. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Interpreting an experiment
  • Vetting a claim

Example prompts

  • “/statistical-analyst”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Write down the business question and the decision it changes. If the answer
  2. Fill in the outcome type, group structure, baseline, and the minimum effect
  3. Run the selector. It returns one test, its load-bearing assumptions, the
  4. If the achievable sample is below the required sample, say so before
  5. Record everything in assets/experiment_design_template.md.

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Statistical Analyst loads about 3.5k tokens when it runs, and up to ~8.3k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 1,798 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 1,798 words, ~3,458 tokens.

Download SKILL.mdSave it as .claude/skills/statistical-analyst/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
statistical-analyst
description
Applied statistics for business and product questions — test selection, assumption checks, power planning, effect sizes with intervals, multiplicity correction. Use when interpreting an experiment, sizing a study, or vetting a claim.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
data-analytics
metadata.domain
statistics
metadata.updated
2026-07-21
metadata.tags
statistics, hypothesis-testing, effect-size, power-analysis, experimentation

Statistical Analyst

Most bad statistics in business are not arithmetic errors. They are the wrong test on the right data, a null result reported as "no difference," a p-value mistaken for an effect size, or twenty comparisons run and the one that cleared 0.05 written up. This skill covers the applied path: pick the test the data shape actually calls for, check the assumptions that carry weight, size the study before running it, report effects with intervals rather than bare p-values, and say what you found to people who do not want a statistics lecture.

Everything here runs on the Python standard library — the t, chi-square, and normal distributions are implemented directly, so there is no scipy dependency between a question and an answer. Because those implementations are hand-rolled, stats_core.py --selftest verifies all of them against published reference values; run it once before trusting any result.

When to use this skill

  • An A/B test finished and someone needs to know whether to ship
  • A study is being designed and nobody has computed how many observations it needs
  • A stakeholder is quoting a p-value as if it were an effect size
  • A "no difference" result is about to be reported from a study that was underpowered
  • Several variants, metrics, or segments were compared and no correction was applied
  • The data are skewed or outlier-heavy and the default t test is about to be run anyway

Inputs the skill expects

  • The business question and what decision it will change
  • The outcome variable and its type (binary, continuous, count, ordinal, categorical)
  • Group structure — how many groups, independent or paired, randomized or observational
  • Observation counts per group, and the unit of measurement (user, session, order)
  • The baseline value and the smallest effect worth acting on
  • How many comparisons are in the family, and which metric was named primary before the data arrived

Clarify First

Before analyzing, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • The unit of measurement, and whether each unit appears once — clustering (many sessions per user counted as independent rows) deflates standard errors and manufactures significance; it invalidates every test below
  • The smallest effect worth acting on — without it there is no way to size the study or to say whether a significant result matters
  • How many comparisons are in the family, and which metric was primary — decides the correction and whether the result is confirmatory or exploratory
  • Whether anyone has already looked at the data while it accumulated — peeking invalidates a fixed-horizon test and changes the whole approach

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Workflows

Workflow 1 — Choose the test and size the study before collecting data
  1. Write down the business question and the decision it changes. If the answer changes nothing, stop — do not run the study.
  2. Fill in the outcome type, group structure, baseline, and the minimum effect worth shipping into a question spec.
  3. Run the selector. It returns one test, its load-bearing assumptions, the fallback when they fail, and the sample size the stated effect requires.
  4. If the achievable sample is below the required sample, say so before running. An underpowered study should be a conscious decision, not a discovery at write-up time.
  5. Record everything in assets/experiment_design_template.md.
bash
python3 data-analytics/statistical-analyst/scripts/test_selector.py \
  --input data-analytics/statistical-analyst/assets/sample_question.json \
  --power 0.9
Workflow 2 — Run the test and report the effect, not the p-value
  1. Check independence first: count units versus count rows. More rows than units means clustering, and no test below is valid until that is handled.
  2. Inspect skew and outliers (mean vs median, top and bottom five values). If the data are heavy-tailed at small n, switch to Mann-Whitney.
  3. Run the test with the number of comparisons in the family declared, so the threshold is corrected.
  4. Read the effect size and the interval first. Report the effect in domain units, the interval in the same units, and the decision implication.
bash
python3 data-analytics/statistical-analyst/scripts/stats_core.py --selftest

python3 data-analytics/statistical-analyst/scripts/run_test.py \
  --input data-analytics/statistical-analyst/assets/sample_experiment.json \
  --comparisons 3
Workflow 3 — Audit someone else's statistical claim
  1. Ask what the unit of measurement was and whether units repeat. This finds more real errors than every distributional check combined.
  2. Ask how many comparisons were run in total, including the ones not reported, and whether the primary metric was named before the data arrived.
  3. Re-run the test from the raw counts, with the true comparison count.
  4. If the claim is a null result, compute the interval and state the largest effect the study could have missed — "no difference" and "we could not tell" look identical in a significance test and are opposite conclusions.
bash
python3 data-analytics/statistical-analyst/scripts/run_test.py \
  --input data-analytics/statistical-analyst/assets/sample_revenue.json \
  --test welch_t --comparisons 4 --format json

Decision frameworks

Test selection
QuestionOutcomeGroupsTestEffect size
DifferenceBinary2 independent[PROVEN] Two-proportion zAbsolute difference; Cohen's h
DifferenceBinary2 paired[PROVEN] McNemarOdds ratio on discordant pairs
DifferenceBinary/categorical3+[PROVEN] Chi-square of independenceCramér's V
DifferenceContinuous, symmetric2 independent[PROVEN] Welch's tMean difference; Hedges' g
DifferenceContinuous, skewed or n<152 independent[PROVEN] Mann-Whitney URank-biserial r
DifferenceContinuous3+[RECOMMENDED] One-way ANOVAEta-squared
DifferenceCount per exposure2[RECOMMENDED] Poisson rate ratioRate ratio
AssociationTwo continuous—[PROVEN] Pearson, or Spearman if skewedr, r²
Change over timeAny—[RECOMMENDED] Interrupted time seriesLevel and slope change

Use Welch's t, never Student's t, as the two-group default. It does not assume equal variances and costs a fraction of a degree of freedom when they are equal. Testing for equal variance first and then choosing is worse than always using Welch — the pre-test inflates the error rate of the whole procedure.

Reading an interval against your decision threshold
Interval positionReadingAction
Entirely above the thresholdReal and big enoughShip
Above zero, straddles the thresholdReal, unclear if it clears the barCollect more, or decide on cost
Straddles zero, narrowGenuinely no meaningful effectDo not ship — and this is the only case where "no difference" is honest
Straddles zero, wideStudy could not answer the questionReport as inconclusive, state the upper bound

The last two are identical in a significance test and are opposite conclusions. That is the strongest single argument for reporting intervals.

Sample size reality check

Per group, α = 0.05, power = 0.80:

Baseline rateRelative effect to detectn per group
2%+10%~78,000
8%+10%~28,500
8%+25%~4,900
20%+10%~9,000
20%+25%~1,600

Most product experiments are sized by "how long can we wait," which is how underpowered studies get written up as "no difference."

Show full SKILL.md (714 more words)Show less

Anti-Patterns

P-hacking by exploration

Mistake: Twenty metrics, six segments, and three time windows get compared; the one combination that clears p < 0.05 becomes the headline. Why it happens: It rarely feels like cheating. Each individual comparison is a reasonable question, the analyst is genuinely curious, and the tooling makes slicing free. Nobody counts the comparisons because nobody wrote them down. Instead: Name one primary metric before the data arrive and pre-register the subgroups you will examine. Everything else is exploratory, gets Benjamini-Hochberg correction, and is reported as hypothesis-generating rather than decisive. With α = 0.05 and 20 uncorrected comparisons, the chance of at least one false positive is 64% — a coin flip dressed as a finding.

Peeking at a running experiment

Mistake: The dashboard is checked daily and the test is stopped the moment p dips below 0.05. Why it happens: The data are right there, stopping early saves time and traffic, and each individual look feels harmless. The intuition that "more data can only help" is exactly backwards here. Instead: Fix the horizon, compute n up front, and do not look — or use a method built for continuous monitoring (group sequential with O'Brien-Fleming spending, or always-valid confidence sequences). Repeated peeking at an uncorrected fixed-horizon test drives the real false-positive rate to 20-30%: a random walk crosses the threshold eventually even when nothing is happening. If it has already happened, report the result as exploratory and re-run with a fixed horizon.

Reading a null result as "no effect"

Mistake: p = 0.31, so the memo says the change made no difference and the feature is killed. Why it happens: "Not significant" sounds like "no effect," and the alternative sentence — "we ran a study that could not answer the question" — is uncomfortable to write. Instead: Report the interval. If it runs from −0.2% to +4.1%, the study is consistent with a substantial gain and has ruled out almost nothing. State the largest effect you could have missed. Only a narrow interval around zero supports "no meaningful effect," and that distinction is invisible in the p-value.

Confusing significance with importance

Mistake: At n = 400,000 a 0.02% conversion difference reaches p < 0.001 and gets a roadmap slot. Why it happens: p-values conflate effect size with sample size, so at large n everything is significant. The number looks impressive precisely because the sample is large. Instead: Set the decision threshold from the economics — cost to ship divided by value per unit — before the analysis. Then compare the interval to that threshold, not to zero. The reciprocal error matters equally: at small n, an important effect can miss significance and get discarded.

Ignoring the unit of measurement

Mistake: A test on 50,000 sessions from 4,000 users treats every session as an independent observation. Why it happens: The event table has 50,000 rows, the tooling counts rows, and the resulting p-value is gratifyingly small. It is invisible unless someone explicitly compares row count to unit count. Instead: Aggregate to the unit of assignment before testing, or use a cluster-robust method. Treating clustered rows as independent can understate the standard error several-fold and turn pure noise into a highly significant result. This is the single most common invalidating error in applied product analytics, and the cheapest to check.

Files

FilePurpose
scripts/test_selector.pyRecommends one test from question type, outcome type, group structure, and distribution; returns assumptions, fallback, required sample size for a stated MDE, and design warnings
scripts/run_test.pyRuns two-proportion z, Welch's t, chi-square, and Mann-Whitney with effect sizes, confidence intervals, Bonferroni-adjusted thresholds, and assumption warnings
scripts/test_impl.pyThe four test implementations behind run_test.py, with their effect-size magnitude readings; --selftest verifies each against a worked example
scripts/stats_core.pyNormal, Student's t, and chi-square distributions plus the Wilson interval, in stdlib math only; --selftest verifies all 18 against published reference values
references/test-selection-and-assumptions.mdSelection tree, which assumptions are load-bearing and how each fails, power formulas and sizing tables, multiple-comparison corrections, sequential testing
references/effect-sizes-and-communication.mdEffect sizes per test with thresholds, interval choice and interpretation, language for non-statisticians, practical-vs-statistical significance, reporting checklist
assets/sample_question.jsonQuestion spec for the selector — an underpowered 3-variant conversion test
assets/sample_experiment.jsonTwo-proportion conversion data
assets/sample_revenue.jsonContinuous order-value data for Welch's t
assets/sample_contingency.json3x4 contingency table for chi-square
assets/sample_session_times.jsonRight-skewed time-on-task data for Mann-Whitney
assets/experiment_design_template.mdPre-registration template: question, primary metric, decision threshold, sizing, stopping rule, deviations log

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references, assets) in data-analytics/statistical-analyst of borghei/Claude-Skills.

  • SKILL.md
  • assets/experiment_design_template.md
  • assets/sample_contingency.json
  • assets/sample_experiment.json
  • assets/sample_question.json
  • assets/sample_revenue.json
  • assets/sample_session_times.json
  • references/effect-sizes-and-communication.md
  • references/test-selection-and-assumptions.md
  • scripts/run_test.py
  • scripts/stats_core.py
  • scripts/test_impl.py
  • scripts/test_selector.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Statistical Analyst next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Statistical Analyst compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Statistical Analyst this skillborghei/Claude-Skills881—~3.5kAutomated safety check: PassMIT
MatlabzLanqing/codex-claude-academic-skills4.6k9 repos~2.3kAutomated safety check: NotesGPL-3.0
Eqtl Catalogue Region FetchClawBio/ClawBio1.2k1 repos~4.3kAutomated safety check: PassMIT
CSV Data Analysis5zjk5/prompt-engineering127—~2.6kAutomated safety check: PassNone
Meridian MMM Model Buildinggoogle/meridian1.6k—~2.5kAutomated safety check: PassApache-2.0
Gwas Catalog Region FetchClawBio/ClawBio1.2k1 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Matlab

    zLanqing/codex-claude-academic-skills

    MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing.

    4.6k GitHub starsUsed in 9 repos~2.3k tokens
    Data & AnalyticsAuto-check: notes
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Data & AnalyticsAuto-check passed
  • CSV Data Analysis

    5zjk5/prompt-engineering

    This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.

    127 GitHub stars~2.6k tokensUpdated 23 days ago
    Data & AnalyticsAuto-check passed
  • Official

    Takes a user through building a Meridian marketing mix model, from loading CSV data and mapping columns to running EDA, fitting and saving the model.

    1.6k GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    384 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Statistical Analyst

What does Statistical Analyst do?

Applied statistics for business and product questions — test selection, assumption checks, power planning, effect sizes with intervals, multiplicity correction. Statistical Analyst is an agent skill from borghei/Claude-Skills. Applied statistics for business and product questions — test selection, assumption checks, power planning, effect sizes with intervals, multiplicity correction.

When should I use Statistical Analyst?

Statistical Analyst fits situations like: interpreting an experiment; vetting a claim.

How do I install Statistical Analyst in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill statistical-analyst -a claude-code`. Or copy the skill folder (data-analytics/statistical-analyst in borghei/Claude-Skills) into .claude/skills/statistical-analyst in your project. Claude Code loads it when a task matches its description.

How do I install Statistical Analyst in Codex?

Run `npx skills add borghei/Claude-Skills --skill statistical-analyst -a codex`. Or copy the skill folder (data-analytics/statistical-analyst in borghei/Claude-Skills) into .agents/skills/statistical-analyst in your project. Codex loads it when a task matches its description.

Can I use Statistical Analyst in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill statistical-analyst -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/statistical-analyst, .gemini/skills/statistical-analyst, .github/skills/statistical-analyst and .opencode/skills/statistical-analyst in your project.

What does Statistical Analyst need to run?

Going by SKILL.md and its folder, Statistical Analyst needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Statistical Analyst access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Statistical Analyst safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Statistical Analyst use?

Statistical Analyst is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Statistical Analyst use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to Statistical Analyst?

Skills that share tags, products or a category with Statistical Analyst: Matlab (zLanqing/codex-claude-academic-skills, 4.6k stars), Eqtl Catalogue Region Fetch (ClawBio/ClawBio, 1.2k stars), CSV Data Analysis (5zjk5/prompt-engineering, 127 stars) and Meridian MMM Model Building (google/meridian, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Statistical Analyst?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.