Agent skill

Statistical Analysis

by w95 in w95/awesome-claude-corporate-skills

Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing.

MITAuto-check passedData & Analytics

Install Statistical Analysis

skills CLI
$ npx skills add w95/awesome-claude-corporate-skills --skill statistical-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install w95/awesome-claude-corporate-skills statistical-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/w95/awesome-claude-corporate-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/10-data-analytics/statistical-analysis .claude/skills/statistical-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
statistical-analysis
GitHub stars
244
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
1,218 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing.

  • Works in 4 steps: Plot the raw time series -- visual… → Compute day-of-week averages: is there a… → Compute month-of-year averages: is there… → …
  • Analyzing distributions
  • SKILL.md covers Descriptive Statistics…, Trend Analysis and Forecasting, Outlier and Anomaly Detection and Hypothesis Testing Basics, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Statistical Analysis is an agent skill from w95/awesome-claude-corporate-skills. Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing. Use when analyzing distributions, testing for significance, detecting anomalies, computing correlations, or interpreting statistical results.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Forecasting and time series, Statistics and Data cleaning. The repository describes itself as: 166 production-ready Claude AI skills organized by corporate role — executive leadership, finance, HR, marketing, sales, legal, operations, engineering, product, data, customer…. The licence is MIT.

When your agent uses it

  • Analyzing distributions
  • Testing for significance
  • Detecting anomalies
  • Computing correlations

Example prompts

  • “/statistical-analysis”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Plot the raw time series -- visual inspection first
  2. Compute day-of-week averages: is there a clear weekly pattern?
  3. Compute month-of-year averages: is there an annual cycle?
  4. When comparing periods, always use YoY or same-period comparisons to avoid conflating trend with seasonality

What it can do on your machine

Read from SKILL.md and the folder at commit 78dbc7c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Statistical Analysis loads about 2.6k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 1,218 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from w95/awesome-claude-corporate-skills at commit 78dbc7c, republished under its MIT licence (© w95). 1,218 words, ~2,603 tokens.

Download SKILL.mdSave it as .claude/skills/statistical-analysis/SKILL.md (or your agent's skills folder).
name
statistical-analysis
description
Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing. Use when analyzing distributions, testing for significance, detecting anomalies, computing correlations, or interpreting statistical results.

Statistical Analysis Skill

Descriptive statistics, trend analysis, outlier detection, hypothesis testing, and guidance on when to be cautious about statistical claims.

Descriptive Statistics Methodology

Central Tendency

Choose the right measure of center based on the data:

SituationUseWhy
Symmetric distribution, no outliersMeanMost efficient estimator
Skewed distributionMedianRobust to outliers
Categorical or ordinal dataModeOnly option for non-numeric
Highly skewed with outliers (e.g., revenue per user)Median + meanReport both; the gap shows skew

Always report mean and median together for business metrics. If they diverge significantly, the data is skewed and the mean alone is misleading.

Spread and Variability
  • Standard deviation: How far values typically fall from the mean. Use with normally distributed data.
  • Interquartile range (IQR): Distance from p25 to p75. Robust to outliers. Use with skewed data.
  • Coefficient of variation (CV): StdDev / Mean. Use to compare variability across metrics with different scales.
  • Range: Max minus min. Sensitive to outliers but gives a quick sense of data extent.
Percentiles for Business Context

Report key percentiles to tell a richer story than mean alone:

p1:   Bottom 1% (floor / minimum typical value)
p5:   Low end of normal range
p25:  First quartile
p50:  Median (typical user)
p75:  Third quartile
p90:  Top 10% / power users
p95:  High end of normal range
p99:  Top 1% / extreme users

Example narrative: "The median session duration is 4.2 minutes, but the top 10% of users spend over 22 minutes per session, pulling the mean up to 7.8 minutes."

Describing Distributions

Characterize every numeric distribution you analyze:

  • Shape: Normal, right-skewed, left-skewed, bimodal, uniform, heavy-tailed
  • Center: Mean and median (and the gap between them)
  • Spread: Standard deviation or IQR
  • Outliers: How many and how extreme
  • Bounds: Is there a natural floor (zero) or ceiling (100%)?

Trend Analysis and Forecasting

Moving averages to smooth noise:

python
# 7-day moving average (good for daily data with weekly seasonality)
df['ma_7d'] = df['metric'].rolling(window=7, min_periods=1).mean()

# 28-day moving average (smooths weekly AND monthly patterns)
df['ma_28d'] = df['metric'].rolling(window=28, min_periods=1).mean()

Period-over-period comparison:

  • Week-over-week (WoW): Compare to same day last week
  • Month-over-month (MoM): Compare to same month prior
  • Year-over-year (YoY): Gold standard for seasonal businesses
  • Same-day-last-year: Compare specific calendar day

Growth rates:

Simple growth: (current - previous) / previous
CAGR: (ending / beginning) ^ (1 / years) - 1
Log growth: ln(current / previous)  -- better for volatile series
Seasonality Detection

Check for periodic patterns:

  1. Plot the raw time series -- visual inspection first
  2. Compute day-of-week averages: is there a clear weekly pattern?
  3. Compute month-of-year averages: is there an annual cycle?
  4. When comparing periods, always use YoY or same-period comparisons to avoid conflating trend with seasonality
Forecasting (Simple Methods)

For business analysts (not data scientists), use straightforward methods:

  • Naive forecast: Tomorrow = today. Use as a baseline.
  • Seasonal naive: Tomorrow = same day last week/year.
  • Linear trend: Fit a line to historical data. Only for clearly linear trends.
  • Moving average forecast: Use trailing average as the forecast.

Always communicate uncertainty. Provide a range, not a point estimate:

  • "We expect 10K-12K signups next month based on the 3-month trend"
  • NOT "We will get exactly 11,234 signups next month"

When to escalate to a data scientist: Non-linear trends, multiple seasonalities, external factors (marketing spend, holidays), or when forecast accuracy matters for resource allocation.

Outlier and Anomaly Detection

Statistical Methods

Z-score method (for normally distributed data):

python
z_scores = (df['value'] - df['value'].mean()) / df['value'].std()
outliers = df[abs(z_scores) > 3]  # More than 3 standard deviations

IQR method (robust to non-normal distributions):

python
Q1 = df['value'].quantile(0.25)
Q3 = df['value'].quantile(0.75)
IQR = Q3 - Q1
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
outliers = df[(df['value'] < lower_bound) | (df['value'] > upper_bound)]

Percentile method (simplest):

python
outliers = df[(df['value'] < df['value'].quantile(0.01)) |
              (df['value'] > df['value'].quantile(0.99))]
Handling Outliers

Do NOT automatically remove outliers. Instead:

  1. Investigate: Is this a data error, a genuine extreme value, or a different population?
  2. Data errors: Fix or remove (e.g., negative ages, timestamps in year 1970)
  3. Genuine extremes: Keep them but consider using robust statistics (median instead of mean)
  4. Different population: Segment them out for separate analysis (e.g., enterprise vs. SMB customers)

Report what you did: "We excluded 47 records (0.3%) with transaction amounts >$50K, which represent bulk enterprise orders analyzed separately."

Time Series Anomaly Detection

For detecting unusual values in a time series:

  1. Compute expected value (moving average or same-period-last-year)
  2. Compute deviation from expected
  3. Flag deviations beyond a threshold (typically 2-3 standard deviations of the residuals)
  4. Distinguish between point anomalies (single unusual value) and change points (sustained shift)

Hypothesis Testing Basics

When to Use

Use hypothesis testing when you need to determine whether an observed difference is likely real or could be due to random chance. Common scenarios:

  • A/B test results: Is variant B actually better than A?
  • Before/after comparison: Did the product change actually move the metric?
  • Segment comparison: Do enterprise customers really have higher retention?
The Framework
  1. Null hypothesis (H0): There is no difference (the default assumption)
  2. Alternative hypothesis (H1): There is a difference
  3. Choose significance level (alpha): Typically 0.05 (5% chance of false positive)
  4. Compute test statistic and p-value
  5. Interpret: If p < alpha, reject H0 (evidence of a real difference)
Show full SKILL.md (505 more words)Show less
Common Tests
ScenarioTestWhen to Use
Compare two group meanst-test (independent)Normal data, two groups
Compare two group proportionsz-test for proportionsConversion rates, binary outcomes
Compare paired measurementsPaired t-testBefore/after on same entities
Compare 3+ group meansANOVAMultiple segments or variants
Non-normal data, two groupsMann-Whitney U testSkewed metrics, ordinal data
Association between categoriesChi-squared testTwo categorical variables
Practical Significance vs. Statistical Significance

Statistical significance means the difference is unlikely due to chance.

Practical significance means the difference is large enough to matter for business decisions.

A difference can be statistically significant but practically meaningless (common with large samples). Always report:

  • Effect size: How big is the difference? (e.g., "Variant B improved conversion by 0.3 percentage points")
  • Confidence interval: What's the range of plausible true effects?
  • Business impact: What does this translate to in revenue, users, or other business terms?
Sample Size Considerations
  • Small samples produce unreliable results, even with significant p-values
  • Rule of thumb for proportions: Need at least 30 events per group for basic reliability
  • For detecting small effects (e.g., 1% conversion rate change), you may need thousands of observations per group
  • If your sample is small, say so: "With only 200 observations per group, we have limited power to detect effects smaller than X%"

When to Be Cautious About Statistical Claims

Correlation Is Not Causation

When you find a correlation, explicitly consider:

  • Reverse causation: Maybe B causes A, not A causes B
  • Confounding variables: Maybe C causes both A and B
  • Coincidence: With enough variables, spurious correlations are inevitable

What you can say: "Users who use feature X have 30% higher retention" What you cannot say without more evidence: "Feature X causes 30% higher retention"

Multiple Comparisons Problem

When you test many hypotheses, some will be "significant" by chance:

  • Testing 20 metrics at p=0.05 means ~1 will be falsely significant
  • If you looked at many segments before finding one that's different, note that
  • Adjust for multiple comparisons with Bonferroni correction (divide alpha by number of tests) or report how many tests were run
Simpson's Paradox

A trend in aggregated data can reverse when data is segmented:

  • Always check whether the conclusion holds across key segments
  • Example: Overall conversion goes up, but conversion goes down in every segment -- because the mix shifted toward a higher-converting segment
Survivorship Bias

You can only analyze entities that "survived" to be in your dataset:

  • Analyzing active users ignores those who churned
  • Analyzing successful companies ignores those that failed
  • Always ask: "Who is missing from this dataset, and would their inclusion change the conclusion?"
Ecological Fallacy

Aggregate trends may not apply to individuals:

  • "Countries with higher X have higher Y" does NOT mean "individuals with higher X have higher Y"
  • Be careful about applying group-level findings to individual cases
Anchoring on Specific Numbers

Be wary of false precision:

  • "Churn will be 4.73% next quarter" implies more certainty than is warranted
  • Prefer ranges: "We expect churn between 4-6% based on historical patterns"
  • Round appropriately: "About 5%" is often more honest than "4.73%"

© w95, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in 10-data-analytics/statistical-analysis of w95/awesome-claude-corporate-skills.

Open the folder on GitHubat commit 78dbc7c

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in w95/awesome-claude-corporate-skills, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Statistical Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Statistical Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Statistical Analysis this skillw95/awesome-claude-corporate-skills2441 repos~2.6kAutomated safety check: PassMIT
Automl SkillLeoYeAI/openclaw-master-skills2.2k—~3.6kAutomated safety check: PassMIT
Cja Dimension Analysisadobe/skills197—~3.3kAutomated safety check: PassApache-2.0
Stat Edaasgard-ai-platform/skills242—~954Automated safety check: PassMIT
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Automl Skill

    LeoYeAI/openclaw-master-skills

    AutoML 自动化机器学习技能 | Automated Machine Learning Skill. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~3.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Comprehensive dimension analysis and reporting for CJA. An agent skill from adobe/skills.

    197 GitHub stars~3.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Stat Eda

    asgard-ai-platform/skills

    Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks.

    242 GitHub stars~954 tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 10 days ago
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Find Hypertable Candidates

    timescale/pg-aiguide

    A skill your agent uses to analyze an existing PostgreSQL database and identify which tables should be converted to Timescale/TimescaleDB hypertables.

    1.9k GitHub starsUsed in 1 repo~2.6k tokens
    Data & AnalyticsAuto-check passed

More from w95/awesome-claude-corporate-skills

All 42 skills in this repo
  • Data Context Extractor

    w95/awesome-claude-corporate-skills

    Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts.

    244 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Competitive Analysis

    w95/awesome-claude-corporate-skills

    Framework for competitive landscape analysis across any industry.

    244 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed
  • Account Research

    w95/awesome-claude-corporate-skills

    Research a company using Common Room data. An agent skill from w95/awesome-claude-corporate-skills.

    244 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Call Prep

    w95/awesome-claude-corporate-skills

    Prepare for a customer or prospect call using Common Room signals.

    244 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Compose Outreach

    w95/awesome-claude-corporate-skills

    Generate personalized outreach messages using Common Room signals.

    244 GitHub stars~1.4k tokensUpdated 7 mo ago
    Auto-check passed
  • SQL Queries

    w95/awesome-claude-corporate-skills

    Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.).

    244 GitHub starsUsed in 3 repos~2.8k tokens
    Auto-check passed

Questions about Statistical Analysis

What does Statistical Analysis do?

Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing. Statistical Analysis is an agent skill from w95/awesome-claude-corporate-skills. Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing.

When should I use Statistical Analysis?

Statistical Analysis fits situations like: analyzing distributions; testing for significance; detecting anomalies; computing correlations.

How do I install Statistical Analysis in Claude Code?

Run `npx skills add w95/awesome-claude-corporate-skills --skill statistical-analysis -a claude-code`. Or copy the skill folder (10-data-analytics/statistical-analysis in w95/awesome-claude-corporate-skills) into .claude/skills/statistical-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Statistical Analysis in Codex?

Run `npx skills add w95/awesome-claude-corporate-skills --skill statistical-analysis -a codex`. Or copy the skill folder (10-data-analytics/statistical-analysis in w95/awesome-claude-corporate-skills) into .agents/skills/statistical-analysis in your project. Codex loads it when a task matches its description.

Can I use Statistical Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add w95/awesome-claude-corporate-skills --skill statistical-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/statistical-analysis, .gemini/skills/statistical-analysis, .github/skills/statistical-analysis and .opencode/skills/statistical-analysis in your project.

What does Statistical Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Statistical Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Statistical Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Statistical Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Statistical Analysis use?

Statistical Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Statistical Analysis use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Statistical Analysis?

Skills that share tags, products or a category with Statistical Analysis: Automl Skill (LeoYeAI/openclaw-master-skills, 2.2k stars), Cja Dimension Analysis (adobe/skills, 197 stars), Stat Eda (asgard-ai-platform/skills, 242 stars) and TimesFM Forecasting (google-research/timesfm, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Statistical Analysis?

w95 (a GitHub user) maintains it in w95/awesome-claude-corporate-skills, which has 244 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on February 26, 2026.

Source: w95/awesome-claude-corporate-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.