Agent skill

Distribution Profiler

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Single-column distribution deep-dive. An agent skill from ai-analyst-lab/ai-analyst.

MITAuto-check passedDevelopment

Install Distribution Profiler

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill distribution-profiler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst distribution-profiler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/distribution-profiler .claude/skills/distribution-profiler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
distribution-profiler
GitHub stars
304
Token cost
~2.4k tokens
SKILL.md length
853 words
Files
2
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Single-column distribution deep-dive. An agent skill from ai-analyst-lab/ai-analyst.

  • Works in 7 steps: Identify the Target → Extract the Data → Run the Diagnostic Pipeline → …
  • Profile this column
  • SKILL.md covers Purpose, When to Use, Invocation and Instructions, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Distribution Profiler is an agent skill from ai-analyst-lab/ai-analyst. Single-column distribution deep-dive. Profile the statistical distribution of a data column and produce an analytical playbook: distribution identification, valid summary stats, recommended tests, A/B guidance, traps. Trigger on "profile this column", "what distribution is this", "check the distribution", "is this normal", "what test should I use", "check assumptions before A/B test", "how is this data distributed".

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in Development, covering Performance optimization and A/B testing. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • Profile this column
  • What distribution is this
  • Check the distribution
  • What test should I use

Example prompts

  • “profile this column”
  • “what distribution is this”
  • “check the distribution”
  • “/distribution-profiler”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Identify the Target
  2. Extract the Data
  3. Run the Diagnostic Pipeline
  4. Identify the Distribution
  5. Generate the Report
  6. Generate a Diagnostic Chart
  7. Offer Next Steps

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Distribution Profiler loads about 2.4k tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 853 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 853 words, ~2,372 tokens.

Download SKILL.mdSave it as .claude/skills/distribution-profiler/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
distribution-profiler
description
Single-column distribution deep-dive. Profile the statistical distribution of a data column and produce an analytical playbook: distribution identification, valid summary stats, recommended tests, A/B guidance, traps. Trigger on "profile this column", "what distribution is this", "check the distribution", "is this normal", "what test should I use", "check assumptions before A/B test", "how is this data distributed".

Skill: Distribution Profiler

Purpose

Take any numeric data column and produce a complete analytical playbook: identify the distribution, compute the right summary statistics, recommend the correct statistical tests, flag common traps, and give specific A/B testing guidance.

This skill exists because the #1 mistake in product analytics is assuming data is normal when it's not — leading to wrong tests, false positives, and misleading dashboards. The profiler catches this automatically.

When to Use

  • Before any analysis involving a numeric metric
  • When the user asks "what distribution is this?" or "what test should I use?"
  • When checking assumptions for an A/B test
  • When a user says "profile" or "understand" a metric
  • Proactively when you notice an analysis is about to use a t-test or OLS on data that hasn't been checked

Invocation

/distribution-profiler — profile a data column's distribution

Instructions

Step 0: Identify the Target

Figure out what column/metric the user wants profiled. This could be:

  • A specific column name (e.g., "total_amount from orders")
  • A derived metric (e.g., "revenue per user", "sessions per user per month")
  • A SQL query result

If unclear, ask. If the user hasn't specified, look at what they're analyzing and suggest the most relevant metric to profile.

Step 1: Extract the Data

Write and execute a Python script to extract the target column from the active dataset. Use helpers/data/data_helpers.py to resolve the data source:

python
from helpers.data.data_helpers import detect_active_source, check_connection
source = detect_active_source()

Query through ConnectionManager (helpers/data/connection_manager.py) whatever the source; it resolves local vs. remote and auto-logs the query for provenance.

For per-user metrics (revenue per user, sessions per user), aggregate first — the unit of analysis matters. Profile the metric at the level it will be used in the analysis (per-user, per-session, per-day, etc.).

Step 2: Run the Diagnostic Pipeline

Write and execute a single Python script that computes all diagnostics. The script should print results as structured output that you'll use to build the report.

Core diagnostics to compute:

python
import numpy as np
from scipy import stats

# --- Data basics ---
n = len(x)
n_zeros = (x == 0).sum()
pct_zeros = n_zeros / n
n_unique = len(np.unique(x))
is_integer = np.all(x == np.floor(x))
x_min, x_max = x.min(), x.max()

# --- Central tendency ---
mean = np.mean(x)
median = np.median(x)
mean_median_ratio = mean / median if median != 0 else float('inf')

# --- Spread ---
sd = np.std(x, ddof=1)
cv = sd / mean if mean != 0 else float('inf')
iqr = stats.iqr(x)
mad = stats.median_abs_deviation(x)

# --- Shape ---
skewness = stats.skew(x)
excess_kurtosis = stats.kurtosis(x, fisher=True)

# --- Distribution-specific diagnostics ---
vmr = np.var(x, ddof=1) / mean if mean != 0 else float('inf')  # for counts

# --- Formal tests ---
# Normality (on raw data)
if n <= 5000:
    shapiro_stat, shapiro_p = stats.shapiro(x)
anderson_result = stats.anderson(x, dist='norm')

# Normality of log-transformed data (if all positive)
if x_min > 0:
    log_x = np.log(x)
    shapiro_log_stat, shapiro_log_p = stats.shapiro(log_x)

# Bimodality
bimodality_coeff = (skewness**2 + 1) / (excess_kurtosis + 3)

Key decision points the diagnostics must answer:

  1. Is the data discrete (counts) or continuous? → determines count vs continuous branch
  2. Is there zero-inflation? → compare observed zero % to expected under fitted distribution
  3. Is it bounded (0,1)? → beta distribution
  4. For counts: VMR near 1, >> 1, or < 1? → Poisson vs NB vs Binomial
  5. For continuous positive: is log(x) normal? → log-normal check
  6. Is it symmetric? → normal check
  7. Is it bimodal? → dip test + bimodality coefficient
  8. Is the tail extremely heavy? → power law check (mean/median ratio, top 1% share)
Step 3: Identify the Distribution

Use the diagnostic results and the decision flowchart from the reference guide at .knowledge/references/statistical-distributions-guide.md (Section 15: Distribution Identification Flowchart) to identify the most likely distribution.

Read the relevant section of the reference guide for the identified distribution to get the full details on summary stats, tests, transformations, and A/B implications.

Report your confidence level:

  • High confidence: Formal tests agree, visual shape matches, domain makes sense
  • Moderate confidence: Some tests agree, shape is consistent but not perfect
  • Low confidence: Tests disagree or data doesn't fit cleanly into one family. In this case, report the top 2-3 candidates and what would distinguish them.
Show full SKILL.md (345 more words)Show less
Step 4: Generate the Report

Produce a clear, actionable report with these sections. Be specific and concrete — use the actual numbers from the diagnostics, not generic advice.

Report Template
## Distribution Profile: [metric name]

### 1. Distribution: [Name] (confidence: high/moderate/low)

[One sentence: what distribution this is and why it makes sense for this metric.]
[If moderate/low confidence, list alternatives.]

### 2. Key Diagnostics

| Diagnostic | Value | What It Tells Us |
|-----------|-------|------------------|
| n | ... | Sample size |
| Mean | ... | ... |
| Median | ... | ... |
| Mean/Median ratio | ... | [>1.3 = right-skewed; near 1 = symmetric] |
| SD | ... | ... |
| MAD | ... | [Robust alternative to SD] |
| CV (SD/Mean) | ... | [~1 = exponential; >1.5 = heavy tail] |
| Skewness | ... | ... |
| Excess kurtosis | ... | ... |
| % zeros | ... | [if relevant] |
| VMR (Var/Mean) | ... | [for counts: ~1 Poisson, >>1 NB] |

**Shapiro-Wilk (normality):** p = ... → [reject/fail to reject]
**Shapiro-Wilk (log-normality):** p = ... → [reject/fail to reject] (if applicable)

### 3. Valid Summary Statistics

Use these to describe this metric:
- **Central tendency:** [median / geometric mean / mean — with reasoning]
- **Spread:** [IQR / MAD / CV — with reasoning]
- **Avoid:** [what NOT to report and why]

### 4. Recommended Statistical Tests

| Purpose | Recommended Test | Why |
|---------|-----------------|-----|
| Compare two groups | ... | ... |
| Regression | ... | ... |
| Confidence intervals | ... | ... |

### 5. A/B Testing Playbook

- **Recommended approach:** [specific method]
- **Sample size impact:** [how this distribution affects required n]
- **Effect size measure:** [what to use instead of / in addition to Cohen's d]
- **Common traps:** [specific warnings for this distribution]
- **Practical tip:** [one concrete thing to do]

### 6. If You Need to Transform

[Recommended transformation and when to use it vs. GLM vs. non-parametric]
Step 5: Generate a Diagnostic Chart

Create a matplotlib visualization with 4 panels:

  1. Histogram with KDE — shows the shape of the distribution
  2. Q-Q plot (against normal) — shows departure from normality
  3. Box plot — shows median, IQR, and outliers
  4. Log-scale histogram (if data is positive and skewed) — shows whether log-transform normalizes the data

Save as a PNG file. Use plt.savefig() and plt.close() to avoid display issues.

Step 6: Offer Next Steps

After presenting the report, suggest what the user might want to do next:

  • "Want me to run the recommended A/B test on this metric?"
  • "Should I profile another column for comparison?"
  • "Want to see how this distribution changes across user segments?"

Reference

The comprehensive distribution reference guide is at: .knowledge/references/statistical-distributions-guide.md

This 1,700+ line guide covers 12 distributions plus zero-inflated models, with detection heuristics, valid summary statistics, recommended tests, transformations, A/B testing implications, and Python code for each.

When to read it: After identifying the distribution in Step 3, read the specific section for that distribution to get detailed guidance. Don't load the entire guide — read only the relevant section (each is ~100 lines).

Key sections:

  • Sections 1-12: Individual distribution cards (A through G for each)
  • Section 13: Zero-inflated distributions
  • Section 14: Practical guidance (t-test rules, family tree, effect sizes, censoring)
  • Section 15: Quick-reference lookup tables
  • Appendix: Quick diagnostic ratios

Rules

  1. Always show your work — print the diagnostic numbers, don't just state conclusions
  2. Always generate the 4-panel chart — visual evidence is as important as the numbers
  3. Never assume normality without testing — that's the whole point of this skill
  4. If the distribution is unclear, say so — report multiple candidates with reasoning
  5. Use the reference guide for detailed recommendations, not your own memory
  6. Profile at the right level of aggregation — per-user, per-session, etc.
  7. Warn loudly if the user is about to use a test that's wrong for the distribution

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/distribution-profiler of ai-analyst-lab/ai-analyst.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Distribution Profiler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Distribution Profiler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Distribution Profiler this skillai-analyst-lab/ai-analyst304—~2.4kAutomated safety check: PassMIT
Chrome Performance Optimizernwjs/chromium.src160—~4.2kAutomated safety check: PassBSD-3-Clause
Cache Trace Analyzerben-manes/caffeine18k—~702Automated safety check: NotesApache-2.0
Platform Norm Profileraaron-he-zhu/aaron-marketing-skills2.9k—~3.1kAutomated safety check: PassApache-2.0
Memory Optimizationbenchflow-ai/skillsbench1.8k—~1.5kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Autonomous multi-agent performance optimization loop for Chromium and V8.

    160 GitHub stars~4.2k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed
  • Cache Trace Analyzer

    ben-manes/caffeine

    Analyzes a cache trace file for the Caffeine simulator, characterizes its access pattern and recommends which eviction policies to compare.

    18k GitHub stars~702 tokensUpdated today
    DevelopmentAuto-check: notes
  • Platform Norm Profiler

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "build the norm card for this platform", "what are the char limits and visible-fold cutoffs here", "is the LinkedIn link-in-first-comment thing…

    2.9k GitHub stars~3.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Memory Optimization

    benchflow-ai/skillsbench

    Optimize Python code for reduced memory usage and improved memory efficiency.

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 8 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 8 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 8 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed

Questions about Distribution Profiler

What does Distribution Profiler do?

Single-column distribution deep-dive. An agent skill from ai-analyst-lab/ai-analyst. Distribution Profiler is an agent skill from ai-analyst-lab/ai-analyst. Single-column distribution deep-dive.

When should I use Distribution Profiler?

Distribution Profiler fits situations like: profile this column; what distribution is this; check the distribution; what test should I use.

How do I install Distribution Profiler in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill distribution-profiler -a claude-code`. Or copy the skill folder (.claude/skills/distribution-profiler in ai-analyst-lab/ai-analyst) into .claude/skills/distribution-profiler in your project. Claude Code loads it when a task matches its description.

How do I install Distribution Profiler in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill distribution-profiler -a codex`. Or copy the skill folder (.claude/skills/distribution-profiler in ai-analyst-lab/ai-analyst) into .agents/skills/distribution-profiler in your project. Codex loads it when a task matches its description.

Can I use Distribution Profiler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill distribution-profiler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/distribution-profiler, .gemini/skills/distribution-profiler, .github/skills/distribution-profiler and .opencode/skills/distribution-profiler in your project.

What does Distribution Profiler need to run?

SKILL.md names no scripts, command-line tools or credentials: Distribution Profiler is instructions for the agent only. Our summary lists: Python 3.

Does Distribution Profiler access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Distribution Profiler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Distribution Profiler use?

Distribution Profiler is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Distribution Profiler use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Distribution Profiler?

Skills that share tags, products or a category with Distribution Profiler: Chrome Performance Optimizer (nwjs/chromium.src, 160 stars), Cache Trace Analyzer (ben-manes/caffeine, 18k stars), Platform Norm Profiler (aaron-he-zhu/aaron-marketing-skills, 2.9k stars) and Memory Optimization (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Distribution Profiler?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.