Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks.

MITAuto-check passedData & Analytics

Install Stat Eda

skills CLI
$ npx skills add asgard-ai-platform/skills --skill stat-eda -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install asgard-ai-platform/skills stat-eda --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/asgard-ai-platform/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/stat-eda .claude/skills/stat-eda && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stat-eda
GitHub stars
242
Token cost
~954 tokens
SKILL.md length
240 words
Files
3 (incl. references)
Skills in repo
207
Repo updated
First seen
Licence
MIT

At a glance

Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks.

  • Works in 4 steps: Split BEFORE explore (see IRON LAW above) → Missing data pattern matters more than… → Simpson's paradox check: If a trend… → …
  • The user has a dataset and needs to understand its structure
  • SKILL.md covers Framework, Output Format, Gotchas and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stat Eda is an agent skill from asgard-ai-platform/skills. Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks. Use this skill when the user has a dataset and needs to understand its structure, find patterns, detect anomalies, or prepare data for further analysis — even if they say 'what does this data look like', 'find interesting patterns', 'clean this data', or 'summarize this dataset'.

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `examples/sample_scenario.md` and `references/missing-data.md`).

It sits in Data & Analytics, covering Data analysis, Data cleaning and Anomaly detection. The repository describes itself as: 301 open-source coding agent skills across 22 domains — methodology, judgment & gotchas packaged as Claude Agent Skills for the Asgard AI Platform. The licence is MIT.

When your agent uses it

  • The user has a dataset and needs to understand its structure
  • Detect anomalies
  • Prepare data for further analysis — even if they say what does this data look like
  • Find interesting patterns

Example prompts

  • “what does this data look like”
  • “find interesting patterns”
  • “clean this data”
  • “/stat-eda”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Split BEFORE explore (see IRON LAW above)
  2. Missing data pattern matters more than count: MCAR is safe to impute; MNAR (e.g. high-income respondents skip income question) requires…
  3. Simpson's paradox check: If a trend holds in the aggregate but reverses within subgroups, the aggregate trend is misleading. Always…
  4. Data leakage in features: A feature that perfectly correlates with the target is usually derived FROM the target (e.g. "refund_amount"…

What it can do on your machine

Read from SKILL.md and the folder at commit 4e7f4f8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stat Eda loads about 954 tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 240 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~954
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from asgard-ai-platform/skills at commit 4e7f4f8, republished under its MIT licence (© asgard-ai-platform). 240 words, ~954 tokens.

Download SKILL.mdSave it as .claude/skills/stat-eda/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
stat-eda
description
Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks. Use this skill when the user has a dataset and needs to understand its structure, find patterns, detect anomalies, or prepare data for further analysis — even if they say 'what does this data look like', 'find interesting patterns', 'clean this data', or 'summarize this dataset'.
metadata.category
WP-21 設計/資訊/傳播/公衛
metadata.tags
data-analysis, eda, statistics, visualization

Exploratory Data Analysis (EDA)

Framework

IRON LAW: Perform EDA Only AFTER Train/Test Split — Or You Leak the Future

Agents know "do EDA first." But they almost always do EDA on the FULL
dataset before splitting. This is information leakage: you've seen the
test set's distributions, outliers, and correlations, and your subsequent
modeling choices (feature scaling, outlier treatment, imputation strategy)
are now informed by data the model shouldn't see. Split first, then EDA
only on the training set. Apply the same transformations to the test set
without re-examining it.

Exception: data quality checks (nulls, dtypes, duplicates) CAN run on
the full dataset since they don't inform model hyperparameters.
EDA Workflow

Standard five-phase flow (structure → quality → univariate → bivariate → findings summary). Assume the agent already knows these steps. Focus on the non-obvious traps below instead.

Critical additions most EDA guides miss:

  1. Split BEFORE explore (see IRON LAW above)
  2. Missing data pattern matters more than count: MCAR is safe to impute; MNAR (e.g. high-income respondents skip income question) requires domain modeling, not mean-fill
  3. Simpson's paradox check: If a trend holds in the aggregate but reverses within subgroups, the aggregate trend is misleading. Always stratify by the most obvious confound before reporting a bivariate finding
  4. Data leakage in features: A feature that perfectly correlates with the target is usually derived FROM the target (e.g. "refund_amount" predicting churn — it's an effect, not a cause). Flag any feature with r > 0.95 for causal review

For the visualization selection guide, see references/missing-data.md.

Output Format

markdown
# EDA Report: {Dataset Name}

## Dataset Overview
- Rows: {N}, Columns: {N}
- Date range: {if applicable}
- Key columns: {description}

## Data Quality
| Issue | Columns Affected | Count/% | Action |
|-------|-----------------|---------|--------|
| Missing values | {cols} | {N / %} | {drop / impute / investigate} |
| Outliers | {cols} | {N} | {cap / remove / keep} |
| Duplicates | — | {N} | {remove} |

## Key Statistics
| Variable | Mean | Median | Std | Min | Max | Distribution |
|----------|------|--------|-----|-----|-----|-------------|
| {var} | ... | ... | ... | ... | ... | {normal/skewed/bimodal} |

## Key Findings
1. {insight with supporting data}
2. {insight}
3. {insight}

## Recommendations
- {next analysis step or data issue to resolve}

Gotchas

  • Correlation ≠ causation: EDA finds associations. Establishing causation requires controlled experiments or causal inference methods.
  • Outliers can be data errors OR real signal: Don't auto-remove. Investigate. A transaction amount of $1M might be a typo or your biggest customer.
  • Missing data has meaning: Data missing from one column may be related to values in another. "Missing income" may mean "unemployed", not random. Check patterns.
  • Visualization lies: Truncated Y-axes, cherry-picked time ranges, and misleading scales can distort insights. Always use appropriate scales and note limitations.

References

  • For missing data handling strategies, see references/missing-data.md

© asgard-ai-platform, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in stat-eda of asgard-ai-platform/skills.

  • SKILL.md
  • examples/sample_scenario.md
  • references/missing-data.md

Open the folder on GitHubat commit 4e7f4f8

Compare with similar skills

Stat Eda next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stat Eda compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stat Eda this skillasgard-ai-platform/skills242—~954Automated safety check: PassMIT
Code EngineeropenJiuwen-ai/sciencediscovery156—~2.8kAutomated safety check: PassApache-2.0
Data Analysisxiaoyuge886/aigc198—~794Automated safety check: PassMIT
Data Explorerliangdabiao/claude-data-analysis-ultra-main290—~2.1kAutomated safety check: PassNone
Profiling Tablesastronomer/agents451—~964Automated safety check: PassApache-2.0
Statistical Analysisw95/awesome-claude-corporate-skills2391 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Code Engineer

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

    156 GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Data Analysis

    xiaoyuge886/aigc

    Perform data analysis tasks including data cleaning, statistical analysis, visualization, and insight generation.

    198 GitHub stars~794 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Data Explorer

    liangdabiao/claude-data-analysis-ultra-main

    Performs exploratory data analysis, statistical analysis, and pattern discovery.

    290 GitHub stars~2.1k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Profiling Tables

    astronomer/agents

    Deep-dive data profiling for a specific table. An agent skill from astronomer/agents.

    451 GitHub stars~964 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Statistical Analysis

    w95/awesome-claude-corporate-skills

    Apply statistical methods including descriptive stats, trend analysis, outlier detection, and hypothesis testing.

    239 GitHub starsUsed in 1 repo~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Data Analysis

    revfactory/harness-100

    A full analysis pipeline where an agent team collaborates to perform exploratory data analysis (EDA), data cleaning, statistical analysis, visualization, and report writing.

    1.3k GitHub stars~1.9k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed

More from asgard-ai-platform/skills

All 207 skills in this repo
  • Algo Ecom Bm25

    asgard-ai-platform/skills

    Implement BM25 ranking function for e-commerce product search relevance scoring.

    242 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Mfg Cpk

    asgard-ai-platform/skills

    Calculate Cpk process capability index to assess whether a process meets specification requirements.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Price Elasticity

    asgard-ai-platform/skills

    Calculate price elasticity of demand to quantify how price changes affect sales volume.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Bayesian

    asgard-ai-platform/skills

    Apply Bayesian averaging to rank items by combining observed ratings with prior expectations.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Elo

    asgard-ai-platform/skills

    Implement Elo rating system to rank items or players from pairwise comparison outcomes.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Algo Rank Wilson

    asgard-ai-platform/skills

    Calculate Wilson Score confidence intervals for ranking items by positive proportion with sample size correction.

    242 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Stat Eda

What does Stat Eda do?

Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks. Stat Eda is an agent skill from asgard-ai-platform/skills. Conduct Exploratory Data Analysis (EDA) using descriptive statistics, visualizations, and data quality checks.

When should I use Stat Eda?

Stat Eda fits situations like: the user has a dataset and needs to understand its structure; detect anomalies; prepare data for further analysis — even if they say what does this data look like; find interesting patterns.

How do I install Stat Eda in Claude Code?

Run `npx skills add asgard-ai-platform/skills --skill stat-eda -a claude-code`. Or copy the skill folder (stat-eda in asgard-ai-platform/skills) into .claude/skills/stat-eda in your project. Claude Code loads it when a task matches its description.

How do I install Stat Eda in Codex?

Run `npx skills add asgard-ai-platform/skills --skill stat-eda -a codex`. Or copy the skill folder (stat-eda in asgard-ai-platform/skills) into .agents/skills/stat-eda in your project. Codex loads it when a task matches its description.

Can I use Stat Eda in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add asgard-ai-platform/skills --skill stat-eda -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stat-eda, .gemini/skills/stat-eda, .github/skills/stat-eda and .opencode/skills/stat-eda in your project.

What does Stat Eda need to run?

SKILL.md names no scripts, command-line tools or credentials: Stat Eda is instructions for the agent only.

Does Stat Eda access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Stat Eda safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stat Eda use?

Stat Eda is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stat Eda use?

About 954 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.

What are the alternatives to Stat Eda?

Skills that share tags, products or a category with Stat Eda: Code Engineer (openJiuwen-ai/sciencediscovery, 156 stars), Data Analysis (xiaoyuge886/aigc, 198 stars), Data Explorer (liangdabiao/claude-data-analysis-ultra-main, 290 stars) and Profiling Tables (astronomer/agents, 451 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stat Eda?

asgard-ai-platform (a GitHub organization) maintains it in asgard-ai-platform/skills, which has 242 GitHub stars. The repository holds 207 skills in this directory. The repository was last updated on June 6, 2026.

Source: asgard-ai-platform/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.