Agent skill

Data Analysis

by seb1n in seb1n/awesome-ai-agent-skills

Analyze datasets to answer defined questions through statistical methods, trend identification, hypothesis testing, and correlation analysis.

MITAuto-check passedData & Analytics

Install Data Analysis

skills CLI
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seb1n/awesome-ai-agent-skills data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-and-analytics/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
206
Token cost
~1.7k tokens
SKILL.md length
562 words
Files
1
Skills in repo
101
Repo updated
First seen
Licence
MIT

At a glance

Analyze datasets to answer defined questions through statistical methods, trend identification, hypothesis testing, and correlation analysis.

  • Works in 6 steps: Load and profile the data. Read the… → Compute descriptive statistics. Generate… → Identify trends and patterns. Apply… → …
  • The user needs evidence-backed findings
  • SKILL.md covers Workflow, Supported Technologies, Usage and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Analysis is an agent skill from seb1n/awesome-ai-agent-skills. Analyze datasets to answer defined questions through statistical methods, trend identification, hypothesis testing, and correlation analysis. Use when the user needs evidence-backed findings or decisions from data; use exploratory-data-analysis instead for open-ended first-pass profiling before questions are defined.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data analysis and Statistics. It works with pandas. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.

When your agent uses it

  • The user needs evidence-backed findings
  • Decisions from data
  • Use exploratory-data-analysis instead for open-ended first-pass profiling before questions are defined

Example prompts

  • “/data-analysis”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Load and profile the data. Read the dataset into a pandas DataFrame and inspect its shape, column types, and memory usage. Display the…
  2. Compute descriptive statistics. Generate summary statistics for all numeric columns including mean, median, standard deviation, skewness…
  3. Identify trends and patterns. Apply rolling averages, percentage changes, and seasonal decomposition to time-indexed data. For…
  4. Perform correlation and hypothesis testing. Calculate Pearson and Spearman correlation matrices to quantify relationships between…
  5. Detect anomalies and outliers. Use the IQR method and z-scores to identify data points that deviate significantly from the norm…
  6. Synthesize findings into a report. Summarize the key insights in plain language, supported by specific numbers. Rank findings by business…

What it can do on your machine

Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analysis loads about 1.7k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 562 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 562 words, ~1,700 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder).
name
data-analysis
description
Analyze datasets to answer defined questions through statistical methods, trend identification, hypothesis testing, and correlation analysis. Use when the user needs evidence-backed findings or decisions from data; use exploratory-data-analysis instead for open-ended first-pass profiling before questions are defined.
license
MIT
metadata.author
awesome-ai-agent-skills
metadata.version
1.0.0

Data Analysis

This skill enables an AI agent to perform rigorous statistical analysis on structured datasets. The agent loads data, computes descriptive and inferential statistics, identifies trends and correlations, tests hypotheses, and produces actionable insights. It supports CSV, Excel, Parquet, and JSON inputs and leverages pandas, scipy, and statsmodels for analysis.

Workflow

  1. Load and profile the data. Read the dataset into a pandas DataFrame and inspect its shape, column types, and memory usage. Display the first and last rows to confirm the data loaded correctly. Check for obvious structural issues such as shifted columns or encoding problems.

  2. Compute descriptive statistics. Generate summary statistics for all numeric columns including mean, median, standard deviation, skewness, and kurtosis. For categorical columns, compute value counts and mode. This step establishes a baseline understanding of each variable's distribution.

  3. Identify trends and patterns. Apply rolling averages, percentage changes, and seasonal decomposition to time-indexed data. For non-temporal data, use group-by aggregations and pivot tables to surface patterns across categories. Flag any monotonic trends or cyclical behavior.

  4. Perform correlation and hypothesis testing. Calculate Pearson and Spearman correlation matrices to quantify relationships between variables. Conduct hypothesis tests (t-tests, chi-square, ANOVA) where appropriate to determine statistical significance. Report p-values and confidence intervals alongside effect sizes.

  5. Detect anomalies and outliers. Use the IQR method and z-scores to identify data points that deviate significantly from the norm. Cross-reference outliers with domain context to determine whether they represent errors, rare events, or meaningful signals.

  6. Synthesize findings into a report. Summarize the key insights in plain language, supported by specific numbers. Rank findings by business impact or statistical significance. Include limitations and caveats such as sample size constraints or confounding variables.

Supported Technologies

  • pandas — data loading, manipulation, and aggregation
  • scipy.stats — hypothesis testing, statistical distributions
  • statsmodels — time-series decomposition, regression analysis
  • numpy — numerical computations

Usage

Provide the agent with a file path to the dataset and a description of the analysis goals. Optionally specify which columns to focus on, the significance level for hypothesis tests (default alpha=0.05), and whether time-series methods should be applied.

Show full SKILL.md (218 more words)Show less

Examples

Example 1: Sales CSV analysis with pandas
python
import pandas as pd
from scipy import stats

# Load the dataset
df = pd.read_csv("sales_2024.csv", parse_dates=["order_date"])

# Descriptive statistics
print(df[["revenue", "units_sold", "discount"]].describe())
#          revenue  units_sold  discount
# count   8450.00     8450.00   8450.00
# mean     312.45       4.12      0.08
# std      189.73       2.87      0.05
# min       12.00       1.00      0.00
# max     2450.00      47.00      0.35

# Correlation analysis
corr = df[["revenue", "units_sold", "discount"]].corr(method="pearson")
print(corr)
#             revenue  units_sold  discount
# revenue       1.000       0.847    -0.213
# units_sold    0.847       1.000    -0.089
# discount     -0.213      -0.089     1.000

# Hypothesis test: do discounted orders produce higher revenue?
discounted = df[df["discount"] > 0]["revenue"]
full_price = df[df["discount"] == 0]["revenue"]
t_stat, p_value = stats.ttest_ind(discounted, full_price)
print(f"t={t_stat:.3f}, p={p_value:.4f}")
# t=-3.217, p=0.0013 — discounted orders have significantly lower revenue per order
Example 2: Time-series analysis with seasonal decomposition
python
import pandas as pd
from statsmodels.tsa.seasonal import seasonal_decompose

# Load monthly revenue data
df = pd.read_csv("monthly_revenue.csv", parse_dates=["month"], index_col="month")

# Decompose into trend, seasonal, and residual components
result = seasonal_decompose(df["revenue"], model="additive", period=12)

print("Trend (last 6 months):")
print(result.trend.dropna().tail(6))
# 2024-07    48230.12
# 2024-08    49012.45
# 2024-09    49780.33
# 2024-10    50234.10
# 2024-11    51002.88
# 2024-12    51890.67

print("\nSeasonal peaks:")
seasonal = result.seasonal.groupby(result.seasonal.index.month).mean()
print(seasonal.nlargest(3))
# month
# 11    8923.40   (November — holiday pre-orders)
# 12    7654.20   (December — holiday sales)
# 3     3210.15   (March — spring promotions)

# The upward trend of ~$600/month suggests 14.5% annualized growth.
# Strong Q4 seasonality accounts for roughly 18% of total annual revenue.

Best Practices

  • Always inspect raw data before computing statistics — silent parsing errors (wrong delimiters, encoding issues) can invalidate every downstream result.
  • Report effect sizes alongside p-values; statistical significance alone does not imply practical importance.
  • Use non-parametric tests (Mann-Whitney, Kruskal-Wallis) when data distributions are heavily skewed or sample sizes are small.
  • Segment analysis by meaningful categories (region, product line, customer tier) to avoid Simpson's paradox.
  • Document assumptions explicitly — stationarity for time-series, normality for parametric tests, independence of observations.
  • Validate surprising findings with a holdout sample or alternative methodology before presenting them as conclusions.

Edge Cases

  • Missing values in key columns. If more than 30% of a target column is null, warn the user that imputation may introduce significant bias. Offer to analyze the complete-case subset instead.
  • Extremely skewed distributions. Log-transform or use median-based statistics when skewness exceeds |2.0| to avoid misleading mean values.
  • Multicollinearity. When two predictors correlate above 0.9, flag this and recommend dropping one or using regularized models to avoid inflated coefficients.
  • Small sample sizes (n < 30). Switch to exact tests or bootstrap methods and widen confidence intervals accordingly.
  • Mixed data types in a single column. Coerce carefully and report how many values could not be converted, rather than silently dropping them.

© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in data-and-analytics/data-analysis of seb1n/awesome-ai-agent-skills.

Open the folder on GitHubat commit 75865a5

Compare with similar skills

Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analysis this skillseb1n/awesome-ai-agent-skills206—~1.7kAutomated safety check: PassMIT
Statistical Data Analysislingzhi227/agent-research-skills390—~886Automated safety check: PassNone
Q-EDA Exploratory AnalysisTyrealQ/q-skills108—~1.1kAutomated safety check: PassMIT
Data Explorerliangdabiao/claude-data-analysis-ultra-main290—~2.1kAutomated safety check: PassNone
Data AnalystRightNow-AI/openfang18k—~730Automated safety check: PassApache-2.0
Data AnalysisEXboys/skilllite170—~176Automated safety check: PassMIT

Similar skills

  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    390 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 17 days ago
    Data & AnalyticsAuto-check passed
  • Data Explorer

    liangdabiao/claude-data-analysis-ultra-main

    Performs exploratory data analysis, statistical analysis, and pattern discovery.

    290 GitHub stars~2.1k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Data Analyst

    RightNow-AI/openfang

    Data analysis expert for statistics, visualization, pandas, and exploration

    18k GitHub stars~730 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Data Analysis

    EXboys/skilllite

    Analyze CSV/JSON data with statistics, filtering, and aggregation.

    170 GitHub stars~176 tokensUpdated 15 days ago
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed

More from seb1n/awesome-ai-agent-skills

All 101 skills in this repo
  • Agent Red Teaming

    seb1n/awesome-ai-agent-skills

    Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.

    206 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Eu AI Act Readiness

    seb1n/awesome-ai-agent-skills

    Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…

    206 GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Human In The Loop

    seb1n/awesome-ai-agent-skills

    Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • MCP Server Building

    seb1n/awesome-ai-agent-skills

    Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • PDF Processing

    seb1n/awesome-ai-agent-skills

    Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Skill Supply Chain Audit

    seb1n/awesome-ai-agent-skills

    Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.

    206 GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Data Analysis

What does Data Analysis do?

Analyze datasets to answer defined questions through statistical methods, trend identification, hypothesis testing, and correlation analysis. Data Analysis is an agent skill from seb1n/awesome-ai-agent-skills. Analyze datasets to answer defined questions through statistical methods, trend identification, hypothesis testing, and correlation analysis.

When should I use Data Analysis?

Data Analysis fits situations like: the user needs evidence-backed findings; decisions from data; use exploratory-data-analysis instead for open-ended first-pass profiling before questions are defined.

How do I install Data Analysis in Claude Code?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill data-analysis -a claude-code`. Or copy the skill folder (data-and-analytics/data-analysis in seb1n/awesome-ai-agent-skills) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Data Analysis in Codex?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill data-analysis -a codex`. Or copy the skill folder (data-and-analytics/data-analysis in seb1n/awesome-ai-agent-skills) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Data Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Data Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Analysis use?

Data Analysis is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analysis use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Analysis?

Skills that share tags, products or a category with Data Analysis: Statistical Data Analysis (lingzhi227/agent-research-skills, 390 stars), Q-EDA Exploratory Analysis (TyrealQ/q-skills, 108 stars), Data Explorer (liangdabiao/claude-data-analysis-ultra-main, 290 stars) and Data Analyst (RightNow-AI/openfang, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analysis?

seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 101 skills in this directory. The repository was last updated on August 9, 2026.

Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.