Agent skill

Data Analyst

by borghei in borghei/Claude-Skills

Data analysis across SQL, visualization, statistics, and reporting.

MITAuto-check passedData & Analytics

Install Data Analyst

skills CLI
$ npx skills add borghei/Claude-Skills --skill data-analyst -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills data-analyst --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-analytics/data-analyst .claude/skills/data-analyst && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analyst
GitHub stars
874
Token cost
~3.1k tokens
SKILL.md length
1,053 words
Files
4 (incl. scripts)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

Data analysis across SQL, visualization, statistics, and reporting.

  • Works in 6 steps: Frame the business question -- Restate… → Write and validate SQL -- Use CTEs for… → Explore and profile data -- Compute… → …
  • Writing SQL queries
  • SKILL.md covers Clarify First, Workflow, SQL Patterns and Chart Selection Matrix, plus 11 more sections
  • Runs Python scripts from its folder; calls python

What it does

Data Analyst is an agent skill from borghei/Claude-Skills. Data analysis across SQL, visualization, statistics, and reporting. Use when writing SQL queries, building dashboards, performing cohort or funnel analysis, running hypothesis tests, or presenting data-driven recommendations.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/data_profiler.py`, `scripts/query_optimizer.py` and `scripts/report_generator.py`).

It sits in Data & Analytics, covering Data analysis, Statistics and SQL. It works with SQL. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Writing SQL queries
  • Building dashboards
  • Performing cohort
  • Funnel analysis

Example prompts

  • “/data-analyst”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Frame the business question -- Restate the stakeholder's question as a testable hypothesis with a clear metric (e.g., "Did campaign X…
  2. Write and validate SQL -- Use CTEs for readability. Filter early, aggregate late. Run EXPLAIN ANALYZE on complex queries to verify index…
  3. Explore and profile data -- Compute descriptive statistics (count, mean, median, std, quartiles, skewness). Check for nulls, duplicates…
  4. Analyze -- Apply the appropriate method: cohort analysis for retention, funnel analysis for conversion, hypothesis testing (t-test…
  5. Visualize -- Select chart type from the matrix below. Follow the design rules (Y-axis at zero for bars, <=7 colors, labels on axes…
  6. Deliver the insight -- Structure findings as What / So What / Now What. Lead with the headline, support with a chart, close with a…

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analyst loads about 3.1k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 1,053 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 1,053 words, ~3,085 tokens.

Download SKILL.mdSave it as .claude/skills/data-analyst/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
data-analyst
description
Data analysis across SQL, visualization, statistics, and reporting. Use when writing SQL queries, building dashboards, performing cohort or funnel analysis, running hypothesis tests, or presenting data-driven recommendations.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
data-analytics
metadata.updated
2026-03-31
metadata.tags
analytics, sql, visualization, statistics, reporting

Data Analyst

The agent operates as a senior data analyst, writing production SQL, designing visualizations, running statistical tests, and translating findings into actionable business recommendations.

Clarify First

Before the analysis, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Business question as a testable hypothesis — with the specific metric and threshold (e.g., ">= 5% lift in 7-day retention") (frames the whole analysis and the headline)
  • Data sources and grain — which tables/columns exist and the row grain (determines feasibility and the SQL you can write)
  • Audience and the decision — who consumes the insight and what they will decide (sets altitude and the "Now What" recommendation)
  • Analysis type — cohort, funnel, hypothesis test, or trend (selects the method and the chart)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Workflow

  1. Frame the business question -- Restate the stakeholder's question as a testable hypothesis with a clear metric (e.g., "Did campaign X increase 7-day retention by >= 5%?"). Identify required data sources.
  2. Write and validate SQL -- Use CTEs for readability. Filter early, aggregate late. Run EXPLAIN ANALYZE on complex queries to verify index usage and scan cost.
  3. Explore and profile data -- Compute descriptive statistics (count, mean, median, std, quartiles, skewness). Check for nulls, duplicates, and outliers before drawing conclusions.
  4. Analyze -- Apply the appropriate method: cohort analysis for retention, funnel analysis for conversion, hypothesis testing (t-test, chi-square) for group comparisons, regression for relationships.
  5. Visualize -- Select chart type from the matrix below. Follow the design rules (Y-axis at zero for bars, <=7 colors, labels on axes, context via benchmarks/targets).
  6. Deliver the insight -- Structure findings as What / So What / Now What. Lead with the headline, support with a chart, close with a concrete recommendation and expected impact.

SQL Patterns

Monthly aggregation with growth:

sql
WITH monthly AS (
    SELECT
        date_trunc('month', created_at) AS month,
        COUNT(*)                        AS total_orders,
        COUNT(DISTINCT customer_id)     AS unique_customers,
        SUM(amount)                     AS revenue
    FROM orders
    WHERE created_at >= '2024-01-01'
    GROUP BY 1
),
growth AS (
    SELECT month, revenue,
        LAG(revenue) OVER (ORDER BY month) AS prev_revenue
    FROM monthly
)
SELECT month, revenue,
    ROUND((revenue - prev_revenue) / prev_revenue * 100, 1) AS growth_pct
FROM growth
ORDER BY month;

Cohort retention:

sql
WITH first_orders AS (
    SELECT customer_id,
        date_trunc('month', MIN(created_at)) AS cohort_month
    FROM orders GROUP BY 1
),
cohort_data AS (
    SELECT f.cohort_month,
        date_trunc('month', o.created_at) AS order_month,
        COUNT(DISTINCT o.customer_id)     AS customers
    FROM orders o
    JOIN first_orders f ON o.customer_id = f.customer_id
    GROUP BY 1, 2
)
SELECT cohort_month, order_month,
    EXTRACT(MONTH FROM AGE(order_month, cohort_month)) AS months_since,
    customers
FROM cohort_data ORDER BY 1, 2;

Window functions (running total + previous order):

sql
SELECT customer_id, order_date, amount,
    SUM(amount) OVER (PARTITION BY customer_id ORDER BY order_date) AS running_total,
    LAG(amount) OVER (PARTITION BY customer_id ORDER BY order_date) AS prev_amount
FROM orders;

Chart Selection Matrix

Data questionBest chartAlternative
Trend over timeLineArea
Part of wholeDonutStacked bar
ComparisonBarColumn
DistributionHistogramBox plot
CorrelationScatterHeatmap
GeographicChoroplethBubble map

Design rules: Start Y-axis at zero for bar charts. Use <= 7 colors. Label axes. Include benchmarks or targets for context. Avoid 3D charts and pie charts with > 5 slices.

Dashboard Layout

+------------------------------------------------------------+
| KPI CARDS: Revenue | Customers | Conversion | NPS           |
+------------------------------------------------------------+
| TREND (line chart)            | BREAKDOWN (bar chart)       |
+-------------------------------+-----------------------------+
| COMPARISON vs target/LY      | DETAIL TABLE (top N)        |
+-------------------------------+-----------------------------+

Statistical Methods

Hypothesis testing (t-test):

python
from scipy import stats
import numpy as np

def compare_groups(a: np.ndarray, b: np.ndarray, alpha: float = 0.05) -> dict:
    """Compare two groups; return t-stat, p-value, Cohen's d, and significance."""
    stat, p = stats.ttest_ind(a, b)
    d = (a.mean() - b.mean()) / np.sqrt((a.std()**2 + b.std()**2) / 2)
    return {"t_statistic": stat, "p_value": p, "cohens_d": d, "significant": p < alpha}

Chi-square test for independence:

python
def test_independence(table, alpha=0.05):
    chi2, p, dof, _ = stats.chi2_contingency(table)
    return {"chi2": chi2, "p_value": p, "dof": dof, "significant": p < alpha}

Key Business Metrics

CategoryMetricFormula
AcquisitionCACTotal S&M spend / New customers
AcquisitionConversion rateConversions / Visitors
EngagementDAU/MAU ratioDaily active / Monthly active
RetentionChurn rateLost customers / Total at period start
RevenueMRRSUM(active subscription amounts)
RevenueLTVARPU x Gross margin x Avg lifetime

Insight Delivery Template

markdown
## [Headline: action-oriented finding]

**What:** One-sentence description of the observation.
**So What:** Why this matters to the business (revenue, retention, cost).
**Now What:** Recommended action with expected impact.
**Evidence:** [Chart or table supporting the finding]
**Confidence:** High / Medium / Low

Analysis Framework

markdown
# Analysis: [Topic]
## Business Question -- What are we trying to answer?
## Hypothesis -- What do we expect to find?
## Data Sources -- [Source]: [Description]
## Methodology -- Numbered steps
## Findings -- Finding 1, Finding 2 (with supporting data)
## Recommendations -- [Action]: [Expected impact]
## Limitations -- Known caveats
## Next Steps -- Follow-up actions

Scripts

bash
python scripts/query_optimizer.py --file query.sql
python scripts/query_optimizer.py --sql "SELECT * FROM orders" --json
python scripts/data_profiler.py --file sales.csv
python scripts/data_profiler.py --file data.json --top 10 --json
python scripts/report_generator.py --file sales.csv --title "Monthly Sales Report"
python scripts/report_generator.py --file data.csv --group-by region --format markdown --json

Tool Reference

ToolPurposeKey Flags
query_optimizer.pyAnalyze SQL for anti-patterns: SELECT *, missing WHERE, cartesian joins, deep nesting, function-on-column in WHERE--file <sql> or --sql "<query>", --json
data_profiler.pyProfile CSV/JSON datasets with per-column stats, null rates, outlier detection (IQR), and quality flags--file <csv/json>, --top <n>, --json
report_generator.pyGenerate summary reports with numeric aggregations, group-by breakdowns, and highlights--file <csv/json>, --title, --group-by <col>, --format text/markdown, --json
Show full SKILL.md (538 more words)Show less

Troubleshooting

ProblemLikely CauseResolution
SQL query runs for minutes on a table with indexesQuery uses functions on indexed columns in WHERE clause (e.g., WHERE UPPER(name) = ...)Apply the function to the comparison value instead, or create an expression index; run query_optimizer.py to detect this pattern
data_profiler.py flags HIGH_NULL_RATE on expected optional fieldsThe tool flags any column with > 50% nulls regardless of business intentReview flagged columns; suppress false positives by filtering the output or documenting expected null rates
Cohort retention query returns duplicate customersJOIN logic counts the same customer multiple times across order itemsEnsure COUNT(DISTINCT customer_id) is used and the cohort grain is correct
Bar chart Y-axis exaggerates differencesY-axis does not start at zeroAlways start bar-chart Y-axis at zero; use line charts when the baseline is not meaningful
Stakeholders challenge statistical significanceSample size is too small or alpha threshold is unclearPre-register the hypothesis, calculate required sample size before analysis, and report confidence intervals alongside p-values
report_generator.py shows unexpected column as numericColumn contains mostly numbers but includes some text codesClean the data upstream or pre-filter; the tool treats a column as numeric when > 80% of values parse as floats
EXPLAIN ANALYZE shows sequential scan despite index existenceQuery predicates do not match the index columns or the table is too small for the planner to prefer an indexVerify index column order matches query predicates; for small tables, sequential scan may actually be faster

Success Criteria

  • Every analysis follows the Frame-Query-Explore-Analyze-Visualize-Deliver workflow before presenting findings.
  • SQL queries pass query_optimizer.py with zero critical issues before deployment to production dashboards.
  • Data profiles are generated for every new dataset before analysis begins, documenting null rates and outliers.
  • Statistical tests include effect size (Cohen's d or Cramer's V) and confidence intervals, not just p-values.
  • Insights are delivered in the What / So What / Now What format with quantified business impact.
  • Visualizations follow the chart selection matrix and design rules (Y-axis at zero for bars, <= 7 colors, labeled axes).
  • Reports generated by report_generator.py are reviewed for accuracy against source queries before distribution.

Scope & Limitations

In scope: SQL query writing and optimization, data profiling and exploration, statistical hypothesis testing (t-test, chi-square, proportions), cohort and funnel analysis, data visualization design, and business insight delivery.

Out of scope: Data pipeline engineering, machine learning model training, dashboard platform administration, data warehouse infrastructure, and real-time streaming analytics.

Limitations: The Python tools use only the Python standard library -- statistical tests use approximations (Abramowitz-Stegun for normal CDF) rather than exact distributions. For production-grade statistics, use scipy or statsmodels. query_optimizer.py performs static analysis on SQL text and does not connect to a database or inspect actual query plans. data_profiler.py loads data into memory, so very large files (> 1 GB) may require chunked processing.

Integration Points

  • Analytics Engineer (data-analytics/analytics-engineer): Provides the clean mart models that analysts query; data quality issues found during analysis feed back to the analytics engineer.
  • Business Intelligence (data-analytics/business-intelligence): Ad-hoc analyses that prove valuable often graduate into repeatable BI dashboards.
  • Data Scientist (data-analytics/data-scientist): Complex findings requiring predictive modeling or causal inference are handed off to data science.
  • Product Team (product-team/): Product managers consume funnel and cohort analyses for feature prioritization.
  • Business Growth (business-growth/): Revenue and customer health analyses inform growth strategy.

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in data-analytics/data-analyst of borghei/Claude-Skills.

  • SKILL.md
  • scripts/data_profiler.py
  • scripts/query_optimizer.py
  • scripts/report_generator.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

Data Analyst next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analyst compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analyst this skillborghei/Claude-Skills874—~3.1kAutomated safety check: PassMIT
Data Analysisspytensor/openmozi439—~535Automated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Find Hypertable Candidatestimescale/pg-aiguide1.9k1 repos~2.6kAutomated safety check: PassApache-2.0
Dinobase Business Data Querieskappa90/dinobase263—~1.5kAutomated safety check: PassCustom licence
Code Generatorliangdabiao/claude-data-analysis-ultra-main290—~513Automated safety check: PassNone

Similar skills

  • Data Analysis

    spytensor/openmozi

    Data analysis workflow: ingest, validate quality, explore, analyze, report.

    439 GitHub stars~535 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Find Hypertable Candidates

    timescale/pg-aiguide

    A skill your agent uses to analyze an existing PostgreSQL database and identify which tables should be converted to Timescale/TimescaleDB hypertables.

    1.9k GitHub starsUsed in 1 repo~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.

    263 GitHub stars~1.5k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Code Generator

    liangdabiao/claude-data-analysis-ultra-main

    Generates production-ready analysis code in Python, R, SQL. An agent skill from liangdabiao/claude-data-analysis-ultra-main.

    290 GitHub stars~513 tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Almost Paid

    nestyme/awesome-prompts

    Find the users who almost paid — saw the paywall, started checkout, ran out of free credits, let a trial lapse — size what they are worth, and turn them into a one-screen dashboard with a…

    149 GitHub stars~3.9k tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Questions about Data Analyst

What does Data Analyst do?

Data analysis across SQL, visualization, statistics, and reporting. Data Analyst is an agent skill from borghei/Claude-Skills. Data analysis across SQL, visualization, statistics, and reporting.

When should I use Data Analyst?

Data Analyst fits situations like: writing SQL queries; building dashboards; performing cohort; funnel analysis.

How do I install Data Analyst in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill data-analyst -a claude-code`. Or copy the skill folder (data-analytics/data-analyst in borghei/Claude-Skills) into .claude/skills/data-analyst in your project. Claude Code loads it when a task matches its description.

How do I install Data Analyst in Codex?

Run `npx skills add borghei/Claude-Skills --skill data-analyst -a codex`. Or copy the skill folder (data-analytics/data-analyst in borghei/Claude-Skills) into .agents/skills/data-analyst in your project. Codex loads it when a task matches its description.

Can I use Data Analyst in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill data-analyst -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analyst, .gemini/skills/data-analyst, .github/skills/data-analyst and .opencode/skills/data-analyst in your project.

What does Data Analyst need to run?

Going by SKILL.md and its folder, Data Analyst needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Data Analyst access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Analyst safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Analyst use?

Data Analyst is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analyst use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Analyst?

Skills that share tags, products or a category with Data Analyst: Data Analysis (spytensor/openmozi, 439 stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Find Hypertable Candidates (timescale/pg-aiguide, 1.9k stars) and Dinobase Business Data Queries (kappa90/dinobase, 263 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analyst?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.