Agent skill

Data Analyst

by ailabs-393 in ailabs-393/ai-labs-claude-skills

This skill should be used when analyzing CSV datasets, handling missing values through intelligent imputation, and creating interactive dashboards to visualize data trends.

MITAuto-check passedData & Analytics

Install Data Analyst

skills CLI
$ npx skills add ailabs-393/ai-labs-claude-skills --skill data-analyst -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ailabs-393/ai-labs-claude-skills data-analyst --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ailabs-393/ai-labs-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/skills/data-analyst .claude/skills/data-analyst && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analyst
GitHub stars
454
Token cost
~2.7k tokens
SKILL.md length
1,204 words
Files
8 (incl. scripts, references)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when analyzing CSV datasets, handling missing values through intelligent imputation, and creating interactive dashboards to visualize data trends.

  • Works in 6 steps: Missing Value Analysis → Intelligent Imputation → Interactive Dashboard Creation → …
  • Tasks involving data quality assessment
  • SKILL.md covers Overview, Core Capabilities, Complete Workflow and Individual Use Cases, plus 5 more sections
  • Runs Python and JavaScript scripts from its folder; calls python3 and pip

What it does

Data Analyst is an agent skill from ailabs-393/ai-labs-claude-skills. This skill should be used when analyzing CSV datasets, handling missing values through intelligent imputation, and creating interactive dashboards to visualize data trends. Use this skill for tasks involving data quality assessment, automated missing value detection and filling, statistical analysis, and generating Plotly Dash dashboards for exploratory data analysis.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `index.js`, `package.json` and `references/imputation_methods.md`).

It sits in Data & Analytics, covering Data cleaning and Data analysis. It works with Plotly. The repository describes itself as: This package is use to remove the hustle of finding claudeskills and shift them into any of the user project. This project become a bridge between user's usage and claude skills. The licence is MIT.

When your agent uses it

  • Tasks involving data quality assessment
  • Automated missing value detection and filling
  • Statistical analysis
  • Generating Plotly Dash dashboards for exploratory data analysis

Example prompts

  • “/data-analyst”

Requirements

  • Python 3
  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Missing Value Analysis
  2. Intelligent Imputation
  3. Interactive Dashboard Creation
  4. Analyze Missing Values
  5. Impute Missing Values
  6. Create Interactive Dashboard

What it can do on your machine

Read from SKILL.md and the folder at commit 1a12bc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python and JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analyst loads about 2.7k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 96 tokens; SKILL.md has 1,204 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ailabs-393/ai-labs-claude-skills at commit 1a12bc7, republished under its MIT licence (© ailabs-393). 1,204 words, ~2,716 tokens.

Download SKILL.mdSave it as .claude/skills/data-analyst/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
data-analyst
description
This skill should be used when analyzing CSV datasets, handling missing values through intelligent imputation, and creating interactive dashboards to visualize data trends. Use this skill for tasks involving data quality assessment, automated missing value detection and filling, statistical analysis, and generating Plotly Dash dashboards for exploratory data analysis.

Data Analyst

Overview

This skill provides comprehensive capabilities for data analysis workflows on CSV datasets. It automatically analyzes missing value patterns, intelligently imputes missing data using appropriate statistical methods, and creates interactive Plotly Dash dashboards for visualizing trends and patterns. The skill combines automated missing value handling with rich interactive visualizations to support end-to-end exploratory data analysis.

Core Capabilities

The data-analyst skill provides three main capabilities that can be used independently or as a complete workflow:

1. Missing Value Analysis

Automatically detect and analyze missing values in datasets, identifying patterns and suggesting optimal imputation strategies.

2. Intelligent Imputation

Apply sophisticated imputation methods tailored to each column's data type and distribution characteristics.

3. Interactive Dashboard Creation

Generate comprehensive Plotly Dash dashboards with multiple visualization types for trend analysis and exploration.

Complete Workflow

When a user requests complete data analysis with missing value handling and visualization, follow this workflow:

Step 1: Analyze Missing Values

Run the missing value analysis script to understand the data quality:

bash
python3 scripts/analyze_missing_values.py <input_file.csv> <output_analysis.json>

What this does:

  • Detects missing values in each column
  • Identifies data types (numeric, categorical, temporal, etc.)
  • Calculates missing value statistics
  • Suggests appropriate imputation strategies per column
  • Generates detailed JSON report and console output

Review the output to understand:

  • Which columns have missing data
  • The percentage of missing values
  • The recommended imputation method for each column
  • Why each method was recommended
Step 2: Impute Missing Values

Apply automatic imputation based on the analysis:

bash
python3 scripts/impute_missing_values.py <input_file.csv> <analysis.json> <output_imputed.csv>

What this does:

  • Loads the analysis results (or performs analysis if not provided)
  • Applies the optimal imputation method to each column:
    • Mean: For normally distributed numeric data
    • Median: For skewed numeric data
    • Mode: For categorical variables
    • KNN: For multivariate numeric data with correlations
    • Forward fill: For time series data
    • Constant: For high-cardinality text fields
  • Handles edge cases (drops rows/columns when appropriate)
  • Generates imputation report with before/after statistics
  • Saves cleaned dataset

The script automatically:

  • Drops columns with >70% missing values
  • Drops rows where critical ID columns are missing
  • Performs batch KNN imputation for correlated variables
  • Creates detailed imputation log
Step 3: Create Interactive Dashboard

Generate an interactive Plotly Dash dashboard:

bash
python3 scripts/create_dashboard.py <imputed_file.csv> <output_dir> <port>

Example:

bash
python3 scripts/create_dashboard.py data_imputed.csv ./visualizations 8050

What this does:

  • Automatically detects column types (numeric, categorical, temporal)
  • Creates comprehensive visualizations:
    • Summary statistics table: Descriptive stats for all numeric columns
    • Time series plots: Trend analysis if date/time columns exist
    • Distribution plots: Histograms for understanding data distributions
    • Correlation heatmap: Relationships between numeric variables
    • Categorical analysis: Bar charts for categorical variables
    • Scatter plot matrix: Pairwise relationships between variables
  • Launches interactive Dash web server
  • Optionally saves static HTML visualizations

Access the dashboard at http://127.0.0.1:8050 (or specified port)

Individual Use Cases

Use Case A: Quick Missing Value Assessment

When the user wants to understand data quality without imputation:

bash
python3 scripts/analyze_missing_values.py data.csv

Review the console output to understand missing value patterns and get recommendations.

Use Case B: Imputation Only

When the user has a dataset with missing values and wants cleaned data:

bash
python3 scripts/impute_missing_values.py data.csv

This performs analysis and imputation in one step, producing data_imputed.csv.

Use Case C: Visualization Only

When the user has a clean dataset and wants interactive visualizations:

bash
python3 scripts/create_dashboard.py clean_data.csv ./visualizations 8050

This creates a full dashboard without any preprocessing.

Use Case D: Custom Imputation Strategy

When the user wants to review and adjust imputation strategies:

  1. Run analysis first:

    bash
    python3 scripts/analyze_missing_values.py data.csv analysis.json
  2. Review analysis.json and discuss strategies with the user

  3. If needed, modify the imputation logic or parameters in the script

  4. Run imputation:

    bash
    python3 scripts/impute_missing_values.py data.csv analysis.json data_imputed.csv

Understanding Imputation Methods

The skill uses intelligent imputation strategies based on data characteristics. Key methods include:

  • Mean/Median: For numeric data (mean for normal distributions, median for skewed)
  • Mode: For categorical variables (most frequent value)
  • KNN (K-Nearest Neighbors): For multivariate numeric data where variables are correlated
  • Forward Fill: For time series data (carry last observation forward)
  • Interpolation: For smooth temporal trends
  • Constant Value: For high-cardinality text fields (e.g., "Unknown")
  • Drop: For columns with >70% missing or rows with missing IDs

For detailed information about when each method is appropriate, refer to references/imputation_methods.md.

Dashboard Features

The interactive dashboard includes:

Summary Statistics
  • Count, mean, std, min, max, quartiles for all numeric columns
  • Missing value counts and percentages
  • Sortable table format
Time Series Analysis
  • Line plots with markers for temporal trends
  • Multiple series support (up to 4 primary metrics)
  • Hover details with exact values
  • Unified hover mode for easy comparison
Distribution Analysis
  • Histograms for all numeric variables
  • 30-bin default for granular distribution view
  • Multi-panel layout for easy comparison
Show full SKILL.md (491 more words)Show less
Correlation Analysis
  • Heatmap showing correlation coefficients
  • Color-coded from -1 (negative) to +1 (positive)
  • Annotated with exact correlation values
  • Useful for identifying relationships
Categorical Analysis
  • Bar charts for categorical variables
  • Top 10 categories shown (for high-cardinality variables)
  • Frequency counts displayed
Scatter Plot Matrix
  • Pairwise scatter plots for numeric variables
  • Limited to 5 variables for readability
  • Lower triangle shown (avoiding redundancy)

Setup and Dependencies

Before using the skill, ensure dependencies are installed:

bash
pip install -r requirements.txt

Required packages:

  • pandas - Data manipulation and analysis
  • numpy - Numerical computing
  • scikit-learn - KNN imputation
  • plotly - Interactive visualizations
  • dash - Web dashboard framework
  • dash-bootstrap-components - Dashboard styling

Best Practices

For Analysis:
  1. Always run analysis before imputation to understand data quality
  2. Review suggested imputation methods - they're recommendations, not mandates
  3. Pay attention to missing value percentages (>40% requires careful consideration)
  4. Check data types match expectations (e.g., numeric IDs detected as numeric)
For Imputation:
  1. Save the original dataset before imputation
  2. Review the imputation report to ensure methods make sense
  3. Check imputed values are within reasonable ranges
  4. Consider creating missing indicators for important variables
  5. Document which imputation methods were used for reproducibility
For Dashboards:
  1. Use imputed/cleaned data for most accurate visualizations
  2. Save static HTML plots if sharing with non-technical stakeholders
  3. Use different ports if running multiple dashboards simultaneously
  4. For large datasets (>100k rows), consider sampling for faster rendering

Handling Edge Cases

High Missing Rates (>50%)

The scripts automatically flag columns with >50% missing values. Options:

  • Drop the column if not critical
  • Create a missing indicator variable
  • Investigate why data is missing (may be informative)
Mixed Data Types

If a column contains mixed types (e.g., numbers and text):

  • The script detects the primary type
  • Consider cleaning the column before analysis
  • Use constant imputation for mixed-type text columns
Small Datasets

For datasets with <50 rows:

  • Simple imputation (mean/median/mode) is more stable
  • Avoid KNN (requires sufficient neighbors)
  • Consider dropping rows instead of imputing
Time Series Gaps

For time series with irregular timestamps:

  • Use forward fill for short gaps
  • Use interpolation for longer gaps with smooth trends
  • Consider the sampling frequency when choosing methods

Troubleshooting

Script fails with "module not found"

Install dependencies: pip install -r requirements.txt

Dashboard won't start (port in use)

Specify a different port: python3 scripts/create_dashboard.py data.csv ./viz 8051

KNN imputation is slow

KNN is computationally intensive for large datasets. For >50k rows, consider:

  • Using simpler methods (mean/median)
  • Sampling the data first
  • Using fewer columns in KNN
Imputed values seem incorrect
  • Review the analysis report - check detected data types
  • Verify the column is being detected correctly (numeric vs categorical)
  • Consider manual adjustment or different imputation method
  • Check for outliers that may affect mean/median calculations

Resources

scripts/
  • analyze_missing_values.py - Comprehensive missing value analysis with automatic strategy recommendation
  • impute_missing_values.py - Intelligent imputation using multiple methods tailored to data characteristics
  • create_dashboard.py - Interactive Plotly Dash dashboard generator with multiple visualization types
references/
  • imputation_methods.md - Detailed guide to missing value imputation strategies, decision frameworks, and best practices
Other Files
  • requirements.txt - Python dependencies for the skill

© ailabs-393, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in packages/skills/data-analyst of ailabs-393/ai-labs-claude-skills.

  • SKILL.md
  • index.js
  • package.json
  • references/imputation_methods.md
  • requirements.txt
  • scripts/analyze_missing_values.py
  • scripts/create_dashboard.py
  • scripts/impute_missing_values.py

Open the folder on GitHubat commit 1a12bc7

Compare with similar skills

Data Analyst next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analyst compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analyst this skillailabs-393/ai-labs-claude-skills454—~2.7kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Code EngineeropenJiuwen-ai/sciencediscovery148—~2.8kAutomated safety check: PassApache-2.0
Data Analysisxiaoyuge886/aigc1981 repos~794Automated safety check: PassMIT
Math Modeling Data Cleaning and Chartsyushui2022/MathModel-Skill4521 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Code Engineer

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

    148 GitHub stars~2.8k tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed
  • Data Analysis

    xiaoyuge886/aigc

    Perform data analysis tasks including data cleaning, statistical analysis, visualization, and insight generation.

    198 GitHub starsUsed in 1 repo~794 tokens
    Data & AnalyticsAuto-check passed
  • Math Modeling Data Cleaning and Charts

    yushui2022/MathModel-Skill

    Cleans raw or scraped competition data and produces exploratory charts and a figure plan as one stage of a mathematical modeling paper workflow.

    452 GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Data Analytics

    XiaomiMiMo/MiMo-Code

    A skill your agent uses for quantitative product or business analysis: data quality checks, metric diagnostics, KPI design and reporting, dashboards, analytical reports, charts, notebooks, market…

    14k GitHub stars~961 tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed

More from ailabs-393/ai-labs-claude-skills

All 22 skills in this repo
  • Tech Debt Analyzer

    ailabs-393/ai-labs-claude-skills

    This skill should be used when analyzing technical debt in a codebase, documenting code quality issues, creating technical debt registers, or assessing code maintainability.

    454 GitHub starsUsed in 2 repos~3.9k tokens
    Auto-check passed
  • SEO Optimizer

    ailabs-393/ai-labs-claude-skills

    This skill should be used when analyzing HTML/CSS websites for SEO optimization, fixing SEO issues, generating SEO reports, or implementing SEO best practices.

    454 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Business Document Generator

    ailabs-393/ai-labs-claude-skills

    This skill should be used when the user requests to create professional business documents (proposals, business plans, or budgets) from templates.

    454 GitHub stars~2k tokensUpdated 11 mo ago
    Auto-check passed
  • Finance Manager

    ailabs-393/ai-labs-claude-skills

    Comprehensive personal finance management system for analyzing transaction data, generating insights, creating visualizations, and providing actionable financial recommendations.

    454 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Docker Containerization

    ailabs-393/ai-labs-claude-skills

    This skill should be used when containerizing applications with Docker, creating Dockerfiles, docker-compose configurations, or deploying containers to various platforms.

    454 GitHub stars~2.1k tokensUpdated 11 mo ago
    Auto-check: notes
  • Startup Validator

    ailabs-393/ai-labs-claude-skills

    Comprehensive startup idea validation and market analysis tool.

    454 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Works with

Questions about Data Analyst

What does Data Analyst do?

This skill should be used when analyzing CSV datasets, handling missing values through intelligent imputation, and creating interactive dashboards to visualize data trends. Data Analyst is an agent skill from ailabs-393/ai-labs-claude-skills. This skill should be used when analyzing CSV datasets, handling missing values through intelligent imputation, and creating interactive dashboards to visualize data trends.

When should I use Data Analyst?

Data Analyst fits situations like: tasks involving data quality assessment; automated missing value detection and filling; statistical analysis; generating Plotly Dash dashboards for exploratory data analysis.

How do I install Data Analyst in Claude Code?

Run `npx skills add ailabs-393/ai-labs-claude-skills --skill data-analyst -a claude-code`. Or copy the skill folder (packages/skills/data-analyst in ailabs-393/ai-labs-claude-skills) into .claude/skills/data-analyst in your project. Claude Code loads it when a task matches its description.

How do I install Data Analyst in Codex?

Run `npx skills add ailabs-393/ai-labs-claude-skills --skill data-analyst -a codex`. Or copy the skill folder (packages/skills/data-analyst in ailabs-393/ai-labs-claude-skills) into .agents/skills/data-analyst in your project. Codex loads it when a task matches its description.

Can I use Data Analyst in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ailabs-393/ai-labs-claude-skills --skill data-analyst -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analyst, .gemini/skills/data-analyst, .github/skills/data-analyst and .opencode/skills/data-analyst in your project.

What does Data Analyst need to run?

Going by SKILL.md and its folder, Data Analyst needs Python and JavaScript for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: Python 3; Node.js.

Does Data Analyst access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Data Analyst safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Analyst use?

Data Analyst is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analyst use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Data Analyst?

Skills that share tags, products or a category with Data Analyst: Pandas Pro (Jeffallan/claude-skills, 12k stars), Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars), Code Engineer (openJiuwen-ai/sciencediscovery, 148 stars) and Data Analysis (xiaoyuge886/aigc, 198 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analyst?

ailabs-393 (a GitHub user) maintains it in ailabs-393/ai-labs-claude-skills, which has 454 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on November 11, 2025.

Source: ailabs-393/ai-labs-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.