Agent skill

CSV Data Analyzer

by wentorai in wentorai/research-plugins

Load, explore, clean, and analyze CSV data with statistical summaries

MITAuto-check passedDocuments & Office

Install CSV Data Analyzer

skills CLI
$ npx skills add wentorai/research-plugins --skill csv-data-analyzer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins csv-data-analyzer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/wrangling/csv-data-analyzer .claude/skills/csv-data-analyzer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
csv-data-analyzer
GitHub stars
298
Used in
1 other repo
Token cost
~1.7k tokens
SKILL.md length
362 words
Files
1
Skills in repo
428
Repo updated
First seen
Licence
MIT

At a glance

Load, explore, clean, and analyze CSV data with statistical summaries

  • Works in 6 steps: Remove fully empty rows and columns:… → Standardize column names: Convert to… → Handle missing data: Assess missingness… → …
  • Tasks that involve CSV and tabular files
  • SKILL.md covers Overview, Data Loading and Initial…, Data Cleaning Pipeline and Statistical Summary Generation, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

CSV Data Analyzer is an agent skill from wentorai/research-plugins. Load, explore, clean, and analyze CSV data with statistical summaries

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering CSV and tabular files. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve CSV and tabular files

Example prompts

  • “/csv-data-analyzer”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Remove fully empty rows and columns: Drop rows/columns where all values are NaN.
  2. Standardize column names: Convert to snake_case, remove special characters.
  3. Handle missing data: Assess missingness patterns (MCAR/MAR/MNAR) before choosing imputation strategy.
  4. Detect and handle duplicates: Identify exact and near-duplicates using fuzzy matching.
  5. Validate value ranges: Flag values outside expected domain ranges.
  6. Standardize categorical labels: Merge inconsistent spellings (e.g., "Male", "male", "M").

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CSV Data Analyzer loads about 1.7k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 362 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 362 words, ~1,731 tokens.

Download SKILL.mdSave it as .claude/skills/csv-data-analyzer/SKILL.md (or your agent's skills folder).
name
csv-data-analyzer
description
Load, explore, clean, and analyze CSV data with statistical summaries

CSV Data Analyzer

A comprehensive skill for loading, exploring, cleaning, and analyzing CSV datasets within research workflows. Designed for researchers who need to quickly understand the structure, quality, and statistical properties of tabular data before conducting deeper analysis.

Overview

Research datasets commonly arrive as CSV files from instrument exports, survey platforms, government repositories, and collaborator handoffs. This skill provides a structured approach to the entire CSV analysis pipeline: ingestion, profiling, quality assessment, cleaning, transformation, and summary statistics. It emphasizes reproducibility by generating audit logs of every transformation applied to the raw data.

The skill supports datasets of varying complexity, from single-table survey results to multi-file longitudinal study exports with hundreds of columns. It works with standard Python data science libraries (pandas, numpy, scipy) and produces outputs suitable for inclusion in methods sections and supplementary materials.

Data Loading and Initial Profiling

Loading Strategies
python
import pandas as pd
import numpy as np

def load_and_profile_csv(filepath: str, encoding: str = 'utf-8') -> dict:
    """
    Load a CSV file and generate an initial data profile.
    Handles common encoding issues and delimiter detection.
    """
    # Try multiple encodings if default fails
    encodings = [encoding, 'latin-1', 'utf-8-sig', 'cp1252']
    df = None
    for enc in encodings:
        try:
            df = pd.read_csv(filepath, encoding=enc, low_memory=False)
            break
        except (UnicodeDecodeError, pd.errors.ParserError):
            continue

    if df is None:
        raise ValueError(f"Could not parse {filepath} with any supported encoding")

    profile = {
        'rows': len(df),
        'columns': len(df.columns),
        'memory_mb': df.memory_usage(deep=True).sum() / 1e6,
        'dtypes': df.dtypes.value_counts().to_dict(),
        'missing_pct': (df.isnull().sum() / len(df) * 100).to_dict(),
        'duplicates': df.duplicated().sum(),
        'column_names': df.columns.tolist()
    }
    return df, profile
Column Type Inference
python
def infer_semantic_types(df: pd.DataFrame) -> dict:
    """
    Infer semantic column types beyond pandas dtypes.
    Detects dates, identifiers, categorical, continuous, and text columns.
    """
    semantic_types = {}
    for col in df.columns:
        nunique = df[col].nunique()
        ratio = nunique / len(df) if len(df) > 0 else 0

        if ratio > 0.95 and df[col].dtype == 'object':
            semantic_types[col] = 'identifier'
        elif nunique <= 20 and df[col].dtype in ['object', 'int64']:
            semantic_types[col] = 'categorical'
        elif df[col].dtype in ['float64', 'int64']:
            semantic_types[col] = 'continuous'
        elif pd.to_datetime(df[col], errors='coerce').notna().mean() > 0.8:
            semantic_types[col] = 'datetime'
        else:
            semantic_types[col] = 'text'
    return semantic_types

Data Cleaning Pipeline

Systematic Cleaning Steps
  1. Remove fully empty rows and columns: Drop rows/columns where all values are NaN.
  2. Standardize column names: Convert to snake_case, remove special characters.
  3. Handle missing data: Assess missingness patterns (MCAR/MAR/MNAR) before choosing imputation strategy.
  4. Detect and handle duplicates: Identify exact and near-duplicates using fuzzy matching.
  5. Validate value ranges: Flag values outside expected domain ranges.
  6. Standardize categorical labels: Merge inconsistent spellings (e.g., "Male", "male", "M").
python
def clean_column_names(df: pd.DataFrame) -> pd.DataFrame:
    """Standardize column names to snake_case."""
    import re
    df.columns = [
        re.sub(r'[^a-z0-9]+', '_', col.lower().strip()).strip('_')
        for col in df.columns
    ]
    return df

def assess_missingness(df: pd.DataFrame) -> pd.DataFrame:
    """Generate a missingness report for each column."""
    report = pd.DataFrame({
        'missing_count': df.isnull().sum(),
        'missing_pct': (df.isnull().sum() / len(df) * 100).round(2),
        'dtype': df.dtypes
    })
    report['action'] = report['missing_pct'].apply(
        lambda x: 'drop' if x > 60 else ('impute' if x > 0 else 'ok')
    )
    return report.sort_values('missing_pct', ascending=False)
Show full SKILL.md (142 more words)Show less

Statistical Summary Generation

Descriptive Statistics
python
def generate_statistical_summary(df: pd.DataFrame) -> dict:
    """
    Generate comprehensive descriptive statistics for all columns.
    Includes measures of central tendency, dispersion, and distribution shape.
    """
    numeric_cols = df.select_dtypes(include=[np.number])
    summary = {
        'numeric': numeric_cols.describe().T.assign(
            skewness=numeric_cols.skew(),
            kurtosis=numeric_cols.kurtosis(),
            iqr=numeric_cols.quantile(0.75) - numeric_cols.quantile(0.25),
            cv=numeric_cols.std() / numeric_cols.mean()  # coefficient of variation
        ),
        'categorical': {
            col: df[col].value_counts().head(10).to_dict()
            for col in df.select_dtypes(include=['object']).columns
        },
        'correlations': numeric_cols.corr().round(3)
    }
    return summary
Normality and Distribution Testing
TestUse CaseFunction
Shapiro-WilkNormality test (n < 5000)scipy.stats.shapiro()
D'Agostino-PearsonNormality test (n >= 5000)scipy.stats.normaltest()
Kolmogorov-SmirnovCompare to any distributionscipy.stats.kstest()
Levene's testHomogeneity of variancescipy.stats.levene()

Best Practices for Reproducibility

  • Always save the raw CSV separately; never overwrite original files.
  • Log every cleaning step with timestamps in a transformation audit trail.
  • Export cleaned datasets with a version suffix (e.g., data_v2_cleaned.csv).
  • Include the cleaning script or notebook alongside the published dataset.
  • Report the number of rows removed at each step in your methods section.
  • Use random_state parameters consistently for any stochastic operations.

References

  • McKinney, W. (2022). Python for Data Analysis (3rd ed.). O'Reilly Media.
  • Wickham, H. (2014). Tidy Data. Journal of Statistical Software, 59(10).
  • Van den Broeck, J., et al. (2005). Data Cleaning: Detecting, Diagnosing, and Editing Data Abnormalities. PLoS Medicine, 2(10).

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/analysis/wrangling/csv-data-analyzer of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

CSV Data Analyzer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CSV Data Analyzer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CSV Data Analyzer this skillwentorai/research-plugins2981 repos~1.7kAutomated safety check: PassMIT
Data Cleanupsgharlow/claude-code-recipes388—~566Automated safety check: PassCustom licence
Sn Da Image CaptionMichaelYang-lyx/AIDABench1111 repos~2kAutomated safety check: PassNone
Douban Skilldaymade/claude-code-skills1.4k—~1.1kAutomated safety check: PassMIT
Champion Trackergooseworks-ai/goose-skills1.2k1 repos~1.1kAutomated safety check: NotesMIT
Tabular Cleanupgaasher/Agent-Loop-Skills174—~4kAutomated safety check: PassMIT

Similar skills

  • Data Cleanup

    sgharlow/claude-code-recipes

    Clean and standardize messy tabular data (CSV, spreadsheet paste, system exports) into an analysis-ready dataset — consistent dates and names, typed columns, duplicates identified, missing values…

    388 GitHub stars~566 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Sn Da Image Caption

    MichaelYang-lyx/AIDABench

    图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py…

    111 GitHub starsUsed in 1 repo~2k tokens
    Documents & OfficeAuto-check passed
  • Douban Skill

    daymade/claude-code-skills

    Export and sync Douban (豆瓣) book/movie/music/game collections to local CSV files via Frodo API.

    1.4k GitHub stars~1.1k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Champion Tracker

    gooseworks-ai/goose-skills

    Track product champions for job changes and qualify their new companies against ICP.

    1.2k GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check: notes
  • Tabular Cleanup

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a messy tabular data dump (CSV/TSV/parquet/Excel/JSON) and wants it iteratively cleaned to an inferred data contract — a checklist of deterministic…

    174 GitHub stars~4k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Meta Sensitivity Plot

    aipoch/medical-research-skills

    Generate leave-one-out sensitivity analysis plots for meta-analysis.

    2k GitHub stars~1.6k tokensUpdated 20 days ago
    Documents & OfficeAuto-check passed

More from wentorai/research-plugins

All 428 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about CSV Data Analyzer

What does CSV Data Analyzer do?

Load, explore, clean, and analyze CSV data with statistical summaries. CSV Data Analyzer is an agent skill from wentorai/research-plugins.

When should I use CSV Data Analyzer?

CSV Data Analyzer fits situations like: tasks that involve CSV and tabular files.

How do I install CSV Data Analyzer in Claude Code?

Run `npx skills add wentorai/research-plugins --skill csv-data-analyzer -a claude-code`. Or copy the skill folder (skills/analysis/wrangling/csv-data-analyzer in wentorai/research-plugins) into .claude/skills/csv-data-analyzer in your project. Claude Code loads it when a task matches its description.

How do I install CSV Data Analyzer in Codex?

Run `npx skills add wentorai/research-plugins --skill csv-data-analyzer -a codex`. Or copy the skill folder (skills/analysis/wrangling/csv-data-analyzer in wentorai/research-plugins) into .agents/skills/csv-data-analyzer in your project. Codex loads it when a task matches its description.

Can I use CSV Data Analyzer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill csv-data-analyzer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/csv-data-analyzer, .gemini/skills/csv-data-analyzer, .github/skills/csv-data-analyzer and .opencode/skills/csv-data-analyzer in your project.

What does CSV Data Analyzer need to run?

SKILL.md names no scripts, command-line tools or credentials: CSV Data Analyzer is instructions for the agent only. Our summary lists: Python 3.

Does CSV Data Analyzer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CSV Data Analyzer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does CSV Data Analyzer use?

CSV Data Analyzer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CSV Data Analyzer use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CSV Data Analyzer?

Skills that share tags, products or a category with CSV Data Analyzer: Data Cleanup (sgharlow/claude-code-recipes, 388 stars), Sn Da Image Caption (MichaelYang-lyx/AIDABench, 111 stars), Douban Skill (daymade/claude-code-skills, 1.4k stars) and Champion Tracker (gooseworks-ai/goose-skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CSV Data Analyzer?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 428 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.