Agent skill

Verified Data Analysis with pandas

by pipeshub-ai in pipeshub-ai/pipeshub-ai

Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

Apache-2.0Auto-check passedData & Analytics

Install Verified Data Analysis with pandas

skills CLI
$ npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pipeshub-ai/pipeshub-ai data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
3.8k
Token cost
~1.2k tokens
SKILL.md length
584 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

  • Works in 3 steps: Always run df.info() and df.describe()… → Cross-check any headline number by… → Never report a number your code didn't…
  • Computing totals or averages from a CSV, Excel or JSON export
  • SKILL.md covers Loading, Cleaning, Aggregation, joins, pivots and Verification discipline (do…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill assumes pandas, numpy and scipy are already installed in the sandbox. Its loading rules pin column types when the schema is known so ID-like values keep leading zeros, parse date columns explicitly so sorting and range filters are chronological, and name the sheet when reading Excel files instead of trusting the default first sheet. Cleaning rules ask for null and duplicate counts before any aggregation, with an explicit decision to drop, fill or flag, explicit numeric coercion for columns that only look numeric, and a before and after row count after every step that removes or changes rows.

For aggregation, each column gets its own named function in `groupby` calls instead of a bare sum over the whole frame. The guiding discipline is that reported figures come only from code output and never from inference or estimation, which suits questions such as totals by region, cleaning a dataset or joining two files on a customer ID.

When your agent uses it

  • Computing totals or averages from a CSV, Excel or JSON export
  • Cleaning a messy dataset without silently dropping rows
  • Joining two files on a shared key and checking the row counts

Example prompts

  • “What is the total revenue by region in sales_2025.csv?”
  • “Join orders.xlsx and customers.csv on customer_id and tell me how many orders have no match.”
  • “Clean up this survey export and report how many rows each step removes.”

Requirements

  • A Python sandbox with pandas, numpy and scipy

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Always run df.info() and df.describe() (or .value_counts() for categorical columns) before analyzing — this catches wrong dtypes…
  2. Cross-check any headline number by computing it a second, independent way — e.g. if you report a total via groupby().sum(), also…
  3. Never report a number your code didn't actually print. If you're describing a trend or a comparison, make sure the specific figures you…

What it can do on your machine

Read from SKILL.md and the folder at commit f5aee03. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verified Data Analysis with pandas loads about 1.2k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 584 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pipeshub-ai/pipeshub-ai at commit f5aee03, republished under its Apache-2.0 licence (© pipeshub-ai). 584 words, ~1,202 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder).
name
data-analysis
description
Use this skill whenever the user asks you to load, clean, aggregate, join, or otherwise analyze tabular data (csv, xlsx, json, or a database export) — 'what's the total revenue by region', 'clean up this dataset', 'join these two files on customer id'. Establishes a verification discipline so reported numbers are always numbers the code actually printed, never numbers you inferred or estimated.

Data analysis with pandas

pandas, numpy, and scipy are already installed in the sandbox — no install_packages call needed for any workflow below.

Loading

  • Pin dtypes on load rather than letting pandas infer them, when you know the schema — pd.read_csv(path, dtype={"customer_id": str, "amount": float}). Inference is usually right but silently wrong in a specific, dangerous way for ID-like columns: a numeric-looking ID (e.g. "00123") inferred as int64 loses its leading zeros.
  • Parse date columns explicitly (parse_dates=["order_date"] or pd.to_datetime(df["col"]) post-load) rather than leaving them as strings — string dates sort and filter lexicographically, not chronologically, which silently produces wrong "most recent N" or date-range results.
  • For Excel sources with multiple sheets, always pass sheet_name= explicitly (or inspect pd.ExcelFile(path).sheet_names first) — read_excel's default of the first sheet is a common source of "the numbers don't match" bugs when the data you actually need is on a different sheet.

Cleaning

  • Check for nulls (df.isnull().sum()) and duplicates (df.duplicated().sum()) before any aggregation — decide explicitly whether to drop, fill, or flag them, rather than letting them silently skew a sum/mean (pandas aggregations skip NaN by default, which is not always the right call — a "total" that silently excludes rows with a missing value is a different number than "total of all rows", and the user should know which one they're getting).
  • Coerce types explicitly after cleaning (pd.to_numeric(df["col"], errors="coerce")) rather than assuming a column that "looks numeric" already is — mixed-type columns from a manually-maintained spreadsheet are common, and a silent string comparison instead of a numeric one produces wrong sort/filter results without raising any error.
  • After any cleaning step that drops or modifies rows, print the before/after row count — a cleaning step that unexpectedly drops 80% of the data is a bug, not a result, and should be caught immediately rather than discovered after the whole analysis is built on the reduced dataset.
Show full SKILL.md (281 more words)Show less

Aggregation, joins, pivots

  • groupby(...).agg(...) for aggregates; always specify the aggregation function explicitly per column (.agg({"amount": "sum", "order_id": "count"})) rather than a bare .sum() across the whole frame, which silently includes every numeric column whether or not it makes sense to sum it.
  • For joins (pd.merge), always specify how= explicitly (don't rely on the "inner" default when the user's intent might be "left"/"outer") and check the resulting row count against expectations — an unexpected row-count change after a merge (either fewer rows than the left frame, or way more) almost always means a many-to-many relationship you didn't account for, or a key-matching problem (mismatched types, whitespace, or casing in the join column).
  • pivot_table for cross-tabulations; pass fill_value=0 explicitly when a missing combination should read as zero rather than NaN, since which one is correct depends on the question being asked.

Verification discipline (do this before reporting any number)

  1. Always run df.info() and df.describe() (or .value_counts() for categorical columns) before analyzing — this catches wrong dtypes, unexpected nulls, and outliers early, before they propagate into a wrong conclusion.
  2. Cross-check any headline number by computing it a second, independent way — e.g. if you report a total via groupby().sum(), also sanity-check it against df["amount"].sum() on the ungrouped frame (should match), or against a manual filter-and-sum on a known subset. Two different code paths agreeing is meaningfully more trustworthy than one path you didn't double-check.
  3. Never report a number your code didn't actually print. If you're describing a trend or a comparison, make sure the specific figures you state in your response are ones that appeared in your program's printed output — do not round, estimate, or extrapolate a number that wasn't literally computed and displayed.

© pipeshub-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis of pipeshub-ai/pipeshub-ai.

Open the folder on GitHubat commit f5aee03

Compare with similar skills

Verified Data Analysis with pandas next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verified Data Analysis with pandas compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verified Data Analysis with pandas this skillpipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent883—~2.5kAutomated safety check: PassApache-2.0
Data AnalysisEXboys/skilllite170—~176Automated safety check: PassMIT
CSV and Excel MergerOneWave-AI/claude-skills323—~1.6kAutomated safety check: PassMIT
Codebookbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~527Automated safety check: NotesCustom licence
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    883 GitHub stars~2.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Data Analysis

    EXboys/skilllite

    Analyze CSV/JSON data with statistics, filtering, and aggregation.

    170 GitHub stars~176 tokensUpdated 13 days ago
    Data & AnalyticsAuto-check passed
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    323 GitHub stars~1.6k tokensUpdated 6 days ago
    Documents & OfficeAuto-check passed
  • Codebook

    brycewang-stanford/Auto-Empirical-Research-Skills

    Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics.

    4.5k GitHub stars~527 tokensUpdated 3 days ago
    Documents & OfficeAuto-check: notes
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed

More from pipeshub-ai/pipeshub-ai

All 9 skills in this repo
  • PipesHub Excel Spreadsheet Builder

    pipeshub-ai/pipeshub-ai

    Creates and edits .xlsx workbooks with real Excel formulas rather than hardcoded computed values, defaulting to exceljs in TypeScript with a static formula-safety check.

    3.8k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Office Open XML Utilities

    pipeshub-ai/pipeshub-ai

    Unpacks a .docx or .pptx into pretty-printed XML, lets you make small targeted edits, and repacks it into a file Office will open.

    3.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • PowerPoint Deck Builder

    pipeshub-ai/pipeshub-ai

    Creates new PowerPoint decks with pptxgenjs in TypeScript, reads existing decks with python-pptx, and applies a design-quality checklist so every slide has real visual hierarchy.

    3.8k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Chart Type Selection Guide

    pipeshub-ai/pipeshub-ai

    Picks the right chart type for a data question and applies readability rules like axis labels, colorblind palettes and legend restraint.

    3.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Word Document Creation And Editing

    pipeshub-ai/pipeshub-ai

    Routes a Word document request to the right approach: a TypeScript library for new files, XML editing for existing ones, and plain reading only.

    3.8k GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Questions about Verified Data Analysis with pandas

What does Verified Data Analysis with pandas do?

Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed. The skill assumes pandas, numpy and scipy are already installed in the sandbox. Its loading rules pin column types when the schema is known so ID-like values keep leading zeros, parse date columns explicitly so sorting and range filters are chronological, and name the sheet when reading Excel files instead of trusting the default first sheet.

When should I use Verified Data Analysis with pandas?

Verified Data Analysis with pandas fits situations like: computing totals or averages from a CSV, Excel or JSON export; cleaning a messy dataset without silently dropping rows; joining two files on a shared key and checking the row counts.

How do I install Verified Data Analysis with pandas in Claude Code?

Run `npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a claude-code`. Or copy the skill folder (backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis in pipeshub-ai/pipeshub-ai) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Verified Data Analysis with pandas in Codex?

Run `npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a codex`. Or copy the skill folder (backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis in pipeshub-ai/pipeshub-ai) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Verified Data Analysis with pandas in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Verified Data Analysis with pandas need to run?

SKILL.md names no scripts, command-line tools or credentials: Verified Data Analysis with pandas is instructions for the agent only. Our summary lists: A Python sandbox with pandas, numpy and scipy.

Does Verified Data Analysis with pandas access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verified Data Analysis with pandas safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verified Data Analysis with pandas use?

Verified Data Analysis with pandas is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verified Data Analysis with pandas use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Verified Data Analysis with pandas?

Skills that share tags, products or a category with Verified Data Analysis with pandas: Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 883 stars), Data Analysis (EXboys/skilllite, 170 stars), CSV and Excel Merger (OneWave-AI/claude-skills, 323 stars) and Codebook (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verified Data Analysis with pandas?

pipeshub-ai (a GitHub organization) maintains it in pipeshub-ai/pipeshub-ai, which has 3,816 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.

Source: pipeshub-ai/pipeshub-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.