Data Table Analysis
NVIDIA-AI-Blueprints/deep-researcher-agent
A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.
Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.
$ npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pipeshub-ai/pipeshub-ai data-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis .claude/skills/data-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-analysis" agent skill from https://github.com/pipeshub-ai/pipeshub-ai/tree/main/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis into .claude/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pipeshub-ai/pipeshub-ai/tree/main/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pipeshub-ai/pipeshub-ai data-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .agents/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis .agents/skills/data-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pipeshub-ai/pipeshub-ai/tree/main/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis into .agents/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pipeshub-ai/pipeshub-ai data-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis .cursor/skills/data-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-analysis" agent skill from https://github.com/pipeshub-ai/pipeshub-ai/tree/main/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis into .cursor/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pipeshub-ai/pipeshub-ai.git --path backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pipeshub-ai/pipeshub-ai data-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis .gemini/skills/data-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pipeshub-ai/pipeshub-ai/tree/main/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis into .gemini/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pipeshub-ai/pipeshub-ai data-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .github/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis .github/skills/data-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pipeshub-ai/pipeshub-ai/tree/main/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis into .github/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pipeshub-ai/pipeshub-ai data-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis .opencode/skills/data-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-analysis" agent skill from https://github.com/pipeshub-ai/pipeshub-ai/tree/main/backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis into .opencode/skills/data-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-analysisLoads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.
The skill assumes pandas, numpy and scipy are already installed in the sandbox. Its loading rules pin column types when the schema is known so ID-like values keep leading zeros, parse date columns explicitly so sorting and range filters are chronological, and name the sheet when reading Excel files instead of trusting the default first sheet. Cleaning rules ask for null and duplicate counts before any aggregation, with an explicit decision to drop, fill or flag, explicit numeric coercion for columns that only look numeric, and a before and after row count after every step that removes or changes rows.
For aggregation, each column gets its own named function in `groupby` calls instead of a bare sum over the whole frame. The guiding discipline is that reported figures come only from code output and never from inference or estimation, which suits questions such as totals by region, cleaning a dataset or joining two files on a customer ID.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f5aee03. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Verified Data Analysis with pandas loads about 1.2k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 584 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from pipeshub-ai/pipeshub-ai at commit f5aee03, republished under its Apache-2.0 licence (© pipeshub-ai). 584 words, ~1,202 tokens.
.claude/skills/data-analysis/SKILL.md (or your agent's skills folder).pandas, numpy, and scipy are already installed in the sandbox — no install_packages call needed for any workflow below.
pd.read_csv(path, dtype={"customer_id": str, "amount": float}). Inference is usually right but silently wrong in a specific, dangerous way for ID-like columns: a numeric-looking ID (e.g. "00123") inferred as int64 loses its leading zeros.parse_dates=["order_date"] or pd.to_datetime(df["col"]) post-load) rather than leaving them as strings — string dates sort and filter lexicographically, not chronologically, which silently produces wrong "most recent N" or date-range results.sheet_name= explicitly (or inspect pd.ExcelFile(path).sheet_names first) — read_excel's default of the first sheet is a common source of "the numbers don't match" bugs when the data you actually need is on a different sheet.df.isnull().sum()) and duplicates (df.duplicated().sum()) before any aggregation — decide explicitly whether to drop, fill, or flag them, rather than letting them silently skew a sum/mean (pandas aggregations skip NaN by default, which is not always the right call — a "total" that silently excludes rows with a missing value is a different number than "total of all rows", and the user should know which one they're getting).pd.to_numeric(df["col"], errors="coerce")) rather than assuming a column that "looks numeric" already is — mixed-type columns from a manually-maintained spreadsheet are common, and a silent string comparison instead of a numeric one produces wrong sort/filter results without raising any error.groupby(...).agg(...) for aggregates; always specify the aggregation function explicitly per column (.agg({"amount": "sum", "order_id": "count"})) rather than a bare .sum() across the whole frame, which silently includes every numeric column whether or not it makes sense to sum it.pd.merge), always specify how= explicitly (don't rely on the "inner" default when the user's intent might be "left"/"outer") and check the resulting row count against expectations — an unexpected row-count change after a merge (either fewer rows than the left frame, or way more) almost always means a many-to-many relationship you didn't account for, or a key-matching problem (mismatched types, whitespace, or casing in the join column).pivot_table for cross-tabulations; pass fill_value=0 explicitly when a missing combination should read as zero rather than NaN, since which one is correct depends on the question being asked.df.info() and df.describe() (or .value_counts() for categorical columns) before analyzing — this catches wrong dtypes, unexpected nulls, and outliers early, before they propagate into a wrong conclusion.groupby().sum(), also sanity-check it against df["amount"].sum() on the ungrouped frame (should match), or against a manual filter-and-sum on a known subset. Two different code paths agreeing is meaningfully more trustworthy than one path you didn't double-check.© pipeshub-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis of pipeshub-ai/pipeshub-ai.
Open the folder on GitHubat commit f5aee03
Verified Data Analysis with pandas next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Verified Data Analysis with pandas this skillpipeshub-ai/pipeshub-ai | 3.8k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent | 883 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Data AnalysisEXboys/skilllite | 170 | — | ~176 | Automated safety check: Pass | MIT | |
| CSV and Excel MergerOneWave-AI/claude-skills | 323 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Codebookbrycewang-stanford/Auto-Empirical-Research-Skills | 4.5k | — | ~527 | Automated safety check: Notes | Custom licence | |
| Python Executorcortega26/chile-hub | 113 | 2 repos | ~1.5k | Automated safety check: Pass | MIT |
NVIDIA-AI-Blueprints/deep-researcher-agent
A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.
EXboys/skilllite
Analyze CSV/JSON data with statistics, filtering, and aggregation.
OneWave-AI/claude-skills
Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.
brycewang-stanford/Auto-Empirical-Research-Skills
Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics.
cortega26/chile-hub
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).
Jeffallan/claude-skills
Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.
pipeshub-ai/pipeshub-ai
Creates and edits .xlsx workbooks with real Excel formulas rather than hardcoded computed values, defaulting to exceljs in TypeScript with a static formula-safety check.
pipeshub-ai/pipeshub-ai
Unpacks a .docx or .pptx into pretty-printed XML, lets you make small targeted edits, and repacks it into a file Office will open.
pipeshub-ai/pipeshub-ai
Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.
pipeshub-ai/pipeshub-ai
Creates new PowerPoint decks with pptxgenjs in TypeScript, reads existing decks with python-pptx, and applies a design-quality checklist so every slide has real visual hierarchy.
pipeshub-ai/pipeshub-ai
Picks the right chart type for a data question and applies readability rules like axis labels, colorblind palettes and legend restraint.
pipeshub-ai/pipeshub-ai
Routes a Word document request to the right approach: a TypeScript library for new files, XML editing for existing ones, and plain reading only.
Works with
Categories
Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed. The skill assumes pandas, numpy and scipy are already installed in the sandbox. Its loading rules pin column types when the schema is known so ID-like values keep leading zeros, parse date columns explicitly so sorting and range filters are chronological, and name the sheet when reading Excel files instead of trusting the default first sheet.
Verified Data Analysis with pandas fits situations like: computing totals or averages from a CSV, Excel or JSON export; cleaning a messy dataset without silently dropping rows; joining two files on a shared key and checking the row counts.
Run `npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a claude-code`. Or copy the skill folder (backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis in pipeshub-ai/pipeshub-ai) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a codex`. Or copy the skill folder (backend/python/app/agents/agent_loop/skills/builtin_packs/data-analysis in pipeshub-ai/pipeshub-ai) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pipeshub-ai/pipeshub-ai --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.
SKILL.md names no scripts, command-line tools or credentials: Verified Data Analysis with pandas is instructions for the agent only. Our summary lists: A Python sandbox with pandas, numpy and scipy.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Verified Data Analysis with pandas is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Verified Data Analysis with pandas: Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 883 stars), Data Analysis (EXboys/skilllite, 170 stars), CSV and Excel Merger (OneWave-AI/claude-skills, 323 stars) and Codebook (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pipeshub-ai (a GitHub organization) maintains it in pipeshub-ai/pipeshub-ai, which has 3,816 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 8, 2026.
Source: pipeshub-ai/pipeshub-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.