Agent skill

Data Analysis

by HezaoHezao in HezaoHezao/poirot

Analyze Excel/CSV files with DuckDB SQL via bash. An agent skill from HezaoHezao/poirot.

MITAuto-check passedDocuments & Office

Install Data Analysis

skills CLI
$ npx skills add HezaoHezao/poirot --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HezaoHezao/poirot data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HezaoHezao/poirot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/poirot/backend/agents/skill/builtin_skills/research/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
250
Token cost
~997 tokens
SKILL.md length
187 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Analyze Excel/CSV files with DuckDB SQL via bash. An agent skill from HezaoHezao/poirot.

  • Works in 4 steps: Inspect File Structure → Statistical Summary → SQL Queries → …
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers Overview, When to Use, Prerequisites and Workflow, plus 2 more sections
  • Calls python3 and pip

What it does

Data Analysis is an agent skill from HezaoHezao/poirot. Analyze Excel/CSV files with DuckDB SQL via bash.

Its SKILL.md is about 1000 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Excel spreadsheets, CSV and tabular files and SQL. It works with SQL, Microsoft Excel, DuckDB and Bash. The repository describes itself as: Poirot is a deep research agent kernel built for those who care about how agents are architected. The licence is MIT.

When your agent uses it

  • Tasks that involve Excel spreadsheets
  • Tasks that involve CSV and tabular files
  • Tasks that involve SQL

Example prompts

  • “/data-analysis”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): bash, read_file, write_file

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Inspect File Structure
  2. Statistical Summary
  3. SQL Queries
  4. Export Results

What it can do on your machine

Read from SKILL.md and the folder at commit 86bf279. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • bash
    • read_file
    • write_file

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Analysis loads about 997 tokens when it runs. Until then it costs about 16 tokens; SKILL.md has 187 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~16
When it runs · the whole SKILL.md, loaded when a task matches
~997

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HezaoHezao/poirot at commit 86bf279, republished under its MIT licence (© HezaoHezao). 187 words, ~997 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder).
name
data-analysis
description
Analyze Excel/CSV files with DuckDB SQL via bash.
allowed-tools
bash, read_file, write_file
enabled
true
related-skills
consulting-analysis
license
MIT
author
Adapted from deer-flow (Bytedance, MIT)

Data Analysis

Overview

Analyzes user-provided Excel (.xlsx/.xls) or CSV files using DuckDB — an in-process analytical SQL engine. Supports schema inspection, SQL querying, statistical summaries, and result export.

Poirot note: The original deer-flow skill uses a bundled scripts/analyze.py helper. Poirot doesn't bundle that script, so this version uses bash with python3 + duckdb directly. Install duckdb first: pip install duckdb.

When to Use

  • User uploads Excel/CSV files and wants analysis
  • User wants statistics, summaries, pivot tables, or SQL queries on data
  • User wants to filter, join, or aggregate structured data

Prerequisites

bash
# Install duckdb if not present
pip install duckdb openpyxl

Workflow

Step 1: Inspect File Structure
bash
python3 -c "
import duckdb
con = duckdb.connect()
# For CSV
result = con.execute(\"DESCRIBE SELECT * FROM read_csv_auto('data.csv')\").fetchall()
for col in result:
    print(f'{col[0]:30s} {col[1]}')

# For Excel (each sheet = a table)
result = con.execute(\"SELECT * FROM st_read('data.xlsx', layer='Sheet1') LIMIT 0\").fetchall()

# Row count
count = con.execute(\"SELECT COUNT(*) FROM read_csv_auto('data.csv')\").fetchone()[0]
print(f'Rows: {count}')
"
Step 2: Statistical Summary
bash
python3 -c "
import duckdb
con = duckdb.connect()
# Describe statistics
print(con.execute(\"SUMMARIZE SELECT * FROM read_csv_auto('data.csv')\").df().to_string())
"
Step 3: SQL Queries
bash
python3 -c "
import duckdb
con = duckdb.connect()

# Aggregation
result = con.execute('''
    SELECT category, COUNT(*) as count, AVG(price) as avg_price
    FROM read_csv_auto('data.csv')
    GROUP BY category
    ORDER BY count DESC
''').fetchall()
for row in result:
    print(row)

# Join two files
result = con.execute('''
    SELECT a.id, a.name, b.amount
    FROM read_csv_auto('orders.csv') a
    JOIN read_csv_auto('payments.csv') b ON a.id = b.order_id
''').fetchall()
"
Step 4: Export Results
bash
python3 -c "
import duckdb
con = duckdb.connect()
# Export to CSV
con.execute(\"COPY (SELECT * FROM read_csv_auto('data.csv') WHERE amount > 100) TO 'filtered.csv' (HEADER, DELIMITER ',')\")
# Export to JSON
con.execute(\"COPY (SELECT * FROM read_csv_auto('data.csv')) TO 'output.json' (FORMAT JSON)\")
"

Common Patterns

Pivot table
sql
SELECT
    product,
    SUM(CASE WHEN month = 'Jan' THEN amount ELSE 0 END) AS jan,
    SUM(CASE WHEN month = 'Feb' THEN amount ELSE 0 END) AS feb,
    SUM(CASE WHEN month = 'Mar' THEN amount ELSE 0 END) AS mar
FROM read_csv_auto('sales.csv')
GROUP BY product
Percentiles
sql
SELECT
    percentile_cont(0.5) WITHIN GROUP (ORDER BY price) AS median,
    percentile_cont(0.95) WITHIN GROUP (ORDER BY price) AS p95
FROM read_csv_auto('data.csv')
Multi-sheet Excel
bash
python3 -c "
import duckdb
con = duckdb.connect()
# List sheets
sheets = con.execute(\"SELECT table_name FROM st_geometry_tables()\").fetchall()
# Query specific sheet
result = con.execute(\"SELECT * FROM st_read('data.xlsx', layer='Sheet2') LIMIT 10\").fetchall()
"

Pitfalls

  • DuckDB not installed: pip install duckdb openpyxl first
  • Large files: DuckDB handles large files well, but SUMMARIZE on very large datasets may be slow. Sample first: SELECT * FROM ... TABLESAMPLE 10%
  • Encoding: CSV with non-UTF-8 encoding may fail. Specify encoding in read_csv_auto options.
  • Date parsing: DuckDB auto-detects dates, but ambiguous formats may need explicit strptime parsing.
  • Excel formulas: st_read reads cell values, not formula results. Use openpyxl directly if you need computed values.

© HezaoHezao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in poirot/backend/agents/skill/builtin_skills/research/data-analysis of HezaoHezao/poirot.

Open the folder on GitHubat commit 86bf279

Compare with similar skills

Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Analysis this skillHezaoHezao/poirot250—~997Automated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Rust SQL Testshencangsheng/easydb_app590—~1.2kAutomated safety check: PassMIT
Mindsdb MCP SkillLeoYeAI/openclaw-master-skills2.2k—~2.5kAutomated safety check: PassMIT
Markdown Exporterbowenliang123/markdown-exporter2711 repos~5.3kAutomated safety check: PassApache-2.0
Gaik ToolkitGAIK-project/gaik-toolkit100—~5.7kAutomated safety check: PassMIT

Similar skills

  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Rust SQL Test

    shencangsheng/easydb_app

    Enforces unit test requirements for the Rust data-processing modules in src-tauri/src/sql/ (generator.rs, parse.rs) and src-tauri/src/reader/ (excel.rs and other readers).

    590 GitHub stars~1.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mindsdb MCP Skill

    LeoYeAI/openclaw-master-skills

    MindsDB MCP服务器交互技能,用于通过自然语言查询和操作200+企业级数据源。当用户需要查询数据库、分析数据、创建AI模型、连接数据源(MySQL、PostgreSQL、MongoDB、Excel、CSV、Gmail、Slack等)、执行SQL查询、进行数据预测、构建知识库(RAG)、智能问答、文档检索或任何与数据库交互的任务时使用此技能。即使没有明确提到MindsDB,只要涉及数据库操…

    2.2k GitHub stars~2.5k tokensUpdated 2 mo ago
    DatabasesAuto-check passed
  • Markdown Exporter

    bowenliang123/markdown-exporter

    Convert Markdown text to DOCX, PPTX, XLSX, PDF, PNG, SVG, HTML, IPYNB, MD, CSV, JSON, JSONL, XML files, and extract code blocks in Markdown to Python, Bash,JS and etc files.

    271 GitHub starsUsed in 1 repo~5.3k tokens
    Documents & OfficeAuto-check passed
  • Gaik Toolkit

    GAIK-project/gaik-toolkit

    GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.

    100 GitHub stars~5.7k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Forge Codegen Crud

    yaomindong1996/forge-admin

    Generate or review Forge project code-generation output for CRUD modules.

    125 GitHub stars~1.3k tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from HezaoHezao/poirot

All 23 skills in this repo
  • Requesting Code Review

    HezaoHezao/poirot

    Pre-commit review: security scan, quality gates, auto-fix. An agent skill from HezaoHezao/poirot.

    250 GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check passed
  • Research Paper Writing

    HezaoHezao/poirot

    ML paper pipeline: experiment design to submission. An agent skill from HezaoHezao/poirot.

    250 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Paper Review

    HezaoHezao/poirot

    Structured peer-review of academic papers. An agent skill from HezaoHezao/poirot.

    250 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Blogwatcher

    HezaoHezao/poirot

    Monitor blogs and RSS/Atom feeds via blogwatcher-cli. An agent skill from HezaoHezao/poirot.

    250 GitHub stars~854 tokensUpdated 2 mo ago
    Auto-check passed
  • Bootstrap

    HezaoHezao/poirot

    Onboarding conversation to generate a user profile. An agent skill from HezaoHezao/poirot.

    250 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Chart Visualization

    HezaoHezao/poirot

    Generate charts: select type, extract data, render image. An agent skill from HezaoHezao/poirot.

    250 GitHub stars~1k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Data Analysis

What does Data Analysis do?

Analyze Excel/CSV files with DuckDB SQL via bash. An agent skill from HezaoHezao/poirot. Data Analysis is an agent skill from HezaoHezao/poirot. Analyze Excel/CSV files with DuckDB SQL via bash.

When should I use Data Analysis?

Data Analysis fits situations like: tasks that involve Excel spreadsheets; tasks that involve CSV and tabular files; tasks that involve SQL.

How do I install Data Analysis in Claude Code?

Run `npx skills add HezaoHezao/poirot --skill data-analysis -a claude-code`. Or copy the skill folder (poirot/backend/agents/skill/builtin_skills/research/data-analysis in HezaoHezao/poirot) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Data Analysis in Codex?

Run `npx skills add HezaoHezao/poirot --skill data-analysis -a codex`. Or copy the skill folder (poirot/backend/agents/skill/builtin_skills/research/data-analysis in HezaoHezao/poirot) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HezaoHezao/poirot --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Data Analysis need to run?

Going by SKILL.md and its folder, Data Analysis needs the command-line tools its instructions call (python3 and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: bash, read_file, write_file.

Does Data Analysis access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Analysis use?

Data Analysis is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Analysis use?

About 997 tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Analysis?

Skills that share tags, products or a category with Data Analysis: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Rust SQL Test (shencangsheng/easydb_app, 590 stars), Mindsdb MCP Skill (LeoYeAI/openclaw-master-skills, 2.2k stars) and Markdown Exporter (bowenliang123/markdown-exporter, 271 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Analysis?

HezaoHezao (a GitHub user) maintains it in HezaoHezao/poirot, which has 250 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on July 28, 2026.

Source: HezaoHezao/poirot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.