Agent skill

Excel and CSV Data Analysis

by bytedance in bytedance/deer-flow

Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

MITAuto-check passedData & Analytics

Install Excel and CSV Data Analysis

skills CLI
$ npx skills add bytedance/deer-flow --skill data-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bytedance/deer-flow data-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public/data-analysis .claude/skills/data-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-analysis
GitHub stars
83k
Used in
4 other repos
Token cost
~2.2k tokens
SKILL.md length
649 words
Files
2 (incl. scripts)
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

  • Works in 7 steps: Understand Requirements → Inspect File Structure → Perform Analysis → …
  • Exploring an uploaded spreadsheet to see its sheets, columns and types
  • SKILL.md covers Overview, Core Capabilities, Workflow and Table Naming Rules, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Everything runs through one Python script, `scripts/analyze.py`, which loads uploaded workbooks or CSV files into DuckDB, an in-process analytical SQL engine. Each sheet of a multi-sheet Excel workbook becomes its own table, so the agent can filter, aggregate and join across sheets with ordinary SQL.

The workflow has three actions. Inspect returns sheet names, column names and types, non-null counts, row counts and a first look at the data. Query runs SQL written from that schema. Summary reports count, mean, standard deviation, percentiles and null counts for numeric columns, and count, unique values, top value and frequency for text columns. Results can be saved as CSV, JSON or Markdown, chosen by the file extension. Uploaded files are expected under `/mnt/user-data/uploads/`.

When your agent uses it

  • Exploring an uploaded spreadsheet to see its sheets, columns and types
  • Answering a question about Excel data with a SQL aggregation
  • Joining two sheets or files and exporting the result
  • Getting descriptive statistics for every column in a CSV file

Example prompts

  • “Inspect quarterly-sales.xlsx and tell me what each sheet contains.”
  • “Total revenue by region from the Orders sheet, sorted high to low, saved as a CSV.”
  • “Give me a statistical summary of customers.csv, including how many values are missing per column.”

Requirements

  • Python with DuckDB available to `scripts/analyze.py`
  • Excel or CSV files in the uploads folder

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Understand Requirements
  2. Inspect File Structure
  3. Perform Analysis
  4. Inspect the file
  5. Top products by revenue
  6. Monthly revenue trends
  7. Statistical summary

What it can do on your machine

Read from SKILL.md and the folder at commit 35cdcab. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Excel and CSV Data Analysis loads about 2.2k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 649 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bytedance/deer-flow at commit 35cdcab, republished under its MIT licence (© bytedance). 649 words, ~2,212 tokens.

Download SKILL.mdSave it as .claude/skills/data-analysis/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
data-analysis
description
Use this skill when the user uploads Excel (.xlsx/.xls) or CSV files and wants to perform data analysis, generate statistics, create summaries, pivot tables, SQL queries, or any form of structured data exploration. Supports multi-sheet Excel workbooks, aggregation, filtering, joins, and exporting results to CSV/JSON/Markdown.

Data Analysis Skill

Overview

This skill analyzes user-uploaded Excel/CSV files using DuckDB — an in-process analytical SQL engine. It supports schema inspection, SQL-based querying, statistical summaries, and result export, all through a single Python script.

Core Capabilities

  • Inspect Excel/CSV file structure (sheets, columns, types, row counts)
  • Execute arbitrary SQL queries against uploaded data
  • Generate statistical summaries (mean, median, stddev, percentiles, nulls)
  • Support multi-sheet Excel workbooks (each sheet becomes a table)
  • Export query results to CSV, JSON, or Markdown
  • Handle large files efficiently with DuckDB's columnar engine

Workflow

Step 1: Understand Requirements

When a user uploads data files and requests analysis, identify:

  • File location: Path(s) to uploaded Excel/CSV files under /mnt/user-data/uploads/
  • Analysis goal: What insights the user wants (summary, filtering, aggregation, comparison, etc.)
  • Output format: How results should be presented (table, CSV export, JSON, etc.)
  • You don't need to check the folder under /mnt/user-data
Step 2: Inspect File Structure

First, inspect the uploaded file to understand its schema:

bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/data.xlsx \
  --action inspect

This returns:

  • Sheet names (for Excel) or filename (for CSV)
  • Column names, data types, and non-null counts
  • Row count per sheet/file
  • Sample data (first 5 rows)
Step 3: Perform Analysis

Based on the schema, construct SQL queries to answer the user's questions.

Run SQL Query
bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/data.xlsx \
  --action query \
  --sql "SELECT category, COUNT(*) as count, AVG(amount) as avg_amount FROM Sheet1 GROUP BY category ORDER BY count DESC"
Generate Statistical Summary
bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/data.xlsx \
  --action summary \
  --table Sheet1

This returns for each numeric column: count, mean, std, min, 25%, 50%, 75%, max, null_count. For string columns: count, unique, top value, frequency, null_count.

Export Results
bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/data.xlsx \
  --action query \
  --sql "SELECT * FROM Sheet1 WHERE amount > 1000" \
  --output-file /mnt/user-data/outputs/filtered-results.csv

Supported output formats (auto-detected from extension):

  • .csv — Comma-separated values
  • .json — JSON array of records
  • .md — Markdown table
Parameters
ParameterRequiredDescription
--filesYesSpace-separated paths to Excel/CSV files
--actionYesOne of: inspect, query, summary
--sqlFor querySQL query to execute
--tableFor summaryTable/sheet name to summarize
--output-fileNoPath to export results (CSV/JSON/MD)

[!NOTE] Do NOT read the Python file, just call it with the parameters.

Table Naming Rules

  • Excel files: Each sheet becomes a table named after the sheet (e.g., Sheet1, Sales, Revenue)
  • CSV files: Table name is the filename without extension (e.g., data.csv → data)
  • Multiple files: All tables from all files are available in the same query context, enabling cross-file joins
  • Special characters: Sheet/file names with spaces or special characters are auto-sanitized (spaces → underscores). Use double quotes for names that start with numbers or contain special characters, e.g., "2024_Sales"

Analysis Patterns

Basic Exploration
sql
-- Row count
SELECT COUNT(*) FROM Sheet1

-- Distinct values in a column
SELECT DISTINCT category FROM Sheet1

-- Value distribution
SELECT category, COUNT(*) as cnt FROM Sheet1 GROUP BY category ORDER BY cnt DESC

-- Date range
SELECT MIN(date_col), MAX(date_col) FROM Sheet1
Aggregation & Grouping
sql
-- Revenue by category and month
SELECT category, DATE_TRUNC('month', order_date) as month,
       SUM(revenue) as total_revenue
FROM Sales
GROUP BY category, month
ORDER BY month, total_revenue DESC

-- Top 10 customers by spend
SELECT customer_name, SUM(amount) as total_spend
FROM Orders GROUP BY customer_name
ORDER BY total_spend DESC LIMIT 10
Cross-file Joins
sql
-- Join sales with customer info from different files
SELECT s.order_id, s.amount, c.customer_name, c.region
FROM sales s
JOIN customers c ON s.customer_id = c.id
WHERE s.amount > 500
Window Functions
sql
-- Running total and rank
SELECT order_date, amount,
       SUM(amount) OVER (ORDER BY order_date) as running_total,
       RANK() OVER (ORDER BY amount DESC) as amount_rank
FROM Sales
Pivot-style Analysis
sql
-- Pivot: monthly revenue by category
SELECT category,
       SUM(CASE WHEN MONTH(date) = 1 THEN revenue END) as Jan,
       SUM(CASE WHEN MONTH(date) = 2 THEN revenue END) as Feb,
       SUM(CASE WHEN MONTH(date) = 3 THEN revenue END) as Mar
FROM Sales
GROUP BY category
Show full SKILL.md (263 more words)Show less

Complete Example

User uploads sales_2024.xlsx (with sheets: Orders, Products, Customers) and asks: "Analyze my sales data — show top products by revenue and monthly trends."

Step 1: Inspect the file
bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/sales_2024.xlsx \
  --action inspect
Step 2: Top products by revenue
bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/sales_2024.xlsx \
  --action query \
  --sql "SELECT p.product_name, SUM(o.quantity * o.unit_price) as total_revenue, SUM(o.quantity) as total_units FROM Orders o JOIN Products p ON o.product_id = p.id GROUP BY p.product_name ORDER BY total_revenue DESC LIMIT 10"
bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/sales_2024.xlsx \
  --action query \
  --sql "SELECT DATE_TRUNC('month', order_date) as month, SUM(quantity * unit_price) as revenue FROM Orders GROUP BY month ORDER BY month" \
  --output-file /mnt/user-data/outputs/monthly-trends.csv
Step 4: Statistical summary
bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/sales_2024.xlsx \
  --action summary \
  --table Orders

Present results to the user with clear explanations of findings, trends, and actionable insights.

Multi-file Example

User uploads orders.csv and customers.xlsx and asks: "Which region has the highest average order value?"

bash
python /mnt/skills/public/data-analysis/scripts/analyze.py \
  --files /mnt/user-data/uploads/orders.csv /mnt/user-data/uploads/customers.xlsx \
  --action query \
  --sql "SELECT c.region, AVG(o.amount) as avg_order_value, COUNT(*) as order_count FROM orders o JOIN Customers c ON o.customer_id = c.id GROUP BY c.region ORDER BY avg_order_value DESC"

Output Handling

After analysis:

  • Present query results directly in conversation as formatted tables
  • For large results, export to file and share via present_files tool
  • Always explain findings in plain language with key takeaways
  • Suggest follow-up analyses when patterns are interesting
  • Offer to export results if the user wants to keep them

Caching

The script automatically caches loaded data to avoid re-parsing files on every call:

  • On first load, files are parsed and stored in a persistent DuckDB database under /mnt/user-data/workspace/.data-analysis-cache/
  • The cache key is a SHA256 hash of all input file contents — if files change, a new cache is created
  • Subsequent calls with the same files will use the cached database directly (near-instant startup)
  • Cache is transparent — no extra parameters needed

This is especially useful when running multiple queries against the same data files (inspect → query → summary).

Notes

  • DuckDB supports full SQL including window functions, CTEs, subqueries, and advanced aggregations
  • Excel date columns are automatically parsed; use DuckDB date functions (DATE_TRUNC, EXTRACT, etc.)
  • For very large files (100MB+), DuckDB handles them efficiently without loading everything into memory
  • Column names with spaces are accessible using double quotes: "Column Name"

© bytedance, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/public/data-analysis of bytedance/deer-flow.

  • SKILL.md
  • scripts/analyze.py

Open the folder on GitHubat commit 35cdcab

Used in 4 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in bytedance/deer-flow, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Excel and CSV Data Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Excel and CSV Data Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Excel and CSV Data Analysis this skillbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Data AnalysisHezaoHezao/poirot250—~997Automated safety check: PassMIT
Data Analysisspytensor/openmozi439—~535Automated safety check: PassMIT
Raccoon DataanalysisSenseTime-Copilot/raccoon-dataanalysis-skill137—~1.9kAutomated safety check: PassNone
CSV Data Analysis5zjk5/prompt-engineering127—~2.6kAutomated safety check: PassNone
Rust SQL Testshencangsheng/easydb_app590—~1.2kAutomated safety check: PassMIT

Similar skills

  • Data Analysis

    HezaoHezao/poirot

    Analyze Excel/CSV files with DuckDB SQL via bash. An agent skill from HezaoHezao/poirot.

    250 GitHub stars~997 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Data Analysis

    spytensor/openmozi

    Data analysis workflow: ingest, validate quality, explore, analyze, report.

    439 GitHub stars~535 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Raccoon Dataanalysis

    SenseTime-Copilot/raccoon-dataanalysis-skill

    Raccoon (小浣熊) Data Analysis - Remote code interpreter and data visualization service powered by SenseTime.

    137 GitHub stars~1.9k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • CSV Data Analysis

    5zjk5/prompt-engineering

    This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.

    127 GitHub stars~2.6k tokensUpdated 22 days ago
    Data & AnalyticsAuto-check passed
  • Rust SQL Test

    shencangsheng/easydb_app

    Enforces unit test requirements for the Rust data-processing modules in src-tauri/src/sql/ (generator.rs, parse.rs) and src-tauri/src/reader/ (excel.rs and other readers).

    590 GitHub stars~1.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Gaik Toolkit

    GAIK-project/gaik-toolkit

    GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit.

    100 GitHub stars~5.7k tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from bytedance/deer-flow

All 23 skills in this repo
  • Vercel Deploy

    bytedance/deer-flow

    Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.

    83k GitHub starsUsed in 10 repos~797 tokens
    Auto-check passed
  • Chart Visualization

    bytedance/deer-flow

    Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.

    83k GitHub starsUsed in 2 repos~840 tokens
    Auto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Auto-check passed
  • Structured Image Generation

    bytedance/deer-flow

    Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.

    83k GitHub starsUsed in 5 repos~2.9k tokens
    Auto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    83k GitHub stars~2.5k tokensUpdated today
    Auto-check: notes
  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    83k GitHub starsUsed in 3 repos~1.4k tokens
    Auto-check passed

Questions about Excel and CSV Data Analysis

What does Excel and CSV Data Analysis do?

Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown. py`, which loads uploaded workbooks or CSV files into DuckDB, an in-process analytical SQL engine. Each sheet of a multi-sheet Excel workbook becomes its own table, so the agent can filter, aggregate and join across sheets with ordinary SQL.

When should I use Excel and CSV Data Analysis?

Excel and CSV Data Analysis fits situations like: exploring an uploaded spreadsheet to see its sheets, columns and types; answering a question about Excel data with a SQL aggregation; joining two sheets or files and exporting the result; getting descriptive statistics for every column in a CSV file.

How do I install Excel and CSV Data Analysis in Claude Code?

Run `npx skills add bytedance/deer-flow --skill data-analysis -a claude-code`. Or copy the skill folder (skills/public/data-analysis in bytedance/deer-flow) into .claude/skills/data-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Excel and CSV Data Analysis in Codex?

Run `npx skills add bytedance/deer-flow --skill data-analysis -a codex`. Or copy the skill folder (skills/public/data-analysis in bytedance/deer-flow) into .agents/skills/data-analysis in your project. Codex loads it when a task matches its description.

Can I use Excel and CSV Data Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bytedance/deer-flow --skill data-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-analysis, .gemini/skills/data-analysis, .github/skills/data-analysis and .opencode/skills/data-analysis in your project.

What does Excel and CSV Data Analysis need to run?

Going by SKILL.md and its folder, Excel and CSV Data Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python with DuckDB available to `scripts/analyze.py`; Excel or CSV files in the uploads folder.

Does Excel and CSV Data Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Excel and CSV Data Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Excel and CSV Data Analysis use?

Excel and CSV Data Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Excel and CSV Data Analysis use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Excel and CSV Data Analysis?

Skills that share tags, products or a category with Excel and CSV Data Analysis: Data Analysis (HezaoHezao/poirot, 250 stars), Data Analysis (spytensor/openmozi, 439 stars), Raccoon Dataanalysis (SenseTime-Copilot/raccoon-dataanalysis-skill, 137 stars) and CSV Data Analysis (5zjk5/prompt-engineering, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Excel and CSV Data Analysis?

bytedance (a GitHub organization) maintains it in bytedance/deer-flow, which has 83,441 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 7, 2026.

Source: bytedance/deer-flow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.