Agent skill

Excel Parser

by Harryoung in Harryoung/efka

Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

Apache-2.0Auto-check passedDocuments & Office

Install Excel Parser

skills CLI
$ npx skills add Harryoung/efka --skill excel-parser -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Harryoung/efka excel-parser --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/excel-parser .claude/skills/excel-parser && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
excel-parser
GitHub stars
104
Token cost
~2.3k tokens
SKILL.md length
1,038 words
Files
3 (incl. scripts, references)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

  • Works in 8 steps: Analyze File Complexity → Route to Optimal Strategy → Execute Processing → …
  • Processing Excel/CSV files with unknown
  • SKILL.md covers Table of Contents, Overview, Core Philosophy: Scout Pattern and When to Use This Skill, plus 7 more sections
  • Runs Python scripts from its folder; calls python

What it does

Excel Parser is an agent skill from Harryoung/efka. Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis. Analyzes file structure (merged cells, row count, table layout) using lightweight metadata scanning, then recommends optimal processing strategy - either high-speed Pandas mode for standard tables or semantic HTML mode for complex reports. Use when processing Excel/CSV files with unknown or varying structure where optimization between speed and accuracy is needed.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/smart_excel_router.py` and `scripts/complexity_analyzer.py`).

It sits in Documents & Office, covering Excel spreadsheets, CSV and tabular files and DataFrames. It works with Microsoft Excel and pandas. The repository describes itself as: AI-powered knowledge management without vector embeddings. Built upon Claude Agent SDK, File system based, Agent driven. Maybe slower, but results are much more reliable! The licence is Apache-2.0.

When your agent uses it

  • Processing Excel/CSV files with unknown
  • Varying structure where optimization between speed and accuracy is needed

Example prompts

  • “/excel-parser”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Analyze File Complexity
  2. Route to Optimal Strategy
  3. Execute Processing
  4. Trust the Scout
  5. Respect the Row Count Rule
  6. Pandas First for Unknown Files
  7. Cache Analysis Results
  8. Preserve Original Files

What it can do on your machine

Read from SKILL.md and the folder at commit 9e6d3a4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Excel Parser loads about 2.3k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 117 tokens; SKILL.md has 1,038 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Harryoung/efka at commit 9e6d3a4, republished under its Apache-2.0 licence (© Harryoung). 1,038 words, ~2,295 tokens.

Download SKILL.mdSave it as .claude/skills/excel-parser/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
excel-parser
description
Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis. Analyzes file structure (merged cells, row count, table layout) using lightweight metadata scanning, then recommends optimal processing strategy - either high-speed Pandas mode for standard tables or semantic HTML mode for complex reports. Use when processing Excel/CSV files with unknown or varying structure where optimization between speed and accuracy is needed.

Excel Parser

Table of Contents

Overview

Provide intelligent routing strategies for parsing Excel/CSV files by analyzing complexity and choosing the optimal processing path. The skill implements a "Scout Pattern" that scans file metadata before processing to balance speed (Pandas) with accuracy (semantic extraction).

Core Philosophy: Scout Pattern

Before processing data, deploy a lightweight "scout" to analyze file metadata and make intelligent routing decisions:

  1. Metadata Scanning - Use openpyxl to scan file structure without loading data
  2. Complexity Scoring - Calculate score based on merged cells, row count, and layout
  3. Path Selection - Choose between Pandas (fast) or HTML (accurate) processing
  4. Optimized Execution - Execute with the most appropriate tool for the file type

Key Principle: "LLM handles metadata decisions, Pandas/HTML processes bulk data"

When to Use This Skill

Use excel-parser when:

  • Processing Excel/CSV files with unknown structure or varying complexity
  • Handling files ranging from simple data tables to complex financial reports
  • Need to optimize between processing speed and extraction accuracy
  • Working with files that may contain merged cells, multi-level headers, or irregular layouts

Skip this skill when:

  • File structure is already known and documented
  • Processing simple, well-structured tables with confirmed format
  • Using predefined scripts for specific file formats

Processing Workflow

Step 1: Analyze File Complexity

Use the scripts/complexity_analyzer.py to scan file metadata:

bash
python scripts/complexity_analyzer.py <file_path> [sheet_name]

What it analyzes (without loading data):

  • Merged cell distribution (shallow vs deep in the table)
  • Row count and data continuity
  • Empty row interruptions (indicates multi-table layouts)

Output (JSON format):

json
{
  "is_complex": false,
  "recommended_strategy": "pandas",
  "reasons": ["No deep merges detected", "Row count exceeds 1000, forcing Pandas mode"],
  "stats": {
    "total_rows": 5000,
    "deep_merges": 0,
    "empty_interruptions": 0
  }
}
Step 2: Route to Optimal Strategy

Based on complexity analysis:

  • is_complex = false → Use Path A (Pandas Standard Mode)
  • is_complex = true → Use Path B (HTML Semantic Mode)
Step 3: Execute Processing

Follow the selected path's workflow to extract data.

Complexity Scoring Rules

Rule 1: Deep Merged Cells
  • Condition: Merged cells appearing beyond row 5
  • Interpretation: Complex table structure (not just header formatting)
  • Decision: Mark as complex if >2 deep merges detected
  • Example: Financial reports with merged category labels in data region
Rule 2: Empty Row Interruptions
  • Condition: Multiple empty rows within the table
  • Interpretation: Multiple sub-tables in single sheet
  • Decision: Mark as complex if >2 empty row interruptions found
  • Example: Summary table + detail table in one sheet
Rule 3: Row Count Override
  • Condition: Total rows >1000
  • Interpretation: Too large for HTML processing (token explosion)
  • Decision: Force Pandas mode regardless of complexity
  • Rationale: HTML conversion would exceed token limits
Rule 4: Default (Standard Table)
  • Condition: No deep merges, continuous data, moderate size
  • Interpretation: Standard data table
  • Decision: Use Pandas for optimal speed

Path A: Pandas Standard Mode

When: Simple/large tables (most common case)

Strategy: Agent analyzes ONLY the first 20 rows to determine header position, then use Pandas to read full data at native speed.

Workflow:

  1. Sample First 20 Rows

    • Read only the first 20 rows using pd.read_excel(..., nrows=20)
    • Convert to CSV format for analysis
  2. Determine Header Position

    • Examine the sampled rows to identify which row contains the actual column headers
    • Common patterns: Row 0 (standard), Row 1-2 (if title rows exist), Row with distinct column names
  3. Read Full Data

    • Use pd.read_excel(..., header=<detected_row>) to load complete data
    • The header parameter ensures proper column naming

Token Cost: ~500 tokens (only 20 rows analyzed) Processing Speed: Very fast (Pandas native speed)

For implementation details, see references/smart_excel_router.py

Path B: HTML Semantic Mode

When: Complex/irregular tables (merged cells, multi-level headers)

Strategy: Convert to semantic HTML preserving structure (rowspan/colspan), then extract data understanding the visual layout.

Workflow:

  1. Convert to Semantic HTML

    • Load workbook with openpyxl
    • Build HTML table preserving merged cell spans
    • Use rowspan and colspan attributes to maintain structure
  2. Extract Structured Data

    • Analyze HTML table structure
    • Identify hierarchical headers from merged cells
    • Extract data preserving semantic relationships

Token Cost: Higher (full HTML structure analyzed) Processing Speed: Slower (semantic extraction) Use Case: Only for small (<1000 rows), complex files where Pandas would fail

For implementation details, see references/smart_excel_router.py

Show full SKILL.md (372 more words)Show less

Best Practices

1. Trust the Scout

Always run complexity analysis before processing. The metadata scan is fast (<1 second) and prevents wasted effort on wrong approach.

2. Respect the Row Count Rule

Never attempt HTML mode on files >1000 rows. Token limits will cause failures.

3. Pandas First for Unknown Files

When in doubt, try Pandas mode first. It fails fast and clearly when structure is incompatible.

4. Cache Analysis Results

If processing multiple sheets from same file, run analysis once and cache results.

5. Preserve Original Files

Never modify the original Excel file during analysis or processing.

Troubleshooting

File Cannot Be Opened
  • Symptom: FileNotFoundError or permission errors
  • Causes: Invalid path, file locked by another process, insufficient permissions
  • Solutions:
    • Verify file path is correct and file exists
    • Close the file if open in Excel or another application
    • Check read permissions on the file
Corrupted File Errors
  • Symptom: BadZipFile or InvalidFileException
  • Causes: Incomplete download, file corruption, wrong file extension
  • Solutions:
    • Re-download or obtain fresh copy of the file
    • Verify file is actual Excel format (not CSV with .xlsx extension)
    • Try opening in Excel to confirm file integrity
Memory Issues with Large Files
  • Symptom: MemoryError or system slowdown
  • Causes: File too large for available RAM
  • Solutions:
    • Use read_only=True mode in openpyxl
    • Process file in chunks using Pandas chunksize parameter
    • Increase system memory or use machine with more RAM
Encoding Problems
  • Symptom: Garbled text or UnicodeDecodeError
  • Causes: Non-UTF8 encoding in source data
  • Solutions:
    • Specify encoding when reading CSV: pd.read_csv(..., encoding='gbk')
    • For Excel, data is usually UTF-8; check source data generation
HTML Mode Token Overflow
  • Symptom: Truncated output or API errors
  • Causes: Complex file exceeds token limits despite row count check
  • Solutions:
    • Force Pandas mode even for complex files
    • Split sheet into smaller ranges and process separately
    • Extract only essential columns before HTML conversion
Incorrect Header Detection
  • Symptom: Wrong columns or data shifted
  • Causes: Unusual header patterns not caught by sampling
  • Solutions:
    • Manually specify header row if known
    • Increase sample size beyond 20 rows
    • Use HTML mode for better structure understanding

Dependencies

Required Python packages:

  • openpyxl - Metadata scanning and Excel file manipulation
  • pandas - High-speed data reading and manipulation

Resources

This skill includes:

  • scripts/complexity_analyzer.py - Standalone executable for complexity analysis
  • references/smart_excel_router.py - Complete implementation reference with both processing paths

© Harryoung, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/excel-parser of Harryoung/efka.

  • SKILL.md
  • references/smart_excel_router.py
  • scripts/complexity_analyzer.py

Open the folder on GitHubat commit 9e6d3a4

Compare with similar skills

Excel Parser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Excel Parser compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Excel Parser this skillHarryoung/efka104—~2.3kAutomated safety check: PassApache-2.0
Codebookbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~527Automated safety check: NotesCustom licence
CSV and Excel MergerOneWave-AI/claude-skills322—~1.6kAutomated safety check: PassMIT
Convert Fileduckdb/duckdb-skills5991 repos~720Automated safety check: NotesMIT
Sn Da Image CaptionMichaelYang-lyx/AIDABench1111 repos~2kAutomated safety check: PassNone
Sn Da Excel WorkflowMichaelYang-lyx/AIDABench1111 repos~2.5kAutomated safety check: PassNone

Similar skills

  • Codebook

    brycewang-stanford/Auto-Empirical-Research-Skills

    Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics.

    4.5k GitHub stars~527 tokensUpdated 2 days ago
    Documents & OfficeAuto-check: notes
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    322 GitHub stars~1.6k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Convert File

    duckdb/duckdb-skills

    Official

    Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more.

    599 GitHub starsUsed in 1 repo~720 tokens
    Documents & OfficeAuto-check: notes
  • Sn Da Image Caption

    MichaelYang-lyx/AIDABench

    图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py…

    111 GitHub starsUsed in 1 repo~2k tokens
    Documents & OfficeAuto-check passed
  • Sn Da Excel Workflow

    MichaelYang-lyx/AIDABench

    Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…

    111 GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed
  • Multi Source Data Integration Extraction

    Drchronx/ai-agent-research-starter-kit

    Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

    134 GitHub stars~671 tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed

More from Harryoung/efka

  • Document Conversion

    Harryoung/efka

    Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka.

    104 GitHub stars~505 tokensUpdated 6 mo ago
    Auto-check passed
  • Batch Notification

    Harryoung/efka

    Send IM messages to users in batch. An agent skill from Harryoung/efka.

    104 GitHub stars~512 tokensUpdated 6 mo ago
    Auto-check passed
  • Expert Routing

    Harryoung/efka

    Domain expert routing. An agent skill from Harryoung/efka.

    104 GitHub stars~325 tokensUpdated 6 mo ago
    Auto-check passed
  • Large File Toc

    Harryoung/efka

    Generate table of contents overview for large files. An agent skill from Harryoung/efka.

    104 GitHub stars~293 tokensUpdated 6 mo ago
    Auto-check passed
  • Satisfaction Feedback

    Harryoung/efka

    Handle user satisfaction feedback. An agent skill from Harryoung/efka.

    104 GitHub stars~346 tokensUpdated 6 mo ago
    Auto-check passed

Questions about Excel Parser

What does Excel Parser do?

Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis. Excel Parser is an agent skill from Harryoung/efka. Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

When should I use Excel Parser?

Excel Parser fits situations like: processing Excel/CSV files with unknown; varying structure where optimization between speed and accuracy is needed.

How do I install Excel Parser in Claude Code?

Run `npx skills add Harryoung/efka --skill excel-parser -a claude-code`. Or copy the skill folder (skills/excel-parser in Harryoung/efka) into .claude/skills/excel-parser in your project. Claude Code loads it when a task matches its description.

How do I install Excel Parser in Codex?

Run `npx skills add Harryoung/efka --skill excel-parser -a codex`. Or copy the skill folder (skills/excel-parser in Harryoung/efka) into .agents/skills/excel-parser in your project. Codex loads it when a task matches its description.

Can I use Excel Parser in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Harryoung/efka --skill excel-parser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/excel-parser, .gemini/skills/excel-parser, .github/skills/excel-parser and .opencode/skills/excel-parser in your project.

What does Excel Parser need to run?

Going by SKILL.md and its folder, Excel Parser needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Excel Parser access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Excel Parser safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Excel Parser use?

Excel Parser is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Excel Parser use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.

What are the alternatives to Excel Parser?

Skills that share tags, products or a category with Excel Parser: Codebook (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), CSV and Excel Merger (OneWave-AI/claude-skills, 322 stars), Convert File (duckdb/duckdb-skills, 599 stars) and Sn Da Image Caption (MichaelYang-lyx/AIDABench, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Excel Parser?

Harryoung (a GitHub user) maintains it in Harryoung/efka, which has 104 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on March 16, 2026.

Source: Harryoung/efka on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.