Codebook
brycewang-stanford/Auto-Empirical-Research-Skills
Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics.
Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.
$ npx skills add Harryoung/efka --skill excel-parser -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Harryoung/efka excel-parser --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/excel-parser .claude/skills/excel-parser && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "excel-parser" agent skill from https://github.com/Harryoung/efka/tree/main/skills/excel-parser into .claude/skills/excel-parser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "excel-parser", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Harryoung/efka/tree/main/skills/excel-parserType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Harryoung/efka --skill excel-parser -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Harryoung/efka excel-parser --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/excel-parser .agents/skills/excel-parser && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "excel-parser" agent skill from https://github.com/Harryoung/efka/tree/main/skills/excel-parser into .agents/skills/excel-parser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "excel-parser", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Harryoung/efka --skill excel-parser -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Harryoung/efka excel-parser --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/excel-parser .cursor/skills/excel-parser && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "excel-parser" agent skill from https://github.com/Harryoung/efka/tree/main/skills/excel-parser into .cursor/skills/excel-parser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "excel-parser", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Harryoung/efka.git --path skills/excel-parser--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Harryoung/efka --skill excel-parser -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Harryoung/efka excel-parser --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/excel-parser .gemini/skills/excel-parser && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "excel-parser" agent skill from https://github.com/Harryoung/efka/tree/main/skills/excel-parser into .gemini/skills/excel-parser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "excel-parser", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Harryoung/efka excel-parserInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Harryoung/efka --skill excel-parser -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/excel-parser .github/skills/excel-parser && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "excel-parser" agent skill from https://github.com/Harryoung/efka/tree/main/skills/excel-parser into .github/skills/excel-parser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "excel-parser", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Harryoung/efka --skill excel-parser -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Harryoung/efka excel-parser --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Harryoung/efka.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/excel-parser .opencode/skills/excel-parser && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "excel-parser" agent skill from https://github.com/Harryoung/efka/tree/main/skills/excel-parser into .opencode/skills/excel-parser/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "excel-parser", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
excel-parserSmart Excel/CSV file parsing with intelligent routing based on file complexity analysis.
Excel Parser is an agent skill from Harryoung/efka. Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis. Analyzes file structure (merged cells, row count, table layout) using lightweight metadata scanning, then recommends optimal processing strategy - either high-speed Pandas mode for standard tables or semantic HTML mode for complex reports. Use when processing Excel/CSV files with unknown or varying structure where optimization between speed and accuracy is needed.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/smart_excel_router.py` and `scripts/complexity_analyzer.py`).
It sits in Documents & Office, covering Excel spreadsheets, CSV and tabular files and DataFrames. It works with Microsoft Excel and pandas. The repository describes itself as: AI-powered knowledge management without vector embeddings. Built upon Claude Agent SDK, File system based, Agent driven. Maybe slower, but results are much more reliable! The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 9e6d3a4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Excel Parser loads about 2.3k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 117 tokens; SKILL.md has 1,038 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Harryoung/efka at commit 9e6d3a4, republished under its Apache-2.0 licence (© Harryoung). 1,038 words, ~2,295 tokens.
.claude/skills/excel-parser/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Provide intelligent routing strategies for parsing Excel/CSV files by analyzing complexity and choosing the optimal processing path. The skill implements a "Scout Pattern" that scans file metadata before processing to balance speed (Pandas) with accuracy (semantic extraction).
Before processing data, deploy a lightweight "scout" to analyze file metadata and make intelligent routing decisions:
openpyxl to scan file structure without loading dataKey Principle: "LLM handles metadata decisions, Pandas/HTML processes bulk data"
Use excel-parser when:
Skip this skill when:
Use the scripts/complexity_analyzer.py to scan file metadata:
python scripts/complexity_analyzer.py <file_path> [sheet_name]What it analyzes (without loading data):
Output (JSON format):
{
"is_complex": false,
"recommended_strategy": "pandas",
"reasons": ["No deep merges detected", "Row count exceeds 1000, forcing Pandas mode"],
"stats": {
"total_rows": 5000,
"deep_merges": 0,
"empty_interruptions": 0
}
}Based on complexity analysis:
Follow the selected path's workflow to extract data.
When: Simple/large tables (most common case)
Strategy: Agent analyzes ONLY the first 20 rows to determine header position, then use Pandas to read full data at native speed.
Workflow:
Sample First 20 Rows
pd.read_excel(..., nrows=20)Determine Header Position
Read Full Data
pd.read_excel(..., header=<detected_row>) to load complete dataToken Cost: ~500 tokens (only 20 rows analyzed) Processing Speed: Very fast (Pandas native speed)
For implementation details, see
references/smart_excel_router.py
When: Complex/irregular tables (merged cells, multi-level headers)
Strategy: Convert to semantic HTML preserving structure (rowspan/colspan), then extract data understanding the visual layout.
Workflow:
Convert to Semantic HTML
openpyxlrowspan and colspan attributes to maintain structureExtract Structured Data
Token Cost: Higher (full HTML structure analyzed) Processing Speed: Slower (semantic extraction) Use Case: Only for small (<1000 rows), complex files where Pandas would fail
For implementation details, see
references/smart_excel_router.py
Always run complexity analysis before processing. The metadata scan is fast (<1 second) and prevents wasted effort on wrong approach.
Never attempt HTML mode on files >1000 rows. Token limits will cause failures.
When in doubt, try Pandas mode first. It fails fast and clearly when structure is incompatible.
If processing multiple sheets from same file, run analysis once and cache results.
Never modify the original Excel file during analysis or processing.
FileNotFoundError or permission errorsBadZipFile or InvalidFileExceptionMemoryError or system slowdownread_only=True mode in openpyxlchunksize parameterUnicodeDecodeErrorpd.read_csv(..., encoding='gbk')Required Python packages:
openpyxl - Metadata scanning and Excel file manipulationpandas - High-speed data reading and manipulationThis skill includes:
scripts/complexity_analyzer.py - Standalone executable for complexity analysisreferences/smart_excel_router.py - Complete implementation reference with both processing paths© Harryoung, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts, references) in skills/excel-parser of Harryoung/efka.
Open the folder on GitHubat commit 9e6d3a4
Excel Parser next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Excel Parser this skillHarryoung/efka | 104 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Codebookbrycewang-stanford/Auto-Empirical-Research-Skills | 4.5k | — | ~527 | Automated safety check: Notes | Custom licence | |
| CSV and Excel MergerOneWave-AI/claude-skills | 322 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Convert Fileduckdb/duckdb-skills | 599 | 1 repos | ~720 | Automated safety check: Notes | MIT | |
| Sn Da Image CaptionMichaelYang-lyx/AIDABench | 111 | 1 repos | ~2k | Automated safety check: Pass | None | |
| Sn Da Excel WorkflowMichaelYang-lyx/AIDABench | 111 | 1 repos | ~2.5k | Automated safety check: Pass | None |
brycewang-stanford/Auto-Empirical-Research-Skills
Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics.
OneWave-AI/claude-skills
Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.
duckdb/duckdb-skills
Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more.
MichaelYang-lyx/AIDABench
图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py…
MichaelYang-lyx/AIDABench
Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…
Drchronx/ai-agent-research-starter-kit
Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.
Harryoung/efka
Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka.
Harryoung/efka
Send IM messages to users in batch. An agent skill from Harryoung/efka.
Harryoung/efka
Domain expert routing. An agent skill from Harryoung/efka.
Harryoung/efka
Generate table of contents overview for large files. An agent skill from Harryoung/efka.
Harryoung/efka
Handle user satisfaction feedback. An agent skill from Harryoung/efka.
Works with
Categories
Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis. Excel Parser is an agent skill from Harryoung/efka. Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.
Excel Parser fits situations like: processing Excel/CSV files with unknown; varying structure where optimization between speed and accuracy is needed.
Run `npx skills add Harryoung/efka --skill excel-parser -a claude-code`. Or copy the skill folder (skills/excel-parser in Harryoung/efka) into .claude/skills/excel-parser in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Harryoung/efka --skill excel-parser -a codex`. Or copy the skill folder (skills/excel-parser in Harryoung/efka) into .agents/skills/excel-parser in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Harryoung/efka --skill excel-parser -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/excel-parser, .gemini/skills/excel-parser, .github/skills/excel-parser and .opencode/skills/excel-parser in your project.
Going by SKILL.md and its folder, Excel Parser needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Excel Parser is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Excel Parser: Codebook (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), CSV and Excel Merger (OneWave-AI/claude-skills, 322 stars), Convert File (duckdb/duckdb-skills, 599 stars) and Sn Da Image Caption (MichaelYang-lyx/AIDABench, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Harryoung (a GitHub user) maintains it in Harryoung/efka, which has 104 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on March 16, 2026.
Source: Harryoung/efka on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.