Agent skill

Doc Parse

by anbeime in anbeime/skill

将 PDF/PPT/Excel/Word 等多格式文档解析为结构化 Markdown,并输出元数据与解析置信度,作为 RAG 与四色卡片的数据底座。

No licenceAuto-check passedDocuments & Office

Install Doc Parse

skills CLI
$ npx skills add anbeime/skill --skill doc-parse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install anbeime/skill doc-parse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/anbeime/skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/antinet-agentteams/skills/doc-parse .claude/skills/doc-parse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doc-parse
GitHub stars
7.6k
Token cost
~294 tokens
SKILL.md length
59 words
Files
2 (incl. scripts)
Skills in repo
57
Repo updated
First seen
Licence
None found

At a glance

将 PDF/PPT/Excel/Word 等多格式文档解析为结构化 Markdown,并输出元数据与解析置信度,作为 RAG 与四色卡片的数据底座。

  • Tasks that involve Excel spreadsheets
  • SKILL.md covers 使用方式, 输入(Input), 输出(Output) and 依赖(Dependencies), plus 3 more sections
  • Runs Python scripts from its folder; calls python
  • Tasks that involve Retrieval-augmented generation

What it does

Doc Parse is an agent skill from anbeime/skill. 将 PDF/PPT/Excel/Word 等多格式文档解析为结构化 Markdown,并输出元数据与解析置信度,作为 RAG 与四色卡片的数据底座。

Its SKILL.md is about 290 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/run_doc_parse.py`).

It sits in Documents & Office, covering Excel spreadsheets and Retrieval-augmented generation. It works with Microsoft Excel and Python. The repository describes itself as: 收录最全、更新最快的技能Skills商店:精选原创技能包(涵盖文档处理、内容创作、编程开发、机器学习、自动化工作流),全部打包好可直接安装使用!同时自动抓取GitHub上万个Skills项目,按分类、更新时间、Star数量整理。The most comprehensive and frequently updated AI Agent skill…

When your agent uses it

  • Tasks that involve Excel spreadsheets
  • Tasks that involve Retrieval-augmented generation

Example prompts

  • “/doc-parse”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit b8dadc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Doc Parse loads about 294 tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 59 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~294

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 59 words (~294 tokens).

name
doc-parse
assign_when
该 Worker 负责把任意格式的原始文档转成机器可读的结构化文本,是下游信息抽取与检索的通用解析入口。

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (scripts) in antinet-agentteams/skills/doc-parse of anbeime/skill.

  • SKILL.md
  • scripts/run_doc_parse.py

Open the folder on GitHubat commit b8dadc0

Compare with similar skills

Doc Parse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doc Parse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doc Parse this skillanbeime/skill7.6k—~294Automated safety check: PassNone
Markitdownaipoch/medical-research-skills2k—~1.3kAutomated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Doc Cleanernotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
XLSXzzhonglei/GeoCode-Release186—~3.1kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    aipoch/medical-research-skills

    Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation).

    2k GitHub stars~1.3k tokensUpdated 20 days ago
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Doc Cleaner

    notoriouslab/doc-cleaner

    Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

    309 GitHub stars~712 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • XLSX

    zzhonglei/GeoCode-Release

    Create, edit, analyze, or convert Excel spreadsheets (.xlsx, .xlsm) where the workbook file is the primary deliverable.

    186 GitHub stars~3.1k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • Python Bridge

    tmustier/pi-for-excel

    Native Python execution via the local Python bridge. An agent skill from tmustier/pi-for-excel.

    434 GitHub stars~820 tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from anbeime/skill

All 57 skills in this repo
  • Chinese-language skill that produces a 25-second vertical video of a digital shopping-guide avatar for e-commerce, chaining AI image, voice and video generation.

    7.6k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Produces a full material pack for a three-minute history-of-science explainer video from a theme, era and core conclusion: narration, storyboard, Veo2 prompts and character design.

    7.6k GitHub stars~480 tokensUpdated yesterday
    Auto-check passed
  • Turns a PPT content outline into styled slide images, an interactive HTML viewer and an optional stitched video from a content-planning partner skill.

    7.6k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check: notes
  • Builds or improves a PowerPoint deck through a seven-role workflow: topic analysis, template choice, content planning, writing, image suggestions, editing and PPTX generation.

    7.6k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Turns a document into a narrated roadshow video through ten staged roles, from document analysis and slide planning to audio, subtitles and final composition.

    7.6k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check: notes
  • PPTX Generator

    anbeime/skill

    产品经理与运营专员在准备汇报展示时,将 JSON 数据一键转换为标准 .pptx 文件。自动应用多种布局、图表与表格样式,快速生成专业 PPT,彻底告别繁琐手工排版,让内容产出与展示更高效!

    7.6k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Doc Parse

What does Doc Parse do?

将 PDF/PPT/Excel/Word 等多格式文档解析为结构化 Markdown,并输出元数据与解析置信度,作为 RAG 与四色卡片的数据底座。. Doc Parse is an agent skill from anbeime/skill.

When should I use Doc Parse?

Doc Parse fits situations like: tasks that involve Excel spreadsheets; tasks that involve Retrieval-augmented generation.

How do I install Doc Parse in Claude Code?

Run `npx skills add anbeime/skill --skill doc-parse -a claude-code`. Or copy the skill folder (antinet-agentteams/skills/doc-parse in anbeime/skill) into .claude/skills/doc-parse in your project. Claude Code loads it when a task matches its description.

How do I install Doc Parse in Codex?

Run `npx skills add anbeime/skill --skill doc-parse -a codex`. Or copy the skill folder (antinet-agentteams/skills/doc-parse in anbeime/skill) into .agents/skills/doc-parse in your project. Codex loads it when a task matches its description.

Can I use Doc Parse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anbeime/skill --skill doc-parse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doc-parse, .gemini/skills/doc-parse, .github/skills/doc-parse and .opencode/skills/doc-parse in your project.

What does Doc Parse need to run?

Going by SKILL.md and its folder, Doc Parse needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Doc Parse access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Doc Parse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Doc Parse use?

No licence was found for Doc Parse or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Doc Parse use?

About 294 tokens (SKILL.md is roughly 1.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doc Parse?

Skills that share tags, products or a category with Doc Parse: Markitdown (aipoch/medical-research-skills, 2k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Doc Cleaner (notoriouslab/doc-cleaner, 309 stars) and Mineru (Nebutra/MinerU-Skill, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doc Parse?

anbeime (a GitHub user) maintains it in anbeime/skill, which has 7,605 GitHub stars. The repository holds 57 skills in this directory. The repository was last updated on October 6, 2026.

Source: anbeime/skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.