Topic · Documents & Office
Best document parsing skills for Claude Code, Codex and other agents.
- skills
- 147
- official
- 9
Document parsing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR. | anthropics/ | 180k | 48 repos | ~2k | Automated safety check: Pass | Proprietary | 2 days ago |
| 2 | Gives the agent command-line and Python recipes for reading, creating, merging and splitting PDF files, plus tips for large and scanned documents. | shareAI-lab/ | 78k | 5 repos | ~646 | Automated safety check: Pass | MIT | 9 days ago |
| 3 | Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws. | ImCa0/ | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | 8 days ago |
| 4 | Converts books and documents in PDF, EPUB, DOCX, HTML, Markdown, text, RTF or MOBI form into agent skills built from frameworks, principles, techniques and anti-patterns. | virgiliojr94/ | 34k | — | ~14k | Automated safety check: Pass | MIT | 2 days ago |
| 5 | Converts PDFs, Office files, HTML, images and other documents into a unified DoclingDocument with Markdown or JSON output, through the docling CLI, Python SDK or a remote service. | docling-project/ | 68k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 6 | Reads, OCRs, searches and cites local documents through the mineru CLI, covering PDF, images, Office files, EPUB, HTML and CSV. | opendatalab/ | 81k | — | ~9.4k | Automated safety check: Warn | Unknown | yesterday |
| 7 | Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox. | HKUDS/ | 41k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 8 | Fetches web pages and PDFs and returns a source-grounded summary, clean Markdown, quotes or citations, routing each kind of link to a suitable fetch method. | tw93/ | 7.2k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 9 | Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data. | TokenRhythm/ | 7.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 10 | Rebuilds slide images, screenshots, scanned PDFs or image-only PPTX files as PowerPoint with editable objects, using a local CLI with per-page manifests and QA. | Yuan1z0825/ | 46k | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 11 | Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy. | NateBJones-Projects/ | 4.7k | — | ~995 | Automated safety check: Pass | Unknown | yesterday |
| 12 | 12.DOCX Toolkit Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task. | XiaomiMiMo/ | 14k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 13 | 处理 Office 文档的一站式 skill:Word(.doc/.docx/.dotx)、Excel(.xls/.xlsx/.xlsm/.csv)、PowerPoint(.ppt/.pptx/.potx) 的创建、读取、编辑、提取、转换、校验。触发:『读取 word 文档』『提取 excel 内容』『看 ppt 讲了什么』、.doc 老格式打不开、生成/编辑 Word… | xstongxue/ | 2.9k | — | ~1.8k | Automated safety check: Pass | Proprietary | 24 days ago |
| 14 | Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates. | alchaincyf/ | 908 | — | ~4.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 15 | Fetches the content of Feishu cloud documents as Markdown, with steps for downloading embedded images, files and whiteboards and for resolving wiki links. | op7418/ | 6.5k | 1 repo | ~554 | Automated safety check: Pass | Unknown | 15 days ago |
| 16 | Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite. | joeseesun/ | 508 | — | ~1.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 17 | Bootstraps and repairs the model-free AutoRAG Lite MCP server: installing it, initializing a config with approved search roots, building indexes and verifying discovery. | Marker-Inc-Korea/ | 5.1k | — | ~3.3k | Automated safety check: Pass | MIT | yesterday |
| 18 | 18.Markitdown Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. | jimmc414/ | 594 | 2 repos | ~1.7k | Automated safety check: Pass | No licence | 3 days ago |
| 19 | Quick reference for wdoc, a command-line and Python tool that summarizes, searches and answers questions over documents of many file types. | thiswillbeyourgithub/ | 545 | — | ~1.1k | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 20 | 20.PDF Toolkit Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling. | XiaomiMiMo/ | 14k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 21 | 21.Pullmd Read any web page, document, or YouTube video as clean Markdown using PullMD. | AeternaLabsHQ/ | 486 | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | 16 days ago |
| 22 | 22.Liteparse A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question… | bastani-inc/ | 846 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 23 | Builds a review grid with one row per document and one column per data point, each cell cited to a verbatim quote, built for M&A diligence and other batch reviews. | anthropics/ | 9.6k | 3 repos | ~4.3k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 24 | Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library. | LigphiDonk/ | 738 | — | ~1.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 25 | Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible. | pipeshub-ai/ | 3.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | Turns a URL, PDF, DOCX, Markdown file, text or screenshots into a designed, shareable single-file HTML article through a staged review workflow. | ConardLi/ | 13k | — | ~4.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 27 | Run a large-scale language migration with the six-step process: create the map and the rules, stress-test the rules, translate everything, compile, run it, match behavior. | anthropics/ | 742 | — | ~983 | Automated safety check: Pass | Unknown | 3 mo ago |
| 28 | 28.Mineru An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents. | Nebutra/ | 122 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 29 | Fetch any URL and convert to markdown using Chrome CDP. An agent skill from freestylefly/canghe-skills. | freestylefly/ | 461 | 4 repos | ~1.1k | Automated safety check: Pass | No licence | 4 mo ago |
| 30 | Converts a tender document into Markdown, extracts scoring criteria and requirements, then drafts an industry-formatted technical bid as a Word file. | Get00/ | 167 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 31 | 31.Lt2md Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text… | libnyx/ | 109 | — | ~4.4k | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 32 | Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool. | Team-Commonly/ | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | today |
| 33 | 33.Mineru An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents. | Nebutra/ | 122 | — | ~504 | Automated safety check: Pass | MIT | 13 days ago |
| 34 | Image-to-code replication pipeline. An agent skill from Yu-369/VibeCurb. | Yu-369/ | 980 | — | ~8.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 35 | 35.Markdrop Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions. | shoryasethia/ | 211 | — | ~1.4k | Automated safety check: Notes | GPL-3.0 | 1 mo ago |
| 36 | Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more. | agentscope-ai/ | 35k | — | ~1.8k | Automated safety check: Pass | Proprietary | 7 days ago |
| 37 | Reads screenshots of brokerage or portfolio transaction tables, checks the rows for duplicates and prints lines you can paste into your operations file by hand. | guilhermecgs/ | 178 | — | ~1.2k | Automated safety check: Pass | MPL-2.0 | 2 mo ago |
| 38 | 38.Lexoid CLI Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI. | oidlabs-com/ | 109 | — | ~2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 39 | Turns course slides, homework files, class code and earlier solutions into concise student-style Markdown answers, with optional subagents for solving and review. | vect-G/ | 127 | — | ~729 | Automated safety check: Pass | MIT | 5 mo ago |
| 40 | Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging. | swyxio/ | 172 | — | ~2.2k | Automated safety check: Pass | MIT | 2 days ago |
| 41 | Turn unstructured text into validated, structured variables at corpus scale with LLMs: codebook design, dev/gold-set construction, DSPy prompt optimization, a hard reliability gate (Cohen κ ≥ 0.70)… | joshzyj/ | 167 | — | ~4.6k | Automated safety check: Pass | Unknown | 19 days ago |
| 42 | 42.Final Review 用于期末复习、考前突击、题目生成、知识点辅导、复习资料清洗等场景。基于用户上传课程资料开展工作,优先参考历年真题、老师PPT、平时作业和 Markdown 资料;先用 markitdown 转 Markdown,再清洗材料,最后生成题目、答案和偏考试得分导向的解析。 | lucianwhy/ | 113 | — | ~1.3k | Automated safety check: Pass | No licence | 13 days ago |
| 43 | Picks the right Python library for PDF jobs, with examples for text and table extraction, merging, splitting and page extraction, plus OCR and memory fixes. | FareedKhan-dev/ | 298 | — | ~785 | Automated safety check: Pass | MIT | 6 mo ago |
| 44 | Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka. | Harryoung/ | 104 | — | ~505 | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 45 | Pulls architecture, method and result figures from an arXiv paper or PDF into an Obsidian vault and writes an index of them. | juliye2025/ | 1.7k | — | ~298 | Automated safety check: Pass | No licence | 22 days ago |
| 46 | Guidance for processing documents, extracting content, and transforming structured information. | vixues/ | 231 | — | ~560 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 47 | Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) inside a Python program using the lexoid library. | oidlabs-com/ | 109 | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 48 | Converts scanned faxes, images, CSV/TSV exports and C-CDA XML into clean text on-device, ready for OpenMed de-identification and named-entity recognition. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
Questions, answered from the data.
What is the best document parsing skill?
PDF Processing (official) from anthropics/skills ranks first of the 147 document parsing skills listed here, with the highest score: its repository has 180k GitHub stars, 48 other GitHub owners carry a copy, its SKILL.md loads about 2k tokens and it passes the automated safety check with no findings. Next come PDF Processing Guide and Markitdown.
Which document parsing skills are official?
9 of the 147 document parsing skills are official, published by the vendor's own GitHub organization: PDF Processing, Tabular Document Review, Code Migration, Excel to Markdown Converter, PDF to Markdown Converter and 4 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.