Search
Document parsing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR. | anthropics/ | 180k | 48 repos | ~2k | Automated safety check: Pass | Proprietary | 3 days ago |
| 2 | Gives the agent command-line and Python recipes for reading, creating, merging and splitting PDF files, plus tips for large and scanned documents. | shareAI-lab/ | 78k | 5 repos | ~646 | Automated safety check: Pass | MIT | 10 days ago |
| 3 | Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws. | ImCa0/ | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | yesterday |
| 4 | Converts books and documents in PDF, EPUB, DOCX, HTML, Markdown, text, RTF or MOBI form into agent skills built from frameworks, principles, techniques and anti-patterns. | virgiliojr94/ | 34k | — | ~14k | Automated safety check: Pass | MIT | 3 days ago |
| 5 | Converts PDFs, Office files, HTML, images and other documents into a unified DoclingDocument with Markdown or JSON output, through the docling CLI, Python SDK or a remote service. | docling-project/ | 69k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 6 | Rebuilds slide images, screenshots, scanned PDFs or image-only PPTX files as PowerPoint with editable objects, using a local CLI with per-page manifests and QA. | Yuan1z0825/ | 46k | 1 repo | ~2.5k | Automated safety check: Pass | MIT | 2 days ago |
| 7 | Reads, OCRs, searches and cites local documents through the mineru CLI, covering PDF, images, Office files, EPUB, HTML and CSV. | opendatalab/ | 81k | — | ~9.4k | Automated safety check: Warn | Unknown | 2 days ago |
| 8 | Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox. | HKUDS/ | 41k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Fetches web pages and PDFs and returns a source-grounded summary, clean Markdown, quotes or citations, routing each kind of link to a suitable fetch method. | tw93/ | 7.2k | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 10 | 10.PDF Toolkit Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data. | TokenRhythm/ | 7.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 11 | Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy. | NateBJones-Projects/ | 4.7k | — | ~995 | Automated safety check: Pass | Unknown | yesterday |
| 12 | 12.DOCX Toolkit Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task. | XiaomiMiMo/ | 14k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 13 | 处理 Office 文档的一站式 skill:Word(.doc/.docx/.dotx)、Excel(.xls/.xlsx/.xlsm/.csv)、PowerPoint(.ppt/.pptx/.potx) 的创建、读取、编辑、提取、转换、校验。触发:『读取 word 文档』『提取 excel 内容』『看 ppt 讲了什么』、.doc 老格式打不开、生成/编辑 Word… | xstongxue/ | 3k | — | ~1.8k | Automated safety check: Pass | Proprietary | 25 days ago |
| 14 | Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates. | alchaincyf/ | 908 | — | ~4.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 15 | Fetches the content of Feishu cloud documents as Markdown, with steps for downloading embedded images, files and whiteboards and for resolving wiki links. | op7418/ | 6.5k | 1 repo | ~554 | Automated safety check: Pass | Unknown | 16 days ago |
| 16 | Converts a URL or PDF into clean Markdown, with dedicated routes for WeChat articles, Feishu docs, arXiv papers and login-gated pages, before any summary or rewrite. | joeseesun/ | 509 | — | ~1.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 17 | Bootstraps and repairs the model-free AutoRAG Lite MCP server: installing it, initializing a config with approved search roots, building indexes and verifying discovery. | Marker-Inc-Korea/ | 5.1k | — | ~3.3k | Automated safety check: Pass | MIT | yesterday |
| 18 | 18.Markitdown Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. | jimmc414/ | 594 | 2 repos | ~1.7k | Automated safety check: Pass | No licence | 4 days ago |
| 19 | Quick reference for wdoc, a command-line and Python tool that summarizes, searches and answers questions over documents of many file types. | thiswillbeyourgithub/ | 545 | — | ~1.1k | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 20 | 20.PDF Toolkit Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling. | XiaomiMiMo/ | 14k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 21 | 21.Pullmd Read any web page, document, or YouTube video as clean Markdown using PullMD. | AeternaLabsHQ/ | 486 | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | 17 days ago |
| 22 | Builds a review grid with one row per document and one column per data point, each cell cited to a verbatim quote, built for M&A diligence and other batch reviews. | anthropics/ | 9.6k | 3 repos | ~4.3k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 23 | Fetch any URL and convert to markdown using Chrome CDP. An agent skill from freestylefly/canghe-skills. | freestylefly/ | 461 | 5 repos | ~1.1k | Automated safety check: Pass | No licence | 4 mo ago |
| 24 | 24.Liteparse A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question… | bastani-inc/ | 846 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 25 | Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library. | LigphiDonk/ | 738 | — | ~1.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 26 | Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible. | pipeshub-ai/ | 3.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 27 | Turns a URL, PDF, DOCX, Markdown file, text or screenshots into a designed, shareable single-file HTML article through a staged review workflow. | ConardLi/ | 13k | — | ~4.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 28 | Run a large-scale language migration with the six-step process: create the map and the rules, stress-test the rules, translate everything, compile, run it, match behavior. | anthropics/ | 743 | — | ~983 | Automated safety check: Pass | Unknown | 3 mo ago |
| 29 | 29.Mineru An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents. | Nebutra/ | 122 | — | ~1.4k | Automated safety check: Pass | MIT | 14 days ago |
| 30 | Converts a tender document into Markdown, extracts scoring criteria and requirements, then drafts an industry-formatted technical bid as a Word file. | Get00/ | 167 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 31 | 31.Lt2md Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text… | libnyx/ | 109 | — | ~4.4k | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 32 | Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool. | Team-Commonly/ | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 33 | 33.Mineru An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents. | Nebutra/ | 122 | — | ~504 | Automated safety check: Pass | MIT | 14 days ago |
| 34 | Image-to-code replication pipeline. An agent skill from Yu-369/VibeCurb. | Yu-369/ | 979 | — | ~8.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 35 | 35.Markdrop Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions. | shoryasethia/ | 211 | — | ~1.4k | Automated safety check: Notes | GPL-3.0 | 2 mo ago |
| 36 | Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more. | agentscope-ai/ | 35k | — | ~1.8k | Automated safety check: Pass | Proprietary | today |
| 37 | Reads screenshots of brokerage or portfolio transaction tables, checks the rows for duplicates and prints lines you can paste into your operations file by hand. | guilhermecgs/ | 179 | — | ~1.2k | Automated safety check: Pass | MPL-2.0 | 3 mo ago |
| 38 | 38.Lexoid CLI Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI. | oidlabs-com/ | 109 | — | ~2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 39 | Turns course slides, homework files, class code and earlier solutions into concise student-style Markdown answers, with optional subagents for solving and review. | vect-G/ | 127 | — | ~729 | Automated safety check: Pass | MIT | 5 mo ago |
| 40 | Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging. | swyxio/ | 175 | — | ~2.2k | Automated safety check: Pass | MIT | 3 days ago |
| 41 | Turn unstructured text into validated, structured variables at corpus scale with LLMs: codebook design, dev/gold-set construction, DSPy prompt optimization, a hard reliability gate (Cohen κ ≥ 0.70)… | joshzyj/ | 168 | — | ~4.6k | Automated safety check: Pass | Unknown | 20 days ago |
| 42 | 42.Final Review 用于期末复习、考前突击、题目生成、知识点辅导、复习资料清洗等场景。基于用户上传课程资料开展工作,优先参考历年真题、老师PPT、平时作业和 Markdown 资料;先用 markitdown 转 Markdown,再清洗材料,最后生成题目、答案和偏考试得分导向的解析。 | lucianwhy/ | 113 | — | ~1.3k | Automated safety check: Pass | No licence | 14 days ago |
| 43 | Picks the right Python library for PDF jobs, with examples for text and table extraction, merging, splitting and page extraction, plus OCR and memory fixes. | FareedKhan-dev/ | 298 | — | ~785 | Automated safety check: Pass | MIT | 6 mo ago |
| 44 | Convert DOC/DOCX/PDF/PPT/PPTX documents to Markdown format. An agent skill from Harryoung/efka. | Harryoung/ | 104 | — | ~505 | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 45 | Pulls architecture, method and result figures from an arXiv paper or PDF into an Obsidian vault and writes an index of them. | juliye2025/ | 1.7k | — | ~298 | Automated safety check: Pass | No licence | 23 days ago |
| 46 | Guidance for processing documents, extracting content, and transforming structured information. | vixues/ | 232 | — | ~560 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 47 | Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) inside a Python program using the lexoid library. | oidlabs-com/ | 109 | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 48 | Converts scanned faxes, images, CSV/TSV exports and C-CDA XML into clean text on-device, ready for OpenMed de-identification and named-entity recognition. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |