Topic · Documents & Office
Best document parsing skills, page 3
Document parsing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | A skill your agent uses when the user needs a local file converted between common image, audio, video, document, or data formats. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 3 days ago |
| 98 | A skill your agent uses when working with Document Mind (DocMind) via Node.js SDK to submit document parsing jobs and poll results. | cinience/ | 397 | — | ~1.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 99 | A skill your agent uses when OCR-specialized extraction is needed with Alibaba Cloud Model Studio Qwen OCR models (qwen-vl-ocr, qwen-vl-ocr-latest, and snapshots), including document parsing, table… | cinience/ | 397 | — | ~953 | Automated safety check: Pass | MIT | 1 mo ago |
| 100 | 100.Aliyun Qwen Vl A skill your agent uses when understanding images with Alibaba Cloud Model Studio Qwen VL models (qwen3-vl-plus/qwen3-vl-flash and latest aliases). | cinience/ | 397 | — | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 101 | Condense a scholarly legal article by moving nonessential detail from the body into substantive footnotes while preserving every word and keeping the argument self-sufficient. | lawve-ai/ | 836 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 102 | Generate, fill, and assemble PDF documents at scale. An agent skill from curiositech/some_claude_skills. | curiositech/ | 243 | — | ~5.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 103 | 103.Tiptap Helps coding agents integrate and work with the Tiptap rich text editor. | fjrevoredo/ | 308 | 1 repo | ~857 | Automated safety check: Pass | MIT | today |
| 104 | Extract text and structure from scans, images, and scanned PDFs. | dkyazzentwatwa/ | 114 | — | ~310 | Automated safety check: Pass | No licence | 6 mo ago |
| 105 | 105.PDF To HTML Converts a PDF into one self-contained, readable HTML file that preserves images, tables, charts and reading order — optionally translating it into another language while keeping every figure. | daymade/ | 1.4k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 106 | 106.Query Files This skill SHOULD be used for structured extraction or batch queries against JSON (use jq), YAML (use yq), or Markdown (use mq), and for advanced text search with ripgrep flags or pipe composition… | kylesnowschwartz/ | 114 | — | ~1.2k | Automated safety check: Pass | No licence | 3 days ago |
| 107 | Create standalone beginner tutorial packages from a topic or supplied references, with adaptive research, course-style outline design, chapter visuals, and Markdown/DOCX/PDF/HTML exports. | yaojingang/ | 1.3k | — | ~698 | Automated safety check: Pass | MIT | 1 mo ago |
| 108 | 108.Compdf Toolkit All-in-one ComPDF workflow for document conversion, OCR, data extraction, PDF editing, protection, compression, and watermarking. | ComPDFKit/ | 109 | — | ~867 | Automated safety check: Pass | No licence | 1 mo ago |
| 109 | 109.Documents Create, inspect, edit, or convert document-style deliverables. | terrense/ | 121 | — | ~131 | Automated safety check: Pass | No licence | 3 mo ago |
| 110 | 110.PDF To Markdown Converts PDF files to Markdown using Microsoft's markitdown package. | pamelafox/ | 125 | — | ~372 | Automated safety check: Pass | MIT | 1 mo ago |
| 111 | 111.PDF Processor Process PDFs - extract text, tables, and structured data from documents | gooseworks-ai/ | 1.2k | 1 repo | ~1.1k | Automated safety check: Pass | MIT | today |
| 112 | 112.Markitdown Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation). | aipoch/ | 2k | — | ~1.3k | Automated safety check: Pass | MIT | 21 days ago |
| 113 | 113.API Assets Call nodetool.assets or nodetool.documents from a code action: find, read and save files in the user's asset library, keep a file another tool downloaded, and convert documents or extract text and… | nodetool-ai/ | 556 | — | ~908 | Automated safety check: Pass | AGPL-3.0 | today |
| 114 | 114.PDF Handle PDF manipulation, form filling, text/table extraction, and high-fidelity generation. | IgorWarzocha/ | 122 | — | ~495 | Automated safety check: Pass | No licence | 8 mo ago |
| 115 | Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC… | BlackBeltTechnology/ | 315 | — | ~999 | Automated safety check: Pass | MIT | today |
| 116 | 116.Liteparse Provides fast document to markdown extraction. An agent skill from sammcj/agentic-coding. | sammcj/ | 162 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 117 | 117.PDF Create new PDFs and handle existing .pdf files safely with bundled Node/JS tools, including text extraction, page rendering, invoice/document parsing, form filling, and overlays. | HybridAIOne/ | 158 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 118 | Execute an Adobe PDF Services operation through explicit asset custody, asynchronous status, output verification, and cleanup. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 119 | 119.Scholar Compute Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction… | joshzyj/ | 168 | — | ~15k | Automated safety check: Pass | Unknown | 19 days ago |
| 120 | 120.Paper To Skill Converts research papers into executable skill packages via document conversion, critical analysis, and co-evolutionary refinement. | Mathews-Tom/ | 328 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 121 | 121.To Markdown Convert any file or URL to clean Markdown: PDF, DOCX, XLSX, PPTX, HTML, images (OCR), audio, CSV, YouTube. | Mathews-Tom/ | 328 | — | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 122 | 122.Data Cleaning A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate… | ericrisco/ | 167 | — | ~3.6k | Automated safety check: Pass | MIT | today |
| 123 | 123.Data Scraper A skill your agent uses when data lives on a website with no usable API — listings, prices, public records — and the scrape must stay legal and not get blocked: legal gate, extraction path, durable… | ericrisco/ | 167 | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 124 | A skill your agent uses when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of… | ericrisco/ | 167 | — | ~2.4k | Automated safety check: Pass | MIT | today |
| 125 | A skill your agent uses when text must become a typed, schema-conformant object you can trust — pulling fields into a fixed JSON shape, extracting line items as typed records, classifying into… | ericrisco/ | 167 | — | ~3.6k | Automated safety check: Pass | MIT | today |
| 126 | 126.Markitdown Skill Convert documents to Markdown using Microsoft's MarkItDown CLI (markitdown). | infometa/ | 344 | — | ~841 | Automated safety check: Notes | No licence | today |
| 127 | 127.Cryptpad Use CryptPad encrypted collaboration for integration, automation, and diagnosis with its browser integration API, instance discovery, document import/export, and session-key lifecycle. | magnus919/ | 113 | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 128 | Use Databricks built-in AI Functions (aiclassify, aiextract, aisummarize, aimask, aitranslate, aifixgrammar, aigen, aianalyzesentiment, aisimilarity, aiparsedocument, aiprepsearch, aiquery… | databricks/ | 345 | — | ~3.9k | Automated safety check: Pass | Unknown | today |
| 129 | 使用 Nutrient DWS API 处理、转换、OCR、提取、脱敏、签名和填充文档。支持 PDF、DOCX、XLSX、PPTX、HTML 和图像。 | xu-xiang/ | 2k | — | ~1.3k | Automated safety check: Pass | MIT | 7 mo ago |
| 130 | 130.Cheerio Parsing Expert guidance for HTML/XML parsing using Cheerio in Node.js with best practices for DOM traversal, data extraction, and efficient scraping pipelines. | Kilo-Org/ | 190 | 1 repo | ~2.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 131 | 131.Duckdb En DuckDB CLI specialist for SQL analysis, data processing and file conversion. | aAAaqwq/ | 105 | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 10 days ago |
| 132 | 132.Teach Teaches the Second Brain to recognize a new external data source. | ccplugins/ | 968 | — | ~5.9k | Automated safety check: Notes | Apache-2.0 | 1 mo ago |
| 133 | 133.PDF Parser Configure and manage - Parse pdf parser operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~550 | Automated safety check: Pass | MIT | today |
| 134 | 134.Chemeagle Guide Multi-agent system for chemical literature information extraction | wentorai/ | 298 | 1 repo | ~1.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 135 | Mine open access full-text repositories for research data extraction | wentorai/ | 298 | 1 repo | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 136 | Build OCR pipelines in MATLAB using the ocr() function. An agent skill from matlab/matlab-agentic-toolkit. | matlab/ | 1.1k | — | ~5.2k | Automated safety check: Pass | Unknown | today |
| 137 | Multi-hop evidence search + structured extraction over enterprise artifact datasets (docs/chats/meetings/PRs/URLs). | benchflow-ai/ | 1.8k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 138 | A skill your agent uses when you need to parse Word/HTML/JSON/Markdown/Excel requirements and produce a structured analysis; triggers include requirements analysis plus and requirement document… | naodeng/ | 245 | — | ~835 | Automated safety check: Pass | Unknown | 12 days ago |
| 139 | Basic Chinese NLP and text analysis workflows for PDF text/table extraction, jieba tokenization, word and sentence frequency, word clouds, TF-IDF, Word2Vec, sentence embeddings, and text similarity. | Drchronx/ | 135 | — | ~737 | Automated safety check: Pass | Unknown | 4 mo ago |
| 140 | 140.Pp Context Dev Printing Press CLI for Context.dev. An agent skill from mvanhorn/printing-press-library. | mvanhorn/ | 2.1k | — | ~4.3k | Automated safety check: Notes | Apache-2.0 | today |
| 141 | 141.Mineru PDF Parse PDFs locally (CPU) into Markdown/JSON using MinerU. An agent skill from sundial-org/awesome-openclaw-skills. | sundial-org/ | 663 | — | ~261 | Automated safety check: Pass | No licence | 7 mo ago |
| 142 | Use ManagedCode.MarkItDown when a .NET application needs deterministic document-to-Markdown conversion for ingestion, indexing, summarization, or content-processing workflows. | managedcode/ | 486 | — | ~550 | Automated safety check: Pass | MIT | today |
| 143 | Converts documents and URLs to markdown via tiered fallback (MCP markitdown, native tools, user notice). | athola/ | 342 | — | ~1.3k | Automated safety check: Pass | MIT | 2 days ago |
| 144 | 144.Auto Paper Skill Find, deduplicate, save, index, and analyze research papers through Codex conversation. | franklee16/ | 223 | — | ~2k | Automated safety check: Pass | No licence | 19 days ago |