Search
Python · Document parsing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR. | anthropics/ | 180k | 48 repos | ~2k | Automated safety check: Pass | Proprietary | today |
| 2 | Gives the agent command-line and Python recipes for reading, creating, merging and splitting PDF files, plus tips for large and scanned documents. | shareAI-lab/ | 78k | 5 repos | ~646 | Automated safety check: Pass | MIT | 10 days ago |
| 3 | Converts PDFs, Office files, HTML, images and other documents into a unified DoclingDocument with Markdown or JSON output, through the docling CLI, Python SDK or a remote service. | docling-project/ | 69k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 4 | Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox. | HKUDS/ | 41k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data. | TokenRhythm/ | 7.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Rebuilds slide images, screenshots, scanned PDFs or image-only PPTX files as PowerPoint with editable objects, using a local CLI with per-page manifests and QA. | Yuan1z0825/ | 47k | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 7 | Quick reference for wdoc, a command-line and Python tool that summarizes, searches and answers questions over documents of many file types. | thiswillbeyourgithub/ | 545 | — | ~1.1k | Automated safety check: Pass | AGPL-3.0 | 1 mo ago |
| 8 | Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling. | XiaomiMiMo/ | 14k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library. | LigphiDonk/ | 738 | — | ~1.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 10 | Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible. | pipeshub-ai/ | 3.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Run a large-scale language migration with the six-step process: create the map and the rules, stress-test the rules, translate everything, compile, run it, match behavior. | anthropics/ | 746 | — | ~983 | Automated safety check: Pass | Unknown | 3 mo ago |
| 12 | 12.Mineru An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents. | Nebutra/ | 122 | — | ~1.4k | Automated safety check: Pass | MIT | 15 days ago |
| 13 | Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool. | Team-Commonly/ | 1.4k | — | ~557 | Automated safety check: Pass | Apache-2.0 | today |
| 14 | 14.Mineru An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents. | Nebutra/ | 122 | — | ~504 | Automated safety check: Pass | MIT | 15 days ago |
| 15 | 15.Markdrop Professional AI skill and usage instructions for the Markdrop package, a Python tool for converting PDFs to Markdown/HTML with AI-powered image/table descriptions. | shoryasethia/ | 211 | — | ~1.4k | Automated safety check: Notes | GPL-3.0 | 2 mo ago |
| 16 | Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more. | agentscope-ai/ | 36k | — | ~1.8k | Automated safety check: Pass | Proprietary | today |
| 17 | Reads screenshots of brokerage or portfolio transaction tables, checks the rows for duplicates and prints lines you can paste into your operations file by hand. | guilhermecgs/ | 179 | — | ~1.2k | Automated safety check: Pass | MPL-2.0 | 3 mo ago |
| 18 | 18.Lexoid CLI Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI. | oidlabs-com/ | 109 | — | ~2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 19 | Turns course slides, homework files, class code and earlier solutions into concise student-style Markdown answers, with optional subagents for solving and review. | vect-G/ | 127 | — | ~729 | Automated safety check: Pass | MIT | 5 mo ago |
| 20 | Picks the right Python library for PDF jobs, with examples for text and table extraction, merging, splitting and page extraction, plus OCR and memory fixes. | FareedKhan-dev/ | 298 | — | ~785 | Automated safety check: Pass | MIT | 6 mo ago |
| 21 | Pulls architecture, method and result figures from an arXiv paper or PDF into an Obsidian vault and writes an index of them. | juliye2025/ | 1.7k | — | ~298 | Automated safety check: Pass | No licence | 24 days ago |
| 22 | Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) inside a Python program using the lexoid library. | oidlabs-com/ | 109 | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 23 | Converts scanned faxes, images, CSV/TSV exports and C-CDA XML into clean text on-device, ready for OpenMed de-identification and named-entity recognition. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | 24.Gaik Toolkit GAIK toolkit overview and reference. An agent skill from GAIK-project/gaik-toolkit. | GAIK-project/ | 100 | — | ~5.7k | Automated safety check: Pass | MIT | today |
| 25 | Creates, edits and reads PowerPoint .pptx files with python-pptx or PptxGenJS, with scripts for XML edits, text dumps, PDF and image rendering, and thumbnails. | XiaomiMiMo/ | 14k | — | ~6.8k | Automated safety check: Notes | Apache-2.0 | today |
| 26 | Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe. | telagod/ | 243 | — | ~532 | Automated safety check: Notes | MIT | 2 mo ago |
| 27 | Parses a math modeling contest problem from PDF, Word or pasted text and produces a task breakdown, paper outline, scoring map and model route for each question. | yushui2022/ | 453 | 1 repo | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 28 | 28.Docling PRIMARY skill for converting .pdf, .docx, .pptx, .xlsx, .doc, .ppt, .xls, images, and audio/video files (.mp3, .wav, .m4a, .mp4, .mov, etc.) to Markdown. | zhuzhaoyun/ | 433 | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 29 | Lists, filters and exports local Codex session transcripts to organized Markdown with a digest, timeline and tool-call details, in full or pick-one mode. | sugarforever/ | 137 | — | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 30 | Converts an EPUB card book into an Obsidian folder with one Markdown file per card, keyword cards, navigation links and an optional shareable zip. | twhsi/ | 259 | — | ~1k | Automated safety check: Pass | No licence | 2 mo ago |
| 31 | 31.PDF Reader Read and extract content from PDF files — text, tables, metadata, and images. | espennilsen/ | 122 | — | ~1.6k | Automated safety check: Pass | MIT | 17 days ago |
| 32 | 32.Markitdown Convert heterogeneous documents and selected URIs to Markdown with Microsoft MarkItDown for text analysis, search, and LLM/RAG ingestion. | K-Dense-AI/ | 2.4k | 1 repo | ~2.9k | Automated safety check: Pass | MIT | today |
| 33 | Extracts figures from a research paper, preferring the arXiv source package for original-quality images and falling back to PDF extraction. | LigphiDonk/ | 738 | 1 repo | ~810 | Automated safety check: Pass | MIT | 5 mo ago |
| 34 | 34.Liteparse Local document and PDF parsing that returns spatial text with bounding boxes. | K-Dense-AI/ | 48k | 1 repo | ~3.1k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 35 | Converts .xlsx workbooks to Markdown with a bundled MarkItDown script so an agent can read, summarize or pull data from spreadsheets, single files or whole folders. | github/ | 40k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 36 | Converts PDF files to Markdown with a bundled MarkItDown-based script before the agent reads, summarizes or extracts from them, and pulls embedded images out separately. | github/ | 40k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 37 | 37.PDF Explore A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs its content: summarize a section, compare sections, read specific pages, check the… | xuzhougeng/ | 1k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 38 | 38.Markitdown Convert files and Office documents into clean Markdown when you need LLM-friendly, token-efficient text (e.g., for summarization, search, RAG ingestion, or dataset preparation). | aipoch/ | 2k | — | ~1.3k | Automated safety check: Pass | MIT | 22 days ago |
| 39 | Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC… | BlackBeltTechnology/ | 315 | — | ~999 | Automated safety check: Pass | MIT | today |