Agent skill

PDF Toolkit

by XiaomiMiMo in XiaomiMiMo/MiMo-Code

Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

Apache-2.0Auto-check passedDocuments & Office

Install PDF Toolkit

skills CLI
$ npx skills add XiaomiMiMo/MiMo-Code --skill pdf-official -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install XiaomiMiMo/MiMo-Code pdf-official --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/XiaomiMiMo/MiMo-Code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/cli/src/skill/builtin/.bundle/pdf-official .claude/skills/pdf-official && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-official
GitHub stars
14k
Token cost
~1.7k tokens
SKILL.md length
638 words
Files
18 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

  • Works in 6 steps: PDF origin is bottom-left, image origin… → pypdf.extract_text() returns nothing for… → **Unicode subscripts / superscripts… → …
  • Extracting text or tables from an existing PDF
  • SKILL.md covers Route the task, First install, One-command triage and Which library for which task, plus 2 more sections
  • Runs Python scripts from its folder; calls brew, apt-get and qpdf

What it does

This toolkit handles four kinds of PDF work and routes by the verb in your request: extracting text, tables, metadata and images; transforming files by combining, carving, rotating, cropping, watermarking, encrypting or shrinking; composing new PDFs such as reports, invoices and certificates; and filling forms, whether AcroForm or scanned. Scanned documents go through OCR.

Every task starts with a probe: survey.py reports the page count, whether the file is encrypted, whether it has an AcroForm and whether page one looks scanned. Mixed tasks follow the order probe, plan, extract or compose, then validate. Separate scripts cover applying form values, carving pages, combining files, overlaying text, rendering pages, reorienting, text dumps and sanity checks, each with argparse and defined exit codes.

It is built on permissively licensed Python libraries (pypdf, pdfplumber, pypdfium2, reportlab) and optionally qpdf, so it can be embedded in commercial projects. Installation is a single pip command, and a bundled runtime is used instead when the MIMO_PYTHON environment variable is set.

When your agent uses it

  • Extracting text or tables from an existing PDF
  • Merging, splitting, rotating or watermarking PDF pages
  • Filling a fillable form or overlaying text onto a scanned one
  • Building a new PDF report, invoice or certificate
  • Running OCR over a scanned document

Example prompts

  • “Extract all tables from ./reports/q3.pdf into CSV files.”
  • “Merge cover.pdf and body.pdf, then rotate the last page.”
  • “Fill in the W-9 form in ./forms/w9.pdf with my details.”
  • “OCR the scanned contract in ./scans/contract.pdf so it becomes searchable.”

Requirements

  • Python 3 with pypdf, pdfplumber, pypdfium2, reportlab and Pillow
  • qpdf, only for merge, split, encrypt and repair tasks

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. PDF origin is bottom-left, image origin is top-left. Every "off by a
  2. pypdf.extract_text() returns nothing for scans. That's not a bug —
  3. **Unicode subscripts / superscripts render as black rectangles in
  4. CJK text renders as black boxes when the font never registered.
  5. XFA forms are not AcroForms. If probe_fields.py returns [] on a
  6. writer.encrypt(pw) in pypdf uses RC4 by default. For real AES-256,

What it can do on your machine

Read from SKILL.md and the folder at commit 6babeb0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 11 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • brew
    • apt-get
    • qpdf
    • python3
    • uv
    • pdftotext

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Toolkit loads about 1.7k tokens when it runs. Until then it costs about 155 tokens; SKILL.md has 638 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~155
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from XiaomiMiMo/MiMo-Code at commit 6babeb0, republished under its Apache-2.0 licence (© XiaomiMiMo). 638 words, ~1,655 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-official/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
pdf-official
description
Use this skill whenever a PDF file is being produced, opened, transformed, filled, or read. That includes: extracting text or tables from an existing PDF; combining, carving, rotating, cropping, or watermarking pages; composing a fresh PDF (report, invoice, certificate); filling AcroForm fields or overlaying text onto a non-fillable scanned form; encrypting or unlocking a PDF; running OCR over a scanned document; rendering pages to PNG/JPEG for visual analysis. Trigger on mentions of 'PDF', a filename ending in .pdf, requests like 'turn this into a PDF report', or references to AcroForm / form fields.
license
Apache-2.0 — see LICENSE for terms and third-party attributions

PDF skill

An Apache-2.0 toolkit for reading, composing, transforming, and filling PDF files. Written from scratch on top of permissively-licensed open-source libraries (pypdf, pdfplumber, pypdfium2, reportlab, pdf-lib, qpdf) so this can be embedded in commercial projects without special agreement.

Route the task

Pick the sub-guide by the verb of the request.

TaskPathRead
Pull text / tables / metadata / images out of an existing PDFExtractextract.md
Combine, carve, rotate, crop, watermark, encrypt, or shrinkTransformtransform.md
Build a PDF that doesn't exist yet (report, invoice, certificate)Composecompose.md
Fill a form (AcroForm or scanned)Interactiveinteractive.md
Scanned / image-only PDF (no selectable text)Extract → OCRextract.md §5

If a task mixes several of these, follow the order: probe → plan → extract or compose → validate.

Every path starts with a probe. scripts/survey.py returns page count, whether the file is encrypted, whether it has an AcroForm, and whether page 1 looks like a scan.

First install

Bundled runtime: when the MIMO_PYTHON environment variable is set, skip the installs below — run every command with python3/uv run replaced by "$MIMO_PYTHON" (pypdf/pypdfium2/reportlab/Pillow preinstalled; pip console scripts unavailable, use "$MIMO_PYTHON" -m <module>). A bundled qpdf is exposed as MIMO_QPDF (picked up automatically by the scripts here); invoke it directly as "$MIMO_QPDF" --check file.pdf.

Python-only path (all BSD / MIT / Apache) — covers 95% of tasks:

bash
python3 -m pip install --upgrade pypdf pdfplumber pypdfium2 reportlab Pillow

Add these external binaries only when you actually need them:

bash
# qpdf — merge/split/encrypt/repair, Apache-2.0
brew install qpdf                # macOS
apt-get install -y qpdf          # Debian / Ubuntu

# Tesseract — OCR for scanned PDFs, Apache-2.0
brew install tesseract
python3 -m pip install pytesseract pdf2image
apt-get install -y tesseract-ocr

# Poppler — pdftotext / pdftoppm / pdfimages, GPL-2.0
# Optional. Only install if you accept a GPL dependency at CLI level.
brew install poppler
apt-get install -y poppler-utils

Every script under scripts/ uses argparse. Exit codes: 0 OK · 1 runtime failure · 2 bad arguments · 3 validation failure (apply_values.py / overlay_text.py; sanity_check.py reports findings with exit 1). Any single script can be lifted into another project — none imports from a shared framework.

One-command triage

bash
scripts/survey.py path/to/file.pdf --pretty

Sample output:

json
{
  "path": "/abs/path/file.pdf",
  "page_count": 12,
  "is_locked": false,
  "form_field_count": 34,
  "looks_scanned": false,
  "metadata": {"Title": "...", "Author": "...", "Producer": "..."}
}

Route by the flags:

  • is_locked: true → unlock first (qpdf --password=… --decrypt). Almost every reader library refuses locked files.
  • form_field_count > 0 → widgets path in interactive.md §1.
  • form_field_count == 0 AND you need to fill it → overlay path in interactive.md §2.
  • looks_scanned: true → skip pypdf text extraction, go straight to OCR (extract.md §5).

Which library for which task

TaskPreferredReasonFallback
Plain textpdftotext -layoutfastest, keeps columnspypdf
Positioned textpdfplumberchar-level bboxespypdfium2.get_text
Tablespdfplumbertunable table_settingspandas over manual CSV
Page → imagepypdfium2Apache/BSD, no GPLpdftoppm (GPL)
Merge / carve / rotatepypdfpure Pythonqpdf --pages (faster on huge files)
Encrypt / repair / lineariseqpdfhandles broken inputpypdf (basic encrypt only)
Compose from scratchreportlabmature, BSDpdf-lib in Node
Fill AcroFormpypdf.update_page_form_field_valuespreserves widget appearancespdf-lib in Node
Overlay on non-fillablereportlab + pypdf.merge_pagetwo-layer merge, see interactive.md—
Show full SKILL.md (228 more words)Show less

Common gotchas

  1. PDF origin is bottom-left, image origin is top-left. Every "off by a few points" bug is one of these two systems misapplied. Coordinate conversion is in one place: interactive.md §2.c.
  2. pypdf.extract_text() returns nothing for scans. That's not a bug — there's no text stream. Use the looks_scanned flag and route to OCR.
  3. Unicode subscripts / superscripts render as black rectangles in reportlab because Helvetica/Times/Courier don't ship those glyphs. Use <sub> / <super> XML in Paragraph, or move the pen manually on canvas. See compose.md §5.
  4. CJK text renders as black boxes when the font never registered. reportlab does not consult the OS font system; a bad font name/path (思源黑体, PingFang, Noto on machines that lack it) plus a swallowed exception means a silent Helvetica fallback — and Helvetica has no CJK glyphs. Resolve fonts with the ladder in compose.md §4 (resolve_cjk_font()); the terminal fallback is the built-in CID font, never Helvetica.
  5. XFA forms are not AcroForms. If probe_fields.py returns [] on a PDF that clearly has widgets in Adobe Reader, it's XFA — flatten it in Acrobat first.
  6. writer.encrypt(pw) in pypdf uses RC4 by default. For real AES-256, pass algorithm="AES-256", or use qpdf --encrypt … 256 --.

What's next

Open the sub-guide from the routing table and work through it end to end. Each sub-guide has a Validation section at the bottom describing how to confirm the result.

© XiaomiMiMo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts) in packages/cli/src/skill/builtin/.bundle/pdf-official of XiaomiMiMo/MiMo-Code.

  • SKILL.md
  • LICENSE
  • README.md
  • compose.md
  • extract.md
  • interactive.md
  • scripts/apply_values.py
  • scripts/carve.py
  • scripts/combine.py
  • scripts/overlay_text.py
  • scripts/probe_fields.py
  • scripts/recognize.py
  • scripts/render_pages.py
  • scripts/reorient.py
  • scripts/sanity_check.py
  • scripts/survey.py
  • scripts/text_dump.py
  • transform.md

Open the folder on GitHubat commit 6babeb0

Compare with similar skills

PDF Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Toolkit this skillXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing with PythonHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0
PDF ToolkitTokenRhythm/opensquilla7.1k—~1.9kAutomated safety check: PassApache-2.0
PDF Processing Guideagentscope-ai/QwenPaw35k—~1.8kAutomated safety check: PassProprietary

Similar skills

  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    agentscope-ai/QwenPaw

    Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

    35k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Toolkit

    telagod/code-abyss

    Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.

    243 GitHub stars~532 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes

More from XiaomiMiMo/MiMo-Code

All 22 skills in this repo
  • Paper Research on arXiv

    XiaomiMiMo/MiMo-Code

    Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.

    14k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Agent Skill Creator

    XiaomiMiMo/MiMo-Code

    Interactive guide for creating, reviewing and fixing agent skills (SKILL.md folders), covering structure, frontmatter rules, trigger phrases and validation before sharing.

    14k GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • DOCX Toolkit

    XiaomiMiMo/MiMo-Code

    Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task.

    14k GitHub stars~2.4k tokensUpdated 5 days ago
    Auto-check passed
  • Drive MiMo Code

    XiaomiMiMo/MiMo-Code

    Lets one MiMoCode process drive another, headless with JSON events or interactively through tmux, to test behavior and visual regressions with parseable evidence.

    14k GitHub stars~3.9k tokensUpdated 5 days ago
    Auto-check passed
  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated 5 days ago
    Auto-check passed
  • Claude Code Delegation

    XiaomiMiMo/MiMo-Code

    Hands coding work to the Claude Code CLI from the terminal in print, interactive tmux or background mode, only when you explicitly ask for Claude Code.

    14k GitHub stars~1.3k tokensUpdated 5 days ago
    Auto-check passed

Works with

Questions about PDF Toolkit

What does PDF Toolkit do?

Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling. This toolkit handles four kinds of PDF work and routes by the verb in your request: extracting text, tables, metadata and images; transforming files by combining, carving, rotating, cropping, watermarking, encrypting or shrinking; composing new PDFs such as reports, invoices and certificates; and filling forms, whether AcroForm or scanned. Scanned documents go through OCR.

When should I use PDF Toolkit?

PDF Toolkit fits situations like: extracting text or tables from an existing PDF; merging, splitting, rotating or watermarking PDF pages; filling a fillable form or overlaying text onto a scanned one; building a new PDF report, invoice or certificate.

How do I install PDF Toolkit in Claude Code?

Run `npx skills add XiaomiMiMo/MiMo-Code --skill pdf-official -a claude-code`. Or copy the skill folder (packages/cli/src/skill/builtin/.bundle/pdf-official in XiaomiMiMo/MiMo-Code) into .claude/skills/pdf-official in your project. Claude Code loads it when a task matches its description.

How do I install PDF Toolkit in Codex?

Run `npx skills add XiaomiMiMo/MiMo-Code --skill pdf-official -a codex`. Or copy the skill folder (packages/cli/src/skill/builtin/.bundle/pdf-official in XiaomiMiMo/MiMo-Code) into .agents/skills/pdf-official in your project. Codex loads it when a task matches its description.

Can I use PDF Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add XiaomiMiMo/MiMo-Code --skill pdf-official -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-official, .gemini/skills/pdf-official, .github/skills/pdf-official and .opencode/skills/pdf-official in your project.

What does PDF Toolkit need to run?

Going by SKILL.md and its folder, PDF Toolkit needs Python for the scripts in its folder and the command-line tools its instructions call (brew, apt-get, qpdf, python3, uv and pdftotext). Our summary lists: Python 3 with pypdf, pdfplumber, pypdfium2, reportlab and Pillow; qpdf, only for merge, split, encrypt and repair tasks.

Does PDF Toolkit access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Toolkit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF Toolkit use?

PDF Toolkit is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Toolkit use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Toolkit?

Skills that share tags, products or a category with PDF Toolkit: PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars), PDF Processing (anthropics/skills, 180k stars), PDF Processing with Python (HKUDS/DeepTutor, 41k stars) and PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Toolkit?

XiaomiMiMo (a GitHub organization) maintains it in XiaomiMiMo/MiMo-Code, which has 13,611 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 3, 2026.

Source: XiaomiMiMo/MiMo-Code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.