PDF Toolkit
XiaomiMiMo/MiMo-Code
Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.
Read and extract content from PDF files — text, tables, metadata, and images.
$ npx skills add espennilsen/pi --skill pdf-reader -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install espennilsen/pi pdf-reader --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf-reader .claude/skills/pdf-reader && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pdf-reader" agent skill from https://github.com/espennilsen/pi/tree/main/skills/pdf-reader into .claude/skills/pdf-reader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-reader", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/espennilsen/pi/tree/main/skills/pdf-readerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add espennilsen/pi --skill pdf-reader -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install espennilsen/pi pdf-reader --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/pdf-reader .agents/skills/pdf-reader && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pdf-reader" agent skill from https://github.com/espennilsen/pi/tree/main/skills/pdf-reader into .agents/skills/pdf-reader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-reader", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add espennilsen/pi --skill pdf-reader -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install espennilsen/pi pdf-reader --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/pdf-reader .cursor/skills/pdf-reader && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pdf-reader" agent skill from https://github.com/espennilsen/pi/tree/main/skills/pdf-reader into .cursor/skills/pdf-reader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-reader", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/espennilsen/pi.git --path skills/pdf-reader--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add espennilsen/pi --skill pdf-reader -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install espennilsen/pi pdf-reader --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/pdf-reader .gemini/skills/pdf-reader && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pdf-reader" agent skill from https://github.com/espennilsen/pi/tree/main/skills/pdf-reader into .gemini/skills/pdf-reader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-reader", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install espennilsen/pi pdf-readerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add espennilsen/pi --skill pdf-reader -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/pdf-reader .github/skills/pdf-reader && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pdf-reader" agent skill from https://github.com/espennilsen/pi/tree/main/skills/pdf-reader into .github/skills/pdf-reader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-reader", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add espennilsen/pi --skill pdf-reader -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install espennilsen/pi pdf-reader --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/pdf-reader .opencode/skills/pdf-reader && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pdf-reader" agent skill from https://github.com/espennilsen/pi/tree/main/skills/pdf-reader into .opencode/skills/pdf-reader/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-reader", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pdf-readerRead and extract content from PDF files — text, tables, metadata, and images.
PDF Reader is an agent skill from espennilsen/pi. Read and extract content from PDF files — text, tables, metadata, and images. Use when asked to read a PDF, extract text from a PDF, summarize a PDF, analyze a PDF document, get tables from a PDF, or check PDF metadata. Also triggers on "open this PDF", "what does this PDF say", "parse PDF", "PDF to text", or when a .pdf file path or URL is provided.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/extract.py`).
It sits in Documents & Office, covering PDF and Document parsing. It works with Python and pypdf. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 79d019b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pdftotextpython3curltesseractbrewFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF Reader loads about 1.6k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 499 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from espennilsen/pi at commit 79d019b, republished under its MIT licence (© espennilsen). 499 words, ~1,575 tokens.
.claude/skills/pdf-reader/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Extract content from PDF files using pdftotext (Poppler) for text and
pdfplumber (Python) for tables and structured extraction.
| Task | Tool | Command |
|---|---|---|
| Full text | pdftotext | pdftotext file.pdf - |
| Text with layout | pdftotext | pdftotext -layout file.pdf - |
| Specific pages | pdftotext | pdftotext -f 3 -l 5 file.pdf - |
| Tables | pdfplumber | python3 scripts/extract.py tables file.pdf |
| Metadata | pdfinfo | pdfinfo file.pdf |
| Page count | pdfinfo | pdfinfo file.pdf | grep Pages |
| Images | pdfimages | pdfimages -list file.pdf |
| Fonts | pdffonts | pdffonts file.pdf |
| OCR (scanned PDF) | tesseract | python3 scripts/extract.py ocr file.pdf |
| Smart text + OCR | extract.py | python3 scripts/extract.py text file.pdf |
| Quick survey | extract.py | python3 scripts/extract.py scan file.pdf |
If the user provides a URL, download it first:
curl -sL "URL" -o /tmp/document.pdfVerify it's a valid PDF:
file /tmp/document.pdf # should say "PDF document"
pdfinfo /tmp/document.pdf # metadata + page countPlain text (most cases):
pdftotext file.pdf -This pipes output to stdout. For large PDFs, use page ranges:
pdftotext -f 1 -l 10 file.pdf - # pages 1-10Layout-preserving text (columns, formatted docs):
pdftotext -layout file.pdf -Use -layout when the PDF has multi-column layouts, tables rendered as text,
or precise spacing that matters.
Tables (structured data):
python3 scripts/extract.py tables file.pdfOr inline with pdfplumber:
import pdfplumber
pdf = pdfplumber.open("file.pdf")
for i, page in enumerate(pdf.pages):
tables = page.extract_tables()
for table in tables:
print(f"\n--- Table on page {i+1} ---")
for row in table:
print(" | ".join(str(cell or "") for cell in row))
pdf.close()Metadata only:
pdfinfo file.pdfReturns: title, author, creator, producer, page count, page size, dates.
For PDFs over ~50 pages, don't dump everything at once:
pdfinfo file.pdf | grep Pagespdftotext -f 1 -l 20 file.pdf -pdftotext -f 21 -l 40 file.pdf -For targeted extraction (searching for specific content):
# Extract all text, grep for relevant sections
pdftotext file.pdf - | grep -n -i "keyword"
# Then extract the specific page range
pdftotext -f PAGE -l PAGE file.pdf -If pdftotext returns empty or garbled output, the PDF is likely scanned.
Detection:
python3 scripts/extract.py scan file.pdf # reports scanned pages
pdffonts file.pdf # empty = image-basedSmart extraction (auto-fallback):
text mode automatically detects scanned pages and OCRs them:
python3 scripts/extract.py text file.pdfPages with selectable text extract normally. Pages without selectable text fall back to OCR via Tesseract. No manual detection needed.
Force OCR on all pages:
python3 scripts/extract.py ocr file.pdf
python3 scripts/extract.py ocr file.pdf --pages 1-5
python3 scripts/extract.py ocr file.pdf --dpi 400 # higher quality
python3 scripts/extract.py ocr file.pdf --lang eng+nor # multi-languageOCR options:
--dpi 300 — resolution for page-to-image conversion (default: 300, higher = slower but better)--lang eng — Tesseract language pack (default: eng). Use + for multiple: eng+nor+deu--pages 1-5 — limit to specific pages (recommended for large PDFs)Available language packs:
tesseract --list-langsInstall additional languages via Homebrew:
brew install tesseract-lang # all languagesMixed PDFs (some pages scanned, some not):
Just use text mode — it handles mixed PDFs automatically:
python3 scripts/extract.py text file.pdfSelectable pages extract instantly, scanned pages get OCR'd. The output is tagged so you know which pages used OCR.
Password-protected PDFs:
pdftotext -upw "password" file.pdf - # user password
pdftotext -opw "password" file.pdf - # owner passwordEncoding issues (garbled output):
pdftotext -enc UTF-8 file.pdf -Extract images:
pdfimages -png file.pdf /tmp/images/img # extracts as PNG
pdfimages -list file.pdf # list images without extractingIs it a URL? → curl -sL "URL" -o /tmp/doc.pdf
↓
Run: python3 scripts/extract.py scan file.pdf
↓
All pages have selectable text?
YES → pdftotext file.pdf - (fast, simple)
NO → python3 scripts/extract.py text file.pdf (auto OCR fallback)
↓
Need tables?
YES → python3 scripts/extract.py tables file.pdfscan on unknown PDFs — it reports pages, tables, scanned detection, and a previewpdftotext is fastest for normal PDFs — try it first-layout for multi-column documents (academic papers, reports)pdfplumber is better for tables — it understands cell boundariestext mode auto-detects scanned pages and OCRs only those — preferred over raw pdftotext for unknown PDFsocr mode is for forcing OCR on everything (useful when text extraction gives garbled output despite appearing selectable)--dpi gives better OCR accuracy but is slower (300 is a good default, 400+ for small text)/tmp/ first — don't pipe curl to tools--pages to extract in ranges© espennilsen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in skills/pdf-reader of espennilsen/pi.
Open the folder on GitHubat commit 79d019b
PDF Reader next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF Reader this skillespennilsen/pi | 122 | — | ~1.6k | Automated safety check: Pass | MIT | |
| PDF ToolkitXiaomiMiMo/MiMo-Code | 14k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai | 3.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| PDF Explorexuzhougeng/wisp-science | 1k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| PDF Processinganthropics/skills | 180k | 48 repos | ~2k | Automated safety check: Pass | Proprietary | |
| PDF Processing with PythonHKUDS/DeepTutor | 41k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 |
XiaomiMiMo/MiMo-Code
Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.
pipeshub-ai/pipeshub-ai
Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.
xuzhougeng/wisp-science
A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs its content: summarize a section, compare sections, read specific pages, check the…
anthropics/skills
Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.
HKUDS/DeepTutor
Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.
TokenRhythm/opensquilla
Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.
espennilsen/pi
Interact with GitHub repos, PRs, issues, CI, and notifications via the pi-github extension commands and gh CLI.
espennilsen/pi
Create, review, and improve skills for Pi agents. An agent skill from espennilsen/pi.
espennilsen/pi
Perform a comprehensive DRY (Don't Repeat Yourself) code review on a codebase.
espennilsen/pi
Reverse-engineer a design system from a live website (public URL or localhost).
espennilsen/pi
Manage Google Workspace via the gws CLI — Drive, Gmail, Sheets, Docs, Slides, People, Chat, Meet, Forms, and cross-service workflows.
espennilsen/pi
A skill your agent uses when inspecting or operating Herdr sessions, workspaces, tabs, panes, agents, terminal output, agent messaging, or waits.
Categories
Read and extract content from PDF files — text, tables, metadata, and images. PDF Reader is an agent skill from espennilsen/pi. Read and extract content from PDF files — text, tables, metadata, and images.
PDF Reader fits situations like: asked to read a PDF; extract text from a PDF; summarize a PDF; analyze a PDF document.
Run `npx skills add espennilsen/pi --skill pdf-reader -a claude-code`. Or copy the skill folder (skills/pdf-reader in espennilsen/pi) into .claude/skills/pdf-reader in your project. Claude Code loads it when a task matches its description.
Run `npx skills add espennilsen/pi --skill pdf-reader -a codex`. Or copy the skill folder (skills/pdf-reader in espennilsen/pi) into .agents/skills/pdf-reader in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add espennilsen/pi --skill pdf-reader -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-reader, .gemini/skills/pdf-reader, .github/skills/pdf-reader and .opencode/skills/pdf-reader in your project.
Going by SKILL.md and its folder, PDF Reader needs Python for the scripts in its folder and the command-line tools its instructions call (pdftotext, python3, curl, tesseract and brew). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
PDF Reader is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with PDF Reader: PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars), PDF Explore (xuzhougeng/wisp-science, 1k stars) and PDF Processing (anthropics/skills, 180k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
espennilsen (a GitHub user) maintains it in espennilsen/pi, which has 122 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 21, 2026.
Source: espennilsen/pi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.