PDF Processing
anthropics/skills
Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.
Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.
$ npx skills add telagod/code-abyss --skill processing-pdfs -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install telagod/code-abyss processing-pdfs --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/telagod/code-abyss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/processing-pdfs .claude/skills/processing-pdfs && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "processing-pdfs" agent skill from https://github.com/telagod/code-abyss/tree/main/skills/processing-pdfs into .claude/skills/processing-pdfs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processing-pdfs", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/telagod/code-abyss/tree/main/skills/processing-pdfsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add telagod/code-abyss --skill processing-pdfs -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install telagod/code-abyss processing-pdfs --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/telagod/code-abyss.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/processing-pdfs .agents/skills/processing-pdfs && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "processing-pdfs" agent skill from https://github.com/telagod/code-abyss/tree/main/skills/processing-pdfs into .agents/skills/processing-pdfs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processing-pdfs", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add telagod/code-abyss --skill processing-pdfs -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install telagod/code-abyss processing-pdfs --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/telagod/code-abyss.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/processing-pdfs .cursor/skills/processing-pdfs && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "processing-pdfs" agent skill from https://github.com/telagod/code-abyss/tree/main/skills/processing-pdfs into .cursor/skills/processing-pdfs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processing-pdfs", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/telagod/code-abyss.git --path skills/processing-pdfs--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add telagod/code-abyss --skill processing-pdfs -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install telagod/code-abyss processing-pdfs --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/telagod/code-abyss.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/processing-pdfs .gemini/skills/processing-pdfs && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "processing-pdfs" agent skill from https://github.com/telagod/code-abyss/tree/main/skills/processing-pdfs into .gemini/skills/processing-pdfs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processing-pdfs", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install telagod/code-abyss processing-pdfsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add telagod/code-abyss --skill processing-pdfs -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/telagod/code-abyss.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/processing-pdfs .github/skills/processing-pdfs && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "processing-pdfs" agent skill from https://github.com/telagod/code-abyss/tree/main/skills/processing-pdfs into .github/skills/processing-pdfs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processing-pdfs", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add telagod/code-abyss --skill processing-pdfs -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install telagod/code-abyss processing-pdfs --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/telagod/code-abyss.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/processing-pdfs .opencode/skills/processing-pdfs && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "processing-pdfs" agent skill from https://github.com/telagod/code-abyss/tree/main/skills/processing-pdfs into .opencode/skills/processing-pdfs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processing-pdfs", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
processing-pdfsPicks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.
A decision matrix maps each task to its best tool: pypdf for merging, splitting, metadata, rotation and encryption; pdfplumber for layout-preserving text and table extraction; reportlab for creating a new PDF from scratch; qpdf or pdftk for batch command-line operations with no Python needed; and pytesseract with pdf2image for OCR on scanned documents. Form filling is split between pdf-lib and pypdf and documented separately in FORMS.md, while deeper pypdfium2 and pdf-lib JS usage sits in REFERENCE.md.
The workflow is to identify which row of the matrix fits the task, load recipes.md for the common 90% of cases or advanced.md for OCR and encryption, copy and adapt the matching recipe, then validate by opening the result in a viewer or grepping the extracted text. Eight bundled Python scripts handle bounding-box checks, fillable-field detection, image conversion and form-filling with or without annotations, alongside a test for the bounding-box checker.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 2544577. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteEditGlobFrom allowed-tools in the SKILL.md frontmatter.
Ships 8 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF Processing Toolkit loads about 532 tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 150 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Write, Edit, GlobAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from telagod/code-abyss at commit 2544577, republished under its MIT licence (© telagod). 150 words, ~532 tokens.
.claude/skills/processing-pdfs/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Essential PDF operations using Python libraries and CLI tools.
| Task | Best Tool | Reference |
|---|---|---|
| Merge / split / metadata / rotate | pypdf | recipes.md |
| Extract text (layout preserved) | pdfplumber | recipes.md |
| Extract tables | pdfplumber | recipes.md |
| Create new PDF | reportlab | recipes.md |
| Batch CLI ops | qpdf / pdftk | recipes.md |
| OCR scanned PDFs | pytesseract + pdf2image | advanced.md |
| Add watermark / extract images / encrypt | pypdf / pdfimages | advanced.md |
| Fill PDF forms | pdf-lib / pypdf | FORMS.md |
| Advanced pypdfium2 / pdf-lib JS | — | REFERENCE.md |
from pypdf import PdfReader
reader = PdfReader("document.pdf")
print(f"Pages: {len(reader.pages)}")
text = "".join(page.extract_text() for page in reader.pages)| Library | Use for |
|---|---|
| pypdf | Merge, split, metadata, encryption, rotation |
| pdfplumber | Text extraction with layout, tables |
| reportlab | Generate PDFs programmatically |
| pdf2image + pytesseract | OCR scanned documents |
| qpdf / pdftk (CLI) | Batch ops, no Python needed |
© telagod, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts, references) in skills/processing-pdfs of telagod/code-abyss.
Open the folder on GitHubat commit 2544577
PDF Processing Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF Processing Toolkit this skilltelagod/code-abyss | 244 | — | ~532 | Automated safety check: Notes | MIT | |
| PDF Processinganthropics/skills | 180k | 48 repos | ~2k | Automated safety check: Pass | Proprietary | |
| PDF Processing with PythonHKUDS/DeepTutor | 41k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| PDF ToolkitTokenRhythm/opensquilla | 7.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| PDF ToolkitXiaomiMiMo/MiMo-Code | 14k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai | 3.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 |
anthropics/skills
Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.
HKUDS/DeepTutor
Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.
TokenRhythm/opensquilla
Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.
XiaomiMiMo/MiMo-Code
Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.
pipeshub-ai/pipeshub-ai
Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.
agentscope-ai/QwenPaw
Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.
telagod/code-abyss
Scans code with a bundled Node scanner for injection, secret leaks and other dangerous patterns, and requires documented decisions for accepted risks.
telagod/code-abyss
Distills a recurring agent voice from conversations into a restricted Persona Voice Card, validates it for schema, safety and distinctness, and prepares it for community submission.
telagod/code-abyss
Distills repeated workflows into new skills, improves existing ones and promotes them through local, project and community tiers after a default-deny safety scan.
telagod/code-abyss
Routes Word document tasks to the right tool: pandoc for text, raw OOXML for structure and comments, docx-js for new files, and a mandatory redlining flow for edits to others' documents.
telagod/code-abyss
Checks what a code change touched, how far its impact reaches and whether design docs, tests and README files kept up, before commit or in review.
telagod/code-abyss
Checks cyclomatic complexity, function and file length, parameter count, nesting depth and naming conventions against fixed thresholds, with a bundled Node.js script and hotspot integration.
Categories
Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe. A decision matrix maps each task to its best tool: pypdf for merging, splitting, metadata, rotation and encryption; pdfplumber for layout-preserving text and table extraction; reportlab for creating a new PDF from scratch; qpdf or pdftk for batch command-line operations with no Python needed; and pytesseract with pdf2image for OCR on scanned documents.md.
PDF Processing Toolkit fits situations like: extracting text or tables from a PDF with its layout preserved; filling in a PDF form's fields programmatically; merging, splitting or rotating PDF files in batch; running OCR on a scanned PDF to make its text searchable.
Run `npx skills add telagod/code-abyss --skill processing-pdfs -a claude-code`. Or copy the skill folder (skills/processing-pdfs in telagod/code-abyss) into .claude/skills/processing-pdfs in your project. Claude Code loads it when a task matches its description.
Run `npx skills add telagod/code-abyss --skill processing-pdfs -a codex`. Or copy the skill folder (skills/processing-pdfs in telagod/code-abyss) into .agents/skills/processing-pdfs in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add telagod/code-abyss --skill processing-pdfs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/processing-pdfs, .gemini/skills/processing-pdfs, .github/skills/processing-pdfs and .opencode/skills/processing-pdfs in your project.
Going by SKILL.md and its folder, PDF Processing Toolkit needs Python for the scripts in its folder. Our summary lists: Python with pypdf, pdfplumber, reportlab and pdf2image/pytesseract as needed; qpdf or pdftk for CLI batch operations. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit, Glob.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
PDF Processing Toolkit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 532 tokens (SKILL.md is roughly 2.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with PDF Processing Toolkit: PDF Processing (anthropics/skills, 180k stars), PDF Processing with Python (HKUDS/DeepTutor, 41k stars), PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars) and PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
telagod (a GitHub user) maintains it in telagod/code-abyss, which has 244 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on July 19, 2026.
Source: telagod/code-abyss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.