Agent skill

PDF Processing Toolkit

by telagod in telagod/code-abyss

Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.

MITAuto-check: notesDocuments & Office

Install PDF Processing Toolkit

skills CLI
$ npx skills add telagod/code-abyss --skill processing-pdfs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install telagod/code-abyss processing-pdfs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/telagod/code-abyss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/processing-pdfs .claude/skills/processing-pdfs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
processing-pdfs
GitHub stars
244
Token cost
~532 tokens
SKILL.md length
150 words
Files
13 (incl. scripts, references)
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.

  • Works in 4 steps: Identify task — text extraction? table?… → Load reference — recipes.md covers 90%… → Implement — copy-adapt recipe; verify… → …
  • Extracting text or tables from a PDF with its layout preserved
  • SKILL.md covers Decision Matrix, Quick Start, Workflow and Library Selection
  • Runs Python scripts from its folder

What it does

A decision matrix maps each task to its best tool: pypdf for merging, splitting, metadata, rotation and encryption; pdfplumber for layout-preserving text and table extraction; reportlab for creating a new PDF from scratch; qpdf or pdftk for batch command-line operations with no Python needed; and pytesseract with pdf2image for OCR on scanned documents. Form filling is split between pdf-lib and pypdf and documented separately in FORMS.md, while deeper pypdfium2 and pdf-lib JS usage sits in REFERENCE.md.

The workflow is to identify which row of the matrix fits the task, load recipes.md for the common 90% of cases or advanced.md for OCR and encryption, copy and adapt the matching recipe, then validate by opening the result in a viewer or grepping the extracted text. Eight bundled Python scripts handle bounding-box checks, fillable-field detection, image conversion and form-filling with or without annotations, alongside a test for the bounding-box checker.

When your agent uses it

  • Extracting text or tables from a PDF with its layout preserved
  • Filling in a PDF form's fields programmatically
  • Merging, splitting or rotating PDF files in batch
  • Running OCR on a scanned PDF to make its text searchable

Example prompts

  • “Extract the tables from ./reports/q3.pdf and save them as CSV.”
  • “Fill in the fields of ./forms/w9.pdf using the values in ./data/applicant.json.”
  • “Merge these five PDFs into one file in the given order.”
  • “Run OCR on this scanned contract and extract the text.”

Requirements

  • Python with pypdf, pdfplumber, reportlab and pdf2image/pytesseract as needed
  • qpdf or pdftk for CLI batch operations
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit, Glob

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Identify task — text extraction? table? creation? form? Pick row from matrix above.
  2. Load reference — recipes.md covers 90% of tasks; advanced.md for OCR / encrypt; FORMS.md for forms.
  3. Implement — copy-adapt recipe; verify output.
  4. Validate — open in a viewer or grep extracted text.

What it can do on your machine

Read from SKILL.md and the folder at commit 2544577. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 8 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Processing Toolkit loads about 532 tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 150 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~532
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from telagod/code-abyss at commit 2544577, republished under its MIT licence (© telagod). 150 words, ~532 tokens.

Download SKILL.mdSave it as .claude/skills/processing-pdfs/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
processing-pdfs
description
Processes PDF files. Extracts text and tables, fills forms, merges and splits documents, batch-processes files, converts to images, and generates PDFs programmatically. Use when working with .pdf files. Do NOT use for Word documents, spreadsheets, or presentations.
allowed-tools
Bash, Read, Write, Edit, Glob
user-invocable
false
argument-hint
<file.pdf | task>

PDF Processing

Essential PDF operations using Python libraries and CLI tools.

Decision Matrix

TaskBest ToolReference
Merge / split / metadata / rotatepypdfrecipes.md
Extract text (layout preserved)pdfplumberrecipes.md
Extract tablespdfplumberrecipes.md
Create new PDFreportlabrecipes.md
Batch CLI opsqpdf / pdftkrecipes.md
OCR scanned PDFspytesseract + pdf2imageadvanced.md
Add watermark / extract images / encryptpypdf / pdfimagesadvanced.md
Fill PDF formspdf-lib / pypdfFORMS.md
Advanced pypdfium2 / pdf-lib JS—REFERENCE.md

Quick Start

python
from pypdf import PdfReader
reader = PdfReader("document.pdf")
print(f"Pages: {len(reader.pages)}")
text = "".join(page.extract_text() for page in reader.pages)

Workflow

  1. Identify task — text extraction? table? creation? form? Pick row from matrix above.
  2. Load reference — recipes.md covers 90% of tasks; advanced.md for OCR / encrypt; FORMS.md for forms.
  3. Implement — copy-adapt recipe; verify output.
  4. Validate — open in a viewer or grep extracted text.

Library Selection

LibraryUse for
pypdfMerge, split, metadata, encryption, rotation
pdfplumberText extraction with layout, tables
reportlabGenerate PDFs programmatically
pdf2image + pytesseractOCR scanned documents
qpdf / pdftk (CLI)Batch ops, no Python needed

© telagod, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references) in skills/processing-pdfs of telagod/code-abyss.

  • SKILL.md
  • FORMS.md
  • REFERENCE.md
  • references/advanced.md
  • references/recipes.md
  • scripts/check_bounding_boxes.py
  • scripts/check_bounding_boxes_test.py
  • scripts/check_fillable_fields.py
  • scripts/convert_pdf_to_images.py
  • scripts/create_validation_image.py
  • scripts/extract_form_field_info.py
  • scripts/fill_fillable_fields.py
  • scripts/fill_pdf_form_with_annotations.py

Open the folder on GitHubat commit 2544577

Compare with similar skills

PDF Processing Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Processing Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Processing Toolkit this skilltelagod/code-abyss244—~532Automated safety check: NotesMIT
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing with PythonHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0
PDF ToolkitTokenRhythm/opensquilla7.1k—~1.9kAutomated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    agentscope-ai/QwenPaw

    Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

    35k GitHub stars~1.8k tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed

More from telagod/code-abyss

All 38 skills in this repo
  • Security Gate Scanner

    telagod/code-abyss

    Scans code with a bundled Node scanner for injection, secret leaks and other dangerous patterns, and requires documented decisions for accepted risks.

    244 GitHub stars~552 tokensUpdated 2 mo ago
    Auto-check: notes
  • Persona Voice Card Builder

    telagod/code-abyss

    Distills a recurring agent voice from conversations into a restricted Persona Voice Card, validates it for schema, safety and distinctness, and prepares it for community submission.

    244 GitHub stars~716 tokensUpdated 2 mo ago
    Auto-check: notes
  • Skill Cultivation Funnel

    telagod/code-abyss

    Distills repeated workflows into new skills, improves existing ones and promotes them through local, project and community tiers after a default-deny safety scan.

    244 GitHub stars~794 tokensUpdated 2 mo ago
    Auto-check: notes
  • DOCX Processing Toolkit

    telagod/code-abyss

    Routes Word document tasks to the right tool: pandoc for text, raw OOXML for structure and comments, docx-js for new files, and a mandatory redlining flow for edits to others' documents.

    244 GitHub stars~668 tokensUpdated 2 mo ago
    Auto-check: notes
  • Code Change Analysis Gate

    telagod/code-abyss

    Checks what a code change touched, how far its impact reaches and whether design docs, tests and README files kept up, before commit or in review.

    244 GitHub stars~532 tokensUpdated 2 mo ago
    Auto-check: notes
  • Code Quality Gate Checker

    telagod/code-abyss

    Checks cyclomatic complexity, function and file length, parameter count, nesting depth and naming conventions against fixed thresholds, with a bundled Node.js script and hotspot integration.

    244 GitHub stars~609 tokensUpdated 2 mo ago
    Auto-check: notes

Works with

Questions about PDF Processing Toolkit

What does PDF Processing Toolkit do?

Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe. A decision matrix maps each task to its best tool: pypdf for merging, splitting, metadata, rotation and encryption; pdfplumber for layout-preserving text and table extraction; reportlab for creating a new PDF from scratch; qpdf or pdftk for batch command-line operations with no Python needed; and pytesseract with pdf2image for OCR on scanned documents.md.

When should I use PDF Processing Toolkit?

PDF Processing Toolkit fits situations like: extracting text or tables from a PDF with its layout preserved; filling in a PDF form's fields programmatically; merging, splitting or rotating PDF files in batch; running OCR on a scanned PDF to make its text searchable.

How do I install PDF Processing Toolkit in Claude Code?

Run `npx skills add telagod/code-abyss --skill processing-pdfs -a claude-code`. Or copy the skill folder (skills/processing-pdfs in telagod/code-abyss) into .claude/skills/processing-pdfs in your project. Claude Code loads it when a task matches its description.

How do I install PDF Processing Toolkit in Codex?

Run `npx skills add telagod/code-abyss --skill processing-pdfs -a codex`. Or copy the skill folder (skills/processing-pdfs in telagod/code-abyss) into .agents/skills/processing-pdfs in your project. Codex loads it when a task matches its description.

Can I use PDF Processing Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add telagod/code-abyss --skill processing-pdfs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/processing-pdfs, .gemini/skills/processing-pdfs, .github/skills/processing-pdfs and .opencode/skills/processing-pdfs in your project.

What does PDF Processing Toolkit need to run?

Going by SKILL.md and its folder, PDF Processing Toolkit needs Python for the scripts in its folder. Our summary lists: Python with pypdf, pdfplumber, reportlab and pdf2image/pytesseract as needed; qpdf or pdftk for CLI batch operations. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit, Glob.

Does PDF Processing Toolkit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF Processing Toolkit safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF Processing Toolkit use?

PDF Processing Toolkit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Processing Toolkit use?

About 532 tokens (SKILL.md is roughly 2.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to PDF Processing Toolkit?

Skills that share tags, products or a category with PDF Processing Toolkit: PDF Processing (anthropics/skills, 180k stars), PDF Processing with Python (HKUDS/DeepTutor, 41k stars), PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars) and PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Processing Toolkit?

telagod (a GitHub user) maintains it in telagod/code-abyss, which has 244 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on July 19, 2026.

Source: telagod/code-abyss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.