Official agent skill

PDF Processing

by anthropics in anthropics/skills

Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

OfficialProprietaryAuto-check passedDocuments & Office

Install PDF Processing

skills CLI
$ npx skills add anthropics/skills --skill pdf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install anthropics/skills pdf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/anthropics/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf .claude/skills/pdf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf
GitHub stars
180k
Used in
48 other repos
Token cost
~2k tokens
SKILL.md length
249 words
Files
12 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
Proprietary

At a glance

Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

  • Pulling text or tables out of a PDF report
  • SKILL.md covers Overview, Quick Start, Python Libraries and Command-Line Tools, plus 3 more sections
  • Runs Python scripts from its folder; calls qpdf and pdftotext
  • Combining, splitting or rotating pages across several PDFs

What it does

This skill gives the agent working recipes for PDF tasks, mainly with Python libraries. pypdf covers merging, splitting, rotating pages and reading metadata; pdfplumber pulls out text with its layout and tables, which can go straight into pandas; reportlab creates new PDFs, from a single canvas page to multi-page documents built from paragraphs.

Form filling has its own instructions in forms.md and a set of scripts that check for fillable fields, extract field information and structure, fill fields or annotations, and render pages to images to check bounding boxes. A longer reference.md holds advanced features and JavaScript options. One practical warning: Unicode subscript and superscript characters render as black boxes in ReportLab, so the skill shows the XML markup to use instead. The SKILL.md is published under a proprietary licence.

When your agent uses it

  • Pulling text or tables out of a PDF report
  • Combining, splitting or rotating pages across several PDFs
  • Filling in a PDF form from structured data
  • Generating a new PDF document from content

Example prompts

  • “Extract every table from quarterly-report.pdf into CSV files.”
  • “Merge these three contracts into one PDF and add page numbers.”
  • “Fill in the W-9 form in ./forms with the details from my profile.”

Requirements

  • Python with pypdf, pdfplumber and reportlab

What it can do on your machine

Read from SKILL.md and the folder at commit 683bc88. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 8 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • qpdf
    • pdftotext

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Processing loads about 2k tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 249 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Proprietary) doesn't allow us to republish the file, so here is its outline and opening line. It has 249 words (~2,009 tokens).

“This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see REFERENCE.md. If you need to fill out a PDF form, read FORMS.md and follow its instructions.”

— opening of SKILL.md by anthropics, Proprietary
name
pdf
license
Proprietary. LICENSE.txt has complete terms

Read the full SKILL.md on GitHub

Files

SKILL.md and 11 other files (scripts) in skills/pdf of anthropics/skills.

  • SKILL.md
  • LICENSE.txt
  • forms.md
  • reference.md
  • scripts/check_bounding_boxes.py
  • scripts/check_fillable_fields.py
  • scripts/convert_pdf_to_images.py
  • scripts/create_validation_image.py
  • scripts/extract_form_field_info.py
  • scripts/extract_form_structure.py
  • scripts/fill_fillable_fields.py
  • scripts/fill_pdf_form_with_annotations.py

Open the folder on GitHubat commit 683bc88

Used in at least 42 other repositories

We found 62 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 48 other GitHub owners. This page covers the copy in anthropics/skills, which our catalogue first saw on October 7, 2026.

…and 12 more copies not listed here.

Compare with similar skills

PDF Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Processing this skillanthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing with PythonHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0
PDF ToolkitTokenRhythm/opensquilla7.1k—~1.9kAutomated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0
PDF Processing Guideagentscope-ai/QwenPaw35k—~1.8kAutomated safety check: PassProprietary

Similar skills

  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    agentscope-ai/QwenPaw

    Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

    35k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Toolkit

    telagod/code-abyss

    Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.

    243 GitHub stars~532 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes

More from anthropics/skills

All 16 skills in this repo
  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Auto-check passed
  • Web Artifacts Builder

    anthropics/skills

    Official

    Builds multi-component claude.ai HTML artifacts as a small React, TypeScript and Tailwind project, then bundles it into one shareable HTML file.

    180k GitHub starsUsed in 41 repos~769 tokens
    Auto-check passed
  • Official

    Creates original generative art in two steps: a written algorithmic philosophy, then a p5.js sketch with seeded randomness and an interactive viewer for exploring parameters.

    180k GitHub starsUsed in 38 repos~4.9k tokens
    Auto-check passed
  • PowerPoint Decks

    anthropics/skills

    Official

    Creates, edits, reads and validates .pptx and .potx files, using pptxgenjs for new decks and direct XML edits for existing ones, with helper scripts for thumbnails and checks.

    180k GitHub starsUsed in 7 repos~5.2k tokens
    Auto-check passed
  • Canvas Design

    anthropics/skills

    Official

    Creates original posters and static art as PNG or PDF by first writing a short design philosophy, then expressing it visually on a canvas.

    180k GitHub starsUsed in 52 repos~3k tokens
    Auto-check passed
  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Auto-check passed

Questions about PDF Processing

What does PDF Processing do?

Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR. This skill gives the agent working recipes for PDF tasks, mainly with Python libraries. pypdf covers merging, splitting, rotating pages and reading metadata; pdfplumber pulls out text with its layout and tables, which can go straight into pandas; reportlab creates new PDFs, from a single canvas page to multi-page documents built from paragraphs.

When should I use PDF Processing?

PDF Processing fits situations like: pulling text or tables out of a PDF report; combining, splitting or rotating pages across several PDFs; filling in a PDF form from structured data; generating a new PDF document from content.

How do I install PDF Processing in Claude Code?

Run `npx skills add anthropics/skills --skill pdf -a claude-code`. Or copy the skill folder (skills/pdf in anthropics/skills) into .claude/skills/pdf in your project. Claude Code loads it when a task matches its description.

How do I install PDF Processing in Codex?

Run `npx skills add anthropics/skills --skill pdf -a codex`. Or copy the skill folder (skills/pdf in anthropics/skills) into .agents/skills/pdf in your project. Codex loads it when a task matches its description.

Can I use PDF Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anthropics/skills --skill pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf, .gemini/skills/pdf, .github/skills/pdf and .opencode/skills/pdf in your project.

What does PDF Processing need to run?

Going by SKILL.md and its folder, PDF Processing needs Python for the scripts in its folder and the command-line tools its instructions call (qpdf and pdftotext). Our summary lists: Python with pypdf, pdfplumber and reportlab.

Does PDF Processing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF Processing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF Processing use?

PDF Processing carries a proprietary licence (declared in SKILL.md). It is published on GitHub, but it is not open source: read the licence file before using or sharing it.

How many tokens does PDF Processing use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Processing?

Skills that share tags, products or a category with PDF Processing: PDF Processing with Python (HKUDS/DeepTutor, 41k stars), PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars), PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars) and PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Processing?

anthropics (a GitHub organization, an official publisher) maintains it in anthropics/skills, which has 180,076 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 5, 2026.

Source: anthropics/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.