Agent skill

PDF Processing with Python

by HKUDS in HKUDS/DeepTutor

Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

Apache-2.0Auto-check passedDocuments & Office

Install PDF Processing with Python

skills CLI
$ npx skills add HKUDS/DeepTutor --skill pdf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/DeepTutor pdf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/DeepTutor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/deeptutor/skills/builtin/pdf .claude/skills/pdf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf
GitHub stars
41k
Token cost
~2.7k tokens
SKILL.md length
608 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

  • Pulling text or tables out of an uploaded PDF
  • SKILL.md covers Extract text and tables…, Scanned / image-only PDFs (be…, Merge / split / rotate / crop… and Create PDFs (reportlab), plus 2 more sections
  • Calls pip
  • Merging, splitting, rotating or watermarking PDF files

What it does

The skill picks a Python library for each PDF job and has the agent run complete scripts in a sandbox. pdfplumber handles text, table, layout and word-coordinate extraction. pypdf covers quick text, merging, splitting, rotating, cropping, watermarking, encryption, decryption and metadata, and fills fillable AcroForm fields. reportlab creates new PDFs from scratch, and flat forms are filled with an annotation overlay.

Tables can be written to Excel with openpyxl, one worksheet per table, and messy tables can be handled by passing line strategies or cropping a region first. If you ask for a size such as 500 words, the agent counts it in the output. Scanned PDFs are treated honestly: with no OCR engine and no network, the agent says the text cannot be recovered instead of making content up or trying to install tools.

When your agent uses it

  • Pulling text or tables out of an uploaded PDF
  • Merging, splitting, rotating or watermarking PDF files
  • Filling in a fillable or flat PDF form
  • Generating a new PDF report from scratch

Example prompts

  • “Extract every table from ./reports/q3.pdf into an Excel file with one sheet per table.”
  • “Merge invoice-1.pdf, invoice-2.pdf and invoice-3.pdf and add a CONFIDENTIAL watermark.”
  • “Fill in the W-9 form in ./forms/w9.pdf with my details.”
  • “Create a one-page PDF memo of about 500 words on our new office hours.”

Requirements

  • A Python sandbox with pdfplumber, pypdf and reportlab
  • openpyxl for exporting tables to Excel

What it can do on your machine

Read from SKILL.md and the folder at commit 6cf793b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Processing with Python loads about 2.7k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 608 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/DeepTutor at commit 6cf793b, republished under its Apache-2.0 licence (© HKUDS). 608 words, ~2,664 tokens.

Download SKILL.mdSave it as .claude/skills/pdf/SKILL.md (or your agent's skills folder).
name
pdf
description
Read, extract (text/tables), create, merge/split/rotate, watermark, encrypt, fill, and render-to-image .pdf files. Use whenever the user uploads a .pdf or asks to produce, edit, or pull data out of one.
tags
tool, office
requires.sandbox
shell

PDF

Work PDFs in the sandbox with preinstalled Python libs. Pick the library by task:

  • Extract text/tables/layout/word-coordinates → pdfplumber; quick raw text or page ops → pypdf.
  • Merge / split / rotate / crop / watermark / encrypt / metadata → pypdf.
  • Fill forms → pypdf (fillable AcroForm fields) or annotation overlay (flat forms).
  • Create from scratch → reportlab.

Use exec with complete Python source (language: python). Prefer creating, reopening, and validating the PDF in one call; later calls can revise the same relative filename. Follow the turn's User workspace instructions for locating inputs, output boundaries, and presenting the finished file. Preserve an explicitly requested quantity (such as 500 words) and verify the count in the output before finishing. If execution fails or the artifact is missing, diagnose stderr/root cause and change strategy; do not retry identical code or reduce the requested scope without asking.

Extract text and tables (pdfplumber)

python
import pdfplumber

with pdfplumber.open("in.pdf") as pdf:
    for i, page in enumerate(pdf.pages, 1):
        print(f"--- page {i} ---")
        print(page.extract_text() or "")  # layout-aware text
        for t in page.extract_tables():  # list of tables; each is list[row]
            for row in t:
                print(row)

Tables → Excel (one worksheet per table):

python
import pdfplumber
from openpyxl import Workbook

workbook = Workbook()
workbook.remove(workbook.active)
table_number = 0
with pdfplumber.open("in.pdf") as pdf:
    for page_number, page in enumerate(pdf.pages, 1):
        for table in page.extract_tables():
            if not table:
                continue
            table_number += 1
            sheet = workbook.create_sheet(f"p{page_number}_table{table_number}"[:31])
            for row in table:
                sheet.append([cell or "" for cell in row])
if table_number:
    workbook.save("tables.xlsx")

Messy tables: pass strategies, or crop a region with page.within_bbox((x0, top, x1, bottom)) first:

python
ts = {
    "vertical_strategy": "lines",
    "horizontal_strategy": "lines",
    "snap_tolerance": 3,
    "intersection_tolerance": 15,
}
page.extract_tables(ts)

For very large PDFs where you only need raw text, pypdf's page.extract_text() is lighter.

Scanned / image-only PDFs (be honest)

If extract_text() returns empty or garbage (e.g. (cid:NN) runs) the page is scanned. No OCR engine (tesseract) is installed and network is off, so you cannot recover that text. Say so plainly and stop — do not fabricate content or attempt pip install.

Merge / split / rotate / crop / metadata (pypdf)

python
from pypdf import PdfReader, PdfWriter

# Merge
w = PdfWriter()
for f in ["a.pdf", "b.pdf"]:
    for p in PdfReader(f).pages:
        w.add_page(p)
w.write("merged.pdf")

# Split: one file per page
r = PdfReader("in.pdf")
for i, p in enumerate(r.pages, 1):
    w = PdfWriter()
    w.add_page(p)
    w.write(f"page_{i}.pdf")

# Rotate page 0 by 90 degrees clockwise
r = PdfReader("in.pdf")
w = PdfWriter()
r.pages[0].rotate(90)
w.add_page(r.pages[0])
w.write("rotated.pdf")
  • Metadata: PdfReader("in.pdf").metadata (.title, .author, ...).
  • Crop: set page.mediabox.left/bottom/right/top (points, origin y=0 at bottom).
  • Encrypt: w = PdfWriter(clone_from=PdfReader("in.pdf")); w.encrypt("userpw", "ownerpw"); w.write("enc.pdf").
  • Decrypt: r = PdfReader("enc.pdf"); r.decrypt("pw") if r.is_encrypted, then read/copy pages.

Watermark (stamp one page over every page):

python
from pypdf import PdfReader, PdfWriter

wm = PdfReader("stamp.pdf").pages[0]
r = PdfReader("in.pdf")
w = PdfWriter()
for p in r.pages:
    p.merge_page(wm)
    w.add_page(p)
w.write("stamped.pdf")

Create PDFs (reportlab)

Flowing document (preferred for text/reports/tables — handles pagination):

python
from reportlab.lib.pagesizes import letter
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.lib import colors
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, Table, TableStyle

styles = getSampleStyleSheet()
story = [
    Paragraph("Report Title", styles["Title"]),
    Spacer(1, 12),
    Paragraph("Body text. " * 20, styles["Normal"]),
]
data = [["Product", "Q1", "Q2"], ["Widgets", "120", "135"]]
tbl = Table(data)
tbl.setStyle(
    TableStyle(
        [
            ("BACKGROUND", (0, 0), (-1, 0), colors.grey),
            ("TEXTCOLOR", (0, 0), (-1, 0), colors.whitesmoke),
            ("GRID", (0, 0), (-1, -1), 0.5, colors.black),
        ]
    )
)
story += [Spacer(1, 12), tbl]
SimpleDocTemplate("out.pdf", pagesize=letter).build(story)

Absolute placement (labels at fixed coordinates): use canvas.Canvas("out.pdf", pagesize=letter), c.drawString(x, y, "...") (origin bottom-left, points), c.showPage() per page, c.save().

Non-Latin text (Chinese / Japanese / Korean, Cyrillic, …)

reportlab's built-in fonts (Helvetica/Times/Courier) carry zero CJK glyphs, so any 中文/日本語/한국어 renders as empty boxes (□) baked permanently into the PDF. reportlab never auto-discovers system fonts — you MUST register a font that has the glyphs and set it on every style. Whenever the document may contain non-Latin text, register a CJK font first (it also covers Latin, so it is safe to use as the only font):

python
import os
from reportlab.pdfbase import pdfmetrics
from reportlab.pdfbase.ttfonts import TTFont


def register_cjk_font(name="CJK"):
    # TrueType ONLY — reportlab cannot embed CFF/OpenType outlines, so a .otf
    # like Noto Sans CJK fails with "postscript outlines are not supported".
    for path in [
        "/usr/share/fonts/truetype/wqy/wqy-zenhei.ttc",  # Linux sandbox (fonts-wqy-zenhei)
        "/usr/share/fonts/truetype/wqy/wqy-microhei.ttc",
        "/System/Library/Fonts/STHeiti Light.ttc",  # macOS
        "/System/Library/Fonts/Hiragino Sans GB.ttc",
        "/System/Library/Fonts/Supplemental/Songti.ttc",
        "/System/Library/Fonts/Supplemental/Arial Unicode.ttf",
        "C:/Windows/Fonts/msyh.ttc",  # Windows
    ]:
        if os.path.exists(path):
            try:
                pdfmetrics.registerFont(TTFont(name, path, subfontIndex=0))
                return name
            except Exception:
                continue
    raise RuntimeError("No CJK-capable TrueType font found — do not emit tofu; say so.")


font = register_cjk_font()
styles = getSampleStyleSheet()
for s in styles.byName.values():  # make the CJK font the default everywhere
    s.fontName = font
# Tables don't read the stylesheet — set the font in the TableStyle too:
#   ("FONTNAME", (0, 0), (-1, -1), font)
# Canvas: c.setFont(font, size) before every drawString.

If register_cjk_font raises (no font on the host), do not ship a tofu PDF — tell the user the sandbox lacks a CJK font instead of producing garbage.

Gotcha: even with a good font, reportlab still needs markup for subscripts/superscripts. In Paragraph use Paragraph("H<sub>2</sub>O", styles["Normal"]), x<super>2</super>.

Markdown/HTML → PDF needs an external converter (soffice/pandoc) that is usually absent — command -v soffice / command -v pandoc and degrade to building the PDF directly with reportlab if neither is present.

Show full SKILL.md (173 more words)Show less

Fill forms (pypdf)

First detect whether the PDF has real fillable (AcroForm) fields:

python
from pypdf import PdfReader

fields = PdfReader("form.pdf").get_fields()
print("fillable" if fields else "flat (no fields)")

Fillable — inspect field names/types, then fill and write:

python
from pypdf import PdfReader, PdfWriter

r = PdfReader("form.pdf")
for name, f in r.get_fields().items():
    print(name, f.get("/FT"), f.get("/_States_"))  # /Tx text, /Btn checkbox/radio, /Ch choice

w = PdfWriter(clone_from=r)
values = {"first_name": "Bart", "agree": "/Yes"}  # checkbox/radio: use its on-state, NOT True/False
for page in w.pages:
    w.update_page_form_field_values(page, values, auto_regenerate=False)
w.set_need_appearances_writer(True)  # force viewers to render the values
w.write("filled.pdf")

Checkbox/radio values are on-state strings, not booleans — read the field's /_States_ (e.g. /Yes, /On); /Off clears it.

Flat form (no fields) — overlay text with FreeText annotations at PDF coordinates. Get real coordinates from the layout with pdfplumber instead of guessing:

python
import pdfplumber

with pdfplumber.open("form.pdf") as pdf:
    pg = pdf.pages[0]
    for wd in pg.extract_words():  # each has x0, top, x1, bottom (TOP-left origin!)
        print(wd["text"], wd["x0"], wd["top"])
    for rc in pg.rects:  # small squares are likely checkboxes
        print("rect", rc["x0"], rc["top"], rc["x1"], rc["bottom"])

pdfplumber top is measured from the page top; pypdf rects are bottom-left, so convert: pdf_y = page_height - top. Place text just right of the matching label:

python
from pypdf import PdfReader, PdfWriter
from pypdf.annotations import FreeText

r = PdfReader("form.pdf")
w = PdfWriter()
w.append(r)
h = float(r.pages[0].mediabox.height)
top = 700  # pdfplumber 'top' of the label's row
w.add_annotation(
    page_number=0,
    annotation=FreeText(
        text="Smith",
        rect=(255, h - top - 14, 720, h - top),  # (x0, y0, x1, y1)
        font="Helvetica",
        font_size="10pt",
        font_color="000000",
        border_color=None,
        background_color=None,
    ),
)
w.write("filled.pdf")

Verify: re-open the output and re-read get_fields() values (fillable) or re-extract text (overlay) to confirm the values landed.

Page → image rendering (PyMuPDF)

PyMuPDF (imported as fitz, preinstalled) rasterizes pages — useful to inspect a PDF visually or to hand a page to an image-capable step. No external tools needed (poppler / pdf2image are absent; don't reach for them).

python
import fitz  # PyMuPDF

doc = fitz.open("in.pdf")
for i, page in enumerate(doc, 1):
    page.get_pixmap(dpi=150).save(f"page_{i}.png")  # higher dpi = sharper + larger

fitz also extracts text (page.get_text()) and can render a sub-region via page.get_pixmap(clip=fitz.Rect(x0, y0, x1, y1)). It does not OCR — a rendered scanned page is still just pixels (see Scanned PDFs above).

© HKUDS, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in deeptutor/skills/builtin/pdf of HKUDS/DeepTutor.

Open the folder on GitHubat commit 6cf793b

Compare with similar skills

PDF Processing with Python next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Processing with Python compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Processing with Python this skillHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF ToolkitTokenRhythm/opensquilla7.1k—~1.9kAutomated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0
PDF Processing Guideagentscope-ai/QwenPaw35k—~1.8kAutomated safety check: PassProprietary

Similar skills

  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    agentscope-ai/QwenPaw

    Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

    35k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Toolkit

    telagod/code-abyss

    Picks the right Python library or CLI tool for a PDF task, text and table extraction, merging, splitting, OCR, watermarking or form filling, and points to a matching recipe.

    243 GitHub stars~532 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes

More from HKUDS/DeepTutor

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Explains how to design and write DeepTutor skills: the SKILL.md anatomy, a trigger-focused description, concise content and progressive disclosure into reference files.

    41k GitHub stars~840 tokensUpdated today
    Auto-check passed
  • Excel Workbook Editor

    HKUDS/DeepTutor

    Reads, creates and edits Excel workbooks with openpyxl, including formulas, styles, charts and CSV or TSV tables, with advice on formula values openpyxl cannot compute.

    41k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about PDF Processing with Python

What does PDF Processing with Python do?

Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox. The skill picks a Python library for each PDF job and has the agent run complete scripts in a sandbox. pdfplumber handles text, table, layout and word-coordinate extraction.

When should I use PDF Processing with Python?

PDF Processing with Python fits situations like: pulling text or tables out of an uploaded PDF; merging, splitting, rotating or watermarking PDF files; filling in a fillable or flat PDF form; generating a new PDF report from scratch.

How do I install PDF Processing with Python in Claude Code?

Run `npx skills add HKUDS/DeepTutor --skill pdf -a claude-code`. Or copy the skill folder (deeptutor/skills/builtin/pdf in HKUDS/DeepTutor) into .claude/skills/pdf in your project. Claude Code loads it when a task matches its description.

How do I install PDF Processing with Python in Codex?

Run `npx skills add HKUDS/DeepTutor --skill pdf -a codex`. Or copy the skill folder (deeptutor/skills/builtin/pdf in HKUDS/DeepTutor) into .agents/skills/pdf in your project. Codex loads it when a task matches its description.

Can I use PDF Processing with Python in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/DeepTutor --skill pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf, .gemini/skills/pdf, .github/skills/pdf and .opencode/skills/pdf in your project.

What does PDF Processing with Python need to run?

Going by SKILL.md and its folder, PDF Processing with Python needs the command-line tools its instructions call (pip). Our summary lists: A Python sandbox with pdfplumber, pypdf and reportlab; openpyxl for exporting tables to Excel.

Does PDF Processing with Python access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Processing with Python safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Processing with Python use?

PDF Processing with Python is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Processing with Python use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Processing with Python?

Skills that share tags, products or a category with PDF Processing with Python: PDF Processing (anthropics/skills, 180k stars), PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars), PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars) and PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Processing with Python?

HKUDS (a GitHub organization) maintains it in HKUDS/DeepTutor, which has 40,905 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.

Source: HKUDS/DeepTutor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.