Picks the right Python library for PDF jobs, with examples for text and table extraction, merging, splitting and page extraction, plus OCR and memory fixes.

MITAuto-check passedDocuments & Office

Install PDF Library Guide

skills CLI
$ npx skills add FareedKhan-dev/claude-code-from-scratch --skill pdf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FareedKhan-dev/claude-code-from-scratch pdf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FareedKhan-dev/claude-code-from-scratch.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf .claude/skills/pdf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf
GitHub stars
298
Token cost
~785 tokens
SKILL.md length
109 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Picks the right Python library for PDF jobs, with examples for text and table extraction, merging, splitting and page extraction, plus OCR and memory fixes.

  • Extracting text or tables from a PDF report
  • SKILL.md covers When to use this skill, Library decision tree, Install and Extract text — pdfplumber, plus 5 more sections
  • Calls pip
  • Merging several PDFs into one file or splitting one apart

What it does

A short working guide for PDF tasks: extracting text or tables, merging, splitting, filling forms, adding watermarks or annotations, converting and searching. A decision tree maps each need to a library, with pdfplumber for text and tables from normal PDFs and pypdf for merging, splitting and page extraction, and code examples show each operation.

A troubleshooting section covers garbled or empty text from scanned files, which calls for OCR with pytesseract and pdf2image, tables that extract badly and can be tuned through pdfplumber table settings, and very large PDFs, which should be processed one page at a time. Installation is a pip install of pdfplumber and pypdf, with the OCR packages added when needed.

When your agent uses it

  • Extracting text or tables from a PDF report
  • Merging several PDFs into one file or splitting one apart
  • Reading a scanned PDF that needs OCR

Example prompts

  • “Extract the tables from ./reports/annual.pdf into CSV files.”
  • “Merge cover.pdf, body.pdf and appendix.pdf into one document.”
  • “This scanned contract returns empty text. OCR it and list its clauses.”

Requirements

  • Python with pdfplumber and pypdf
  • pytesseract and pdf2image for scanned PDFs

What it can do on your machine

Read from SKILL.md and the folder at commit fb9709e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Library Guide loads about 785 tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 109 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~785

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FareedKhan-dev/claude-code-from-scratch at commit fb9709e, republished under its MIT licence (© FareedKhan-dev). 109 words, ~785 tokens.

Download SKILL.mdSave it as .claude/skills/pdf/SKILL.md (or your agent's skills folder).
name
pdf
description
Use when working with PDF files — reading, extracting text or tables, merging, splitting, or filling forms. Provides correct library choices and common patterns.

PDF Skill

When to use this skill

Load when the user asks you to:

  • Extract text or tables from a PDF
  • Merge or split PDF files
  • Fill in a PDF form
  • Add watermarks or annotations
  • Convert PDF to another format
  • Search for content inside a PDF

Library decision tree

Need to extract text from a normal (not scanned) PDF?
    → pdfplumber  (best text + table extraction)

Need to manipulate pages (merge, split, rotate)?
    → pypdf  (formerly PyPDF2)

Need to fill form fields?
    → pypdf  with writer.update_page_form_field_values()

PDF is scanned (images of pages, no selectable text)?
    → pytesseract + pdf2image  (OCR pipeline)

Need to create a PDF from scratch?
    → reportlab  or  fpdf2

Install

bash
pip install pdfplumber pypdf
# For OCR:
pip install pytesseract pdf2image
# Also requires: tesseract-ocr (system package) and poppler-utils

Extract text — pdfplumber

python
import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    for page in pdf.pages:
        text = page.extract_text()
        if text:
            print(text)

Extract tables — pdfplumber

python
import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    for page in pdf.pages:
        tables = page.extract_tables()
        for table in tables:
            for row in table:
                print(row)

Merge PDFs — pypdf

python
from pypdf import PdfWriter

writer = PdfWriter()
for filename in ["part1.pdf", "part2.pdf", "part3.pdf"]:
    writer.append(filename)
with open("merged.pdf", "wb") as f:
    writer.write(f)

Split PDF — pypdf

python
from pypdf import PdfReader, PdfWriter

reader = PdfReader("document.pdf")
for i, page in enumerate(reader.pages):
    writer = PdfWriter()
    writer.add_page(page)
    with open(f"page_{i+1}.pdf", "wb") as f:
        writer.write(f)

Extract specific pages — pypdf

python
from pypdf import PdfReader, PdfWriter

reader = PdfReader("document.pdf")
writer = PdfWriter()
for page_num in [0, 2, 4]:          # 0-indexed
    writer.add_page(reader.pages[page_num])
with open("selected.pdf", "wb") as f:
    writer.write(f)

Common issues

Text comes back garbled or empty The PDF may use a non-standard encoding or be scanned. Try OCR:

python
# pip install pdf2image pytesseract
from pdf2image import convert_from_path
import pytesseract

images = convert_from_path("scanned.pdf")
for img in images:
    print(pytesseract.image_to_string(img))

Tables not extracting correctly Try adjusting the extraction strategy:

python
page.extract_table(table_settings={"vertical_strategy": "lines",
                                    "horizontal_strategy": "lines"})

Large PDF crashes memory Process page by page, don't load entire PDF into memory:

python
with pdfplumber.open("large.pdf") as pdf:
    for page in pdf.pages:
        process(page.extract_text())   # one page at a time

© FareedKhan-dev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/pdf of FareedKhan-dev/claude-code-from-scratch.

Open the folder on GitHubat commit fb9709e

Compare with similar skills

PDF Library Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Library Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Library Guide this skillFareedKhan-dev/claude-code-from-scratch298—~785Automated safety check: PassMIT
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing with PythonHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0
PDF ToolkitTokenRhythm/opensquilla7.1k—~1.9kAutomated safety check: PassApache-2.0
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    agentscope-ai/QwenPaw

    Handles PDF tasks with Python libraries and command-line tools: extract text and tables, merge, split, rotate, create, fill forms and more.

    35k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed

More from FareedKhan-dev/claude-code-from-scratch

  • Agent Harness Builder

    FareedKhan-dev/claude-code-from-scratch

    Gives patterns, a tool design checklist and an architecture decision tree for building agent harnesses, tools and multi-agent setups around a model.

    298 GitHub stars~1.1k tokensUpdated 6 mo ago
    Auto-check passed
  • Structured Code Review

    FareedKhan-dev/claude-code-from-scratch

    Gives the agent a five-step review routine that reads the full file first, labels each finding as bug, security, performance, style or suggestion, and ends with a summary.

    298 GitHub stars~809 tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about PDF Library Guide

What does PDF Library Guide do?

Picks the right Python library for PDF jobs, with examples for text and table extraction, merging, splitting and page extraction, plus OCR and memory fixes. A short working guide for PDF tasks: extracting text or tables, merging, splitting, filling forms, adding watermarks or annotations, converting and searching. A decision tree maps each need to a library, with pdfplumber for text and tables from normal PDFs and pypdf for merging, splitting and page extraction, and code examples show each operation.

When should I use PDF Library Guide?

PDF Library Guide fits situations like: extracting text or tables from a PDF report; merging several PDFs into one file or splitting one apart; reading a scanned PDF that needs OCR.

How do I install PDF Library Guide in Claude Code?

Run `npx skills add FareedKhan-dev/claude-code-from-scratch --skill pdf -a claude-code`. Or copy the skill folder (skills/pdf in FareedKhan-dev/claude-code-from-scratch) into .claude/skills/pdf in your project. Claude Code loads it when a task matches its description.

How do I install PDF Library Guide in Codex?

Run `npx skills add FareedKhan-dev/claude-code-from-scratch --skill pdf -a codex`. Or copy the skill folder (skills/pdf in FareedKhan-dev/claude-code-from-scratch) into .agents/skills/pdf in your project. Codex loads it when a task matches its description.

Can I use PDF Library Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FareedKhan-dev/claude-code-from-scratch --skill pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf, .gemini/skills/pdf, .github/skills/pdf and .opencode/skills/pdf in your project.

What does PDF Library Guide need to run?

Going by SKILL.md and its folder, PDF Library Guide needs the command-line tools its instructions call (pip). Our summary lists: Python with pdfplumber and pypdf; pytesseract and pdf2image for scanned PDFs.

Does PDF Library Guide access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Library Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Library Guide use?

PDF Library Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Library Guide use?

About 785 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Library Guide?

Skills that share tags, products or a category with PDF Library Guide: PDF Processing (anthropics/skills, 180k stars), PDF Processing with Python (HKUDS/DeepTutor, 41k stars), PDF Toolkit (TokenRhythm/opensquilla, 7.1k stars) and PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Library Guide?

FareedKhan-dev (a GitHub user) maintains it in FareedKhan-dev/claude-code-from-scratch, which has 298 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on April 5, 2026.

Source: FareedKhan-dev/claude-code-from-scratch on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.