Agent skill

PDF Reader

by espennilsen in espennilsen/pi

Read and extract content from PDF files — text, tables, metadata, and images.

MITAuto-check passedDocuments & Office

Install PDF Reader

skills CLI
$ npx skills add espennilsen/pi --skill pdf-reader -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install espennilsen/pi pdf-reader --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/espennilsen/pi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf-reader .claude/skills/pdf-reader && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-reader
GitHub stars
122
Token cost
~1.6k tokens
SKILL.md length
499 words
Files
2 (incl. scripts)
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Read and extract content from PDF files — text, tables, metadata, and images.

  • Works in 5 steps: Get the PDF → Choose Extraction Method → Handle Large PDFs → …
  • Asked to read a PDF
  • SKILL.md covers Quick Reference, Workflow, Decision Tree and Tips
  • Runs Python scripts from its folder; calls pdftotext, python3 and curl

What it does

PDF Reader is an agent skill from espennilsen/pi. Read and extract content from PDF files — text, tables, metadata, and images. Use when asked to read a PDF, extract text from a PDF, summarize a PDF, analyze a PDF document, get tables from a PDF, or check PDF metadata. Also triggers on "open this PDF", "what does this PDF say", "parse PDF", "PDF to text", or when a .pdf file path or URL is provided.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/extract.py`).

It sits in Documents & Office, covering PDF and Document parsing. It works with Python and pypdf. The licence is MIT.

When your agent uses it

  • Asked to read a PDF
  • Extract text from a PDF
  • Summarize a PDF
  • Analyze a PDF document

Example prompts

  • “open this PDF”
  • “what does this PDF say”
  • “parse PDF”
  • “/pdf-reader”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Get the PDF
  2. Choose Extraction Method
  3. Handle Large PDFs
  4. Handle Scanned PDFs (OCR)
  5. Handle Other Edge Cases

What it can do on your machine

Read from SKILL.md and the folder at commit 79d019b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pdftotext
    • python3
    • curl
    • tesseract
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Reader loads about 1.6k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 499 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from espennilsen/pi at commit 79d019b, republished under its MIT licence (© espennilsen). 499 words, ~1,575 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-reader/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-reader
description
Read and extract content from PDF files — text, tables, metadata, and images. Use when asked to read a PDF, extract text from a PDF, summarize a PDF, analyze a PDF document, get tables from a PDF, or check PDF metadata. Also triggers on "open this PDF", "what does this PDF say", "parse PDF", "PDF to text", or when a .pdf file path or URL is provided.

PDF Reader

Extract content from PDF files using pdftotext (Poppler) for text and pdfplumber (Python) for tables and structured extraction.

Quick Reference

TaskToolCommand
Full textpdftotextpdftotext file.pdf -
Text with layoutpdftotextpdftotext -layout file.pdf -
Specific pagespdftotextpdftotext -f 3 -l 5 file.pdf -
Tablespdfplumberpython3 scripts/extract.py tables file.pdf
Metadatapdfinfopdfinfo file.pdf
Page countpdfinfopdfinfo file.pdf | grep Pages
Imagespdfimagespdfimages -list file.pdf
Fontspdffontspdffonts file.pdf
OCR (scanned PDF)tesseractpython3 scripts/extract.py ocr file.pdf
Smart text + OCRextract.pypython3 scripts/extract.py text file.pdf
Quick surveyextract.pypython3 scripts/extract.py scan file.pdf

Workflow

Step 1: Get the PDF

If the user provides a URL, download it first:

bash
curl -sL "URL" -o /tmp/document.pdf

Verify it's a valid PDF:

bash
file /tmp/document.pdf  # should say "PDF document"
pdfinfo /tmp/document.pdf  # metadata + page count
Step 2: Choose Extraction Method

Plain text (most cases):

bash
pdftotext file.pdf -

This pipes output to stdout. For large PDFs, use page ranges:

bash
pdftotext -f 1 -l 10 file.pdf -    # pages 1-10

Layout-preserving text (columns, formatted docs):

bash
pdftotext -layout file.pdf -

Use -layout when the PDF has multi-column layouts, tables rendered as text, or precise spacing that matters.

Tables (structured data):

bash
python3 scripts/extract.py tables file.pdf

Or inline with pdfplumber:

python
import pdfplumber

pdf = pdfplumber.open("file.pdf")
for i, page in enumerate(pdf.pages):
    tables = page.extract_tables()
    for table in tables:
        print(f"\n--- Table on page {i+1} ---")
        for row in table:
            print(" | ".join(str(cell or "") for cell in row))
pdf.close()

Metadata only:

bash
pdfinfo file.pdf

Returns: title, author, creator, producer, page count, page size, dates.

Step 3: Handle Large PDFs

For PDFs over ~50 pages, don't dump everything at once:

  1. Get page count: pdfinfo file.pdf | grep Pages
  2. Extract in chunks: pdftotext -f 1 -l 20 file.pdf -
  3. Process chunk, then continue: pdftotext -f 21 -l 40 file.pdf -

For targeted extraction (searching for specific content):

bash
# Extract all text, grep for relevant sections
pdftotext file.pdf - | grep -n -i "keyword"

# Then extract the specific page range
pdftotext -f PAGE -l PAGE file.pdf -
Step 4: Handle Scanned PDFs (OCR)

If pdftotext returns empty or garbled output, the PDF is likely scanned.

Detection:

bash
python3 scripts/extract.py scan file.pdf   # reports scanned pages
pdffonts file.pdf                           # empty = image-based

Smart extraction (auto-fallback):

text mode automatically detects scanned pages and OCRs them:

bash
python3 scripts/extract.py text file.pdf

Pages with selectable text extract normally. Pages without selectable text fall back to OCR via Tesseract. No manual detection needed.

Force OCR on all pages:

bash
python3 scripts/extract.py ocr file.pdf
python3 scripts/extract.py ocr file.pdf --pages 1-5
python3 scripts/extract.py ocr file.pdf --dpi 400        # higher quality
python3 scripts/extract.py ocr file.pdf --lang eng+nor   # multi-language

OCR options:

  • --dpi 300 — resolution for page-to-image conversion (default: 300, higher = slower but better)
  • --lang eng — Tesseract language pack (default: eng). Use + for multiple: eng+nor+deu
  • --pages 1-5 — limit to specific pages (recommended for large PDFs)

Available language packs:

bash
tesseract --list-langs

Install additional languages via Homebrew:

bash
brew install tesseract-lang    # all languages
Show full SKILL.md (173 more words)Show less
Step 5: Handle Other Edge Cases

Mixed PDFs (some pages scanned, some not):

Just use text mode — it handles mixed PDFs automatically:

bash
python3 scripts/extract.py text file.pdf

Selectable pages extract instantly, scanned pages get OCR'd. The output is tagged so you know which pages used OCR.

Password-protected PDFs:

bash
pdftotext -upw "password" file.pdf -   # user password
pdftotext -opw "password" file.pdf -   # owner password

Encoding issues (garbled output):

bash
pdftotext -enc UTF-8 file.pdf -

Extract images:

bash
pdfimages -png file.pdf /tmp/images/img   # extracts as PNG
pdfimages -list file.pdf                  # list images without extracting

Decision Tree

Is it a URL? → curl -sL "URL" -o /tmp/doc.pdf
            ↓
Run: python3 scripts/extract.py scan file.pdf
            ↓
All pages have selectable text?
  YES → pdftotext file.pdf -           (fast, simple)
  NO  → python3 scripts/extract.py text file.pdf  (auto OCR fallback)
            ↓
Need tables?
  YES → python3 scripts/extract.py tables file.pdf

Tips

  • Start with scan on unknown PDFs — it reports pages, tables, scanned detection, and a preview
  • pdftotext is fastest for normal PDFs — try it first
  • Use -layout for multi-column documents (academic papers, reports)
  • pdfplumber is better for tables — it understands cell boundaries
  • text mode auto-detects scanned pages and OCRs only those — preferred over raw pdftotext for unknown PDFs
  • ocr mode is for forcing OCR on everything (useful when text extraction gives garbled output despite appearing selectable)
  • Higher --dpi gives better OCR accuracy but is slower (300 is a good default, 400+ for small text)
  • For PDFs from URLs, always download to /tmp/ first — don't pipe curl to tools
  • Large PDF text output may exceed context limits — use --pages to extract in ranges

© espennilsen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/pdf-reader of espennilsen/pi.

  • SKILL.md
  • scripts/extract.py

Open the folder on GitHubat commit 79d019b

Compare with similar skills

PDF Reader next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Reader compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Reader this skillespennilsen/pi122—~1.6kAutomated safety check: PassMIT
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0
PDF Explorexuzhougeng/wisp-science1k—~1.2kAutomated safety check: PassApache-2.0
PDF Processinganthropics/skills180k48 repos~2kAutomated safety check: PassProprietary
PDF Processing with PythonHKUDS/DeepTutor41k—~2.7kAutomated safety check: PassApache-2.0

Similar skills

  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Explore

    xuzhougeng/wisp-science

    A skill your agent uses when the user has attached a PDF, paper, report, or other document and the answer needs its content: summarize a section, compare sections, read specific pages, check the…

    1k GitHub stars~1.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 48 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    TokenRhythm/opensquilla

    Deterministic PDF operations through bundled scripts: extract text and tables, merge files or page ranges, split by range, fill form fields and build PDFs from data.

    7.1k GitHub stars~1.9k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed

More from espennilsen/pi

All 36 skills in this repo
  • GitHub

    espennilsen/pi

    Interact with GitHub repos, PRs, issues, CI, and notifications via the pi-github extension commands and gh CLI.

    122 GitHub stars~1k tokensUpdated 15 days ago
    Auto-check passed
  • Skill Creator

    espennilsen/pi

    Create, review, and improve skills for Pi agents. An agent skill from espennilsen/pi.

    122 GitHub stars~2.1k tokensUpdated 15 days ago
    Auto-check passed
  • Dry Code Review

    espennilsen/pi

    Perform a comprehensive DRY (Don't Repeat Yourself) code review on a codebase.

    122 GitHub stars~1.7k tokensUpdated 15 days ago
    Auto-check passed
  • Extract Design System

    espennilsen/pi

    Reverse-engineer a design system from a live website (public URL or localhost).

    122 GitHub stars~2k tokensUpdated 15 days ago
    Auto-check passed
  • Google Workspace

    espennilsen/pi

    Manage Google Workspace via the gws CLI — Drive, Gmail, Sheets, Docs, Slides, People, Chat, Meet, Forms, and cross-service workflows.

    122 GitHub stars~3.4k tokensUpdated 15 days ago
    Auto-check passed
  • Herdr Operations

    espennilsen/pi

    A skill your agent uses when inspecting or operating Herdr sessions, workspaces, tabs, panes, agents, terminal output, agent messaging, or waits.

    122 GitHub stars~525 tokensUpdated 15 days ago
    Auto-check passed

Works with

Questions about PDF Reader

What does PDF Reader do?

Read and extract content from PDF files — text, tables, metadata, and images. PDF Reader is an agent skill from espennilsen/pi. Read and extract content from PDF files — text, tables, metadata, and images.

When should I use PDF Reader?

PDF Reader fits situations like: asked to read a PDF; extract text from a PDF; summarize a PDF; analyze a PDF document.

How do I install PDF Reader in Claude Code?

Run `npx skills add espennilsen/pi --skill pdf-reader -a claude-code`. Or copy the skill folder (skills/pdf-reader in espennilsen/pi) into .claude/skills/pdf-reader in your project. Claude Code loads it when a task matches its description.

How do I install PDF Reader in Codex?

Run `npx skills add espennilsen/pi --skill pdf-reader -a codex`. Or copy the skill folder (skills/pdf-reader in espennilsen/pi) into .agents/skills/pdf-reader in your project. Codex loads it when a task matches its description.

Can I use PDF Reader in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add espennilsen/pi --skill pdf-reader -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-reader, .gemini/skills/pdf-reader, .github/skills/pdf-reader and .opencode/skills/pdf-reader in your project.

What does PDF Reader need to run?

Going by SKILL.md and its folder, PDF Reader needs Python for the scripts in its folder and the command-line tools its instructions call (pdftotext, python3, curl, tesseract and brew). Our summary lists: Python 3.

Does PDF Reader access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Reader safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF Reader use?

PDF Reader is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Reader use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Reader?

Skills that share tags, products or a category with PDF Reader: PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars), PDF Explore (xuzhougeng/wisp-science, 1k stars) and PDF Processing (anthropics/skills, 180k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Reader?

espennilsen (a GitHub user) maintains it in espennilsen/pi, which has 122 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 21, 2026.

Source: espennilsen/pi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.