Agent skill

PDF Processing

by majiayu000 in majiayu000/claude-skill-registry

Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity.

MITAuto-check passedDocuments & Office

Install PDF Processing

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill pdf-processing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry pdf-processing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bash/pdf-processing .claude/skills/pdf-processing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-processing
GitHub stars
666
Used in
1 other repo
Token cost
~2.5k tokens
SKILL.md length
1,214 words
Files
2
Skills in repo
1,273
Repo updated
First seen
Licence
MIT

At a glance

Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity.

  • Works in 7 steps: Preserve and inspect → Select an available toolchain → Establish a visual and content baseline → …
  • Working with one
  • SKILL.md covers Inputs, Output contract, Workflow and Safety and permission boundaries, plus 2 more sections
  • Calls python3

What it does

PDF Processing is an agent skill from majiayu000/claude-skill-registry. Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity. Use when working with one or more .pdf files; converting documents to or from PDF; extracting text, tables, images, metadata, forms, or page ranges; applying true redactions or signatures; diagnosing malformed, encrypted, scanned, or inaccessible PDFs; or validating that a PDF transformation preserved the intended content and layout.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in Documents & Office, covering PDF. It works with Bash. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Working with one
  • More .pdf files
  • Converting documents to
  • Extracting text

Example prompts

  • “/pdf-processing”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Preserve and inspect
  2. Select an available toolchain
  3. Establish a visual and content baseline
  4. Apply the smallest transformation
  5. Verify structure and content
  6. Render and inspect
  7. Deliver with evidence

What it can do on your machine

Read from SKILL.md and the folder at commit 2d14a69. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Processing loads about 2.5k tokens when it runs. Until then it costs about 133 tokens; SKILL.md has 1,214 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~133
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 2d14a69, republished under its MIT licence (© majiayu000). 1,214 words, ~2,492 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-processing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-processing
description
Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity. Use when working with one or more .pdf files; converting documents to or from PDF; extracting text, tables, images, metadata, forms, or page ranges; applying true redactions or signatures; diagnosing malformed, encrypted, scanned, or inaccessible PDFs; or validating that a PDF transformation preserved the intended content and layout.

PDF Processing

Preserve the original, choose tools from what is actually available, and verify both document structure and rendered appearance.

Inputs

Collect or state:

  • Source file(s), requested operation, destination, naming convention, and page order/ranges.
  • Whether layout fidelity, searchable text, file size, accessibility, archival quality, or print output is the priority.
  • OCR language, table/image requirements, form fields, annotations, bookmarks, links, and metadata expectations.
  • Passwords or signing credentials supplied through an approved secure channel; never request private keys in chat.
  • Whether content is confidential and whether local-only processing is required.

Clarify ambiguous page references: human page labels may differ from zero- or one-based physical page numbers.

Output contract

Return:

  1. New output file(s) at explicit paths; never silently overwrite the source.
  2. Source and output SHA-256 hashes, byte sizes, page counts, and encryption status when determinable.
  3. A concise operation log: tools used, options, page mapping, OCR language, redactions, and metadata changes.
  4. Verification evidence for structure, content, and rendered appearance, plus limitations.
  5. Any password, signature, accessibility, font, OCR, or active-content caveats.

Do not claim exactness when the check was heuristic or a page could not be rendered.

Workflow

1. Preserve and inspect
  • Resolve exact input paths and calculate hashes before making changes.
  • Write to a new file or a temporary working copy. Never edit the sole source in place.
  • Run the dependency-free structural inspector from this skill directory:
bash
python3 scripts/inspect_pdf.py /path/to/input.pdf --pretty
python3 scripts/inspect_pdf.py /path/to/input.pdf --output /path/to/report.json --pretty

This default scan reads bytes without invoking a PDF parser. With --output, it refuses input aliases and non-regular destinations, then atomically creates or replaces the report via a sibling temporary file. Use --deep only after the file is trusted enough to parse, or after the untrusted-input sandbox profile in step 2 is verified. Declare the decision explicitly:

bash
python3 scripts/inspect_pdf.py /path/to/input.pdf --deep --trust-level trusted --pretty
python3 scripts/inspect_pdf.py /path/to/untrusted.pdf --deep --sandbox-profile-confirmed --pretty

Treat JavaScript, launch actions, automatic actions, embedded files, and rich media as inert hazards; do not activate them.

2. Select an available toolchain

Inventory installed tools and libraries before choosing a method. Read tool-routing.md for capability-based routing and fallback behavior. Prefer a tool that preserves the required features; a text extractor is not a page editor, and rasterization is not a faithful editable conversion.

For an untrusted PDF, parse or render only in a disposable, low-privilege sandbox with no secrets or network access; a read-only source mount; isolated temporary/output storage; CPU, memory, process, file-size, page, and time limits; and no host viewer, shell, URL-handler, or clipboard integration. Configure the parser/renderer to disable JavaScript, OpenAction/additional/launch actions, form submission, external URI fetching, attachment extraction/opening, rich media, and automatic font/resource downloads. Verify these controls for the exact tool/version. If the available parser or renderer cannot satisfy the profile, stop after raw-byte inspection and report the capability gap.

If no suitable dependency exists, explain the missing capability and propose an installation or alternate workflow. Do not install software or upload the PDF without permission.

3. Establish a visual and content baseline

Open or render a trusted original with a trusted local viewer before any layout-sensitive change. For an untrusted original, use only the verified sandbox profile from step 2; if it is unavailable, stop without parsing/rendering. Inspect representative pages and every page that will change. Record page count, dimensions/orientation, searchable text availability, form fields, annotations, bookmarks, attachments, encryption, and obvious font or image issues.

For scanned pages, determine whether OCR is needed. Preserve the original image layer unless the user explicitly requests destructive cleanup.

4. Apply the smallest transformation
  • Extract: preserve reading order uncertainty, page references, table structure, and OCR confidence. Do not invent missing text.
  • Merge/split/reorder/rotate: make the page mapping explicit and retain bookmarks, labels, forms, and metadata when required.
  • Create/convert: embed or substitute fonts deliberately; define page size, margins, links, headings, and accessibility expectations.
  • Fill/annotate: distinguish annotations from flattened page content and preserve an editable copy if useful.
  • Redact: use true object-level redaction, not a colored rectangle. Remove underlying text/images, relevant metadata, comments, attachments, and hidden layers as scoped. Save a fully rewritten new PDF with incremental saving disabled; do not retain prior revisions, original object streams, or appended historical bytes in the output.
  • Compress: measure visual loss and text/searchability changes; avoid lossy rasterization unless authorized.
  • Encrypt/sign: confirm algorithm, permissions, identity, and key handling. Any post-signature change can invalidate a signature.

Work on a temporary output, then move or copy it to the final requested path only after verification.

Show full SKILL.md (502 more words)Show less
5. Verify structure and content

Follow verification-checklist.md. At minimum:

  • Re-open the output with an independent reader or parser when available.
  • Confirm magic bytes, EOF marker, page count/order, dimensions, encryption, and expected metadata.
  • Compare extracted text or OCR page-by-page for content that should remain unchanged.
  • Confirm form values, bookmarks, links, annotations, attachments, and signatures as applicable.
  • For redaction, search extracted text and inspect page objects or a second parser; visual coverage alone is insufficient. Confirm the output was fully rewritten, contains no retained incremental revisions, and does not preserve the original bytes or redacted objects.
6. Render and inspect

Render every changed page and a representative sample of unchanged pages under the applicable trust/sandbox profile. Inspect at readable resolution for clipping, missing glyphs, shifted tables, broken images, blank pages, wrong rotation, and redaction leakage. For short or high-stakes documents, inspect every page. Open the final file in the user's intended viewer only when its trust classification and active-content controls permit it.

7. Deliver with evidence

Use processing-record-template.md for complex or regulated work. Report output paths, hashes, verification performed, failures, and untested behavior. Retain temporary artifacts only as long as needed and delete them only when authorized.

Safety and permission boundaries

  • Keep confidential PDFs local unless the user explicitly approves a specific service and data handling.
  • Never execute embedded scripts, launch actions, macros, media, or attachments.
  • Do not bypass passwords, access controls, digital rights, or signatures.
  • Do not expose passwords, private keys, personal data, or document contents in logs.
  • Treat true redaction, signature application, form submission, and publication as high-impact actions requiring explicit scope.
  • Do not represent OCR as authoritative for legal, medical, financial, or identity data; require human verification.
  • Do not claim PDF/A, PDF/UA, accessibility, evidentiary, or signature compliance without the relevant validator and evidence.

Recovery

If processing fails, preserve the source and failed output separately, capture the exact command/tool/version and error, and retry from the unchanged original. If the output was distributed before a defect was found, identify affected versions and recipients but do not recall or replace files without authorization. For suspected redaction leakage or signature/key exposure, stop distribution and escalate immediately.

Examples

Extract a scanned contract

Request: “Extract this 38-page scanned contract into searchable text with page citations.”

Hash and render the original, OCR with the stated language, preserve the PDF image layer, spot-check names/numbers on every page, flag low-confidence passages, and deliver text with physical page references and an OCR limitation note.

Merge a board packet

Request: “Combine these five PDFs in agenda order and add bookmarks.”

Confirm order and page mapping, merge to a new file, preserve orientation and page sizes, add bookmarks, re-open with a second reader, render boundary pages, and report final page count and hash.

Redact customer identifiers

Request: “Remove account numbers before this PDF is shared publicly.”

Define the identifier pattern and all affected pages, apply true redaction to a copy, remove scoped metadata/comments/attachments, search and inspect with an independent parser, render every redacted page, and state that visual boxes alone were not used.

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/bash/pdf-processing of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 2d14a69

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 8, 2026.

Compare with similar skills

PDF Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Processing this skillmajiayu000/claude-skill-registry6661 repos~2.5kAutomated safety check: PassMIT
Paper2posterQuZhan51496/paper2anything468—~9.2kAutomated safety check: NotesApache-2.0
Paper2wechatQuZhan51496/paper2anything468—~2.6kAutomated safety check: NotesApache-2.0
Markdown Exporterbowenliang123/markdown-exporter2711 repos~5.3kAutomated safety check: PassApache-2.0
Paper2xhsQuZhan51496/paper2anything468—~3.1kAutomated safety check: WarnApache-2.0
Export PDFbluzir/claude-code-design106—~811Automated safety check: PassNone

Similar skills

  • Paper2poster

    QuZhan51496/paper2anything

    Convert academic papers (PDF) into conference posters (HTML/PNG).

    468 GitHub stars~9.2k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • Paper2wechat

    QuZhan51496/paper2anything

    把学术论文 PDF 转成微信公众号深度解读推文(长文 + 配图 + 封面)。你主导设计的协调式:机械活(MinerU 解析 PDF、生成封面、md2wechat 发布草稿箱)交给 scripts/ 下的小工具,论文理解、文章结构、长文撰写由你亲自完成并在关键点与用户确认。当用户说“论文转公众号”、“paper2wechat”、“把论文写成公众号文章”、“论文转微信推文”、“PDF…

    468 GitHub stars~2.6k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • Markdown Exporter

    bowenliang123/markdown-exporter

    Convert Markdown text to DOCX, PPTX, XLSX, PDF, PNG, SVG, HTML, IPYNB, MD, CSV, JSON, JSONL, XML files, and extract code blocks in Markdown to Python, Bash,JS and etc files.

    271 GitHub starsUsed in 1 repo~5.3k tokens
    Documents & OfficeAuto-check passed
  • Paper2xhs

    QuZhan51496/paper2anything

    把学术论文 PDF 转成小红书多图帖(标题 + 正文 + 标签 + 封面 + 论文主图/主实验结果配图)。你主导设计的协调式:机械活(MinerU 解析 PDF、生成封面与配图、半自动发布)交给 scripts/ 下的小工具,论文理解、选题角度、文案撰写由你亲自完成并在关键点与用户确认。当用户说“论文转小红书”、“paper2xhs”、“把这篇论文发小红书”、“论文转社交媒体”、“PDF…

    468 GitHub stars~3.1k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: warnings
  • Export PDF

    bluzir/claude-code-design

    Export an HTML artifact to PDF via headless Chromium (puppeteer Page.pdf).

    106 GitHub stars~811 tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Paper2html

    QuZhan51496/paper2anything

    Convert an academic paper PDF into a publish-ready, self-contained single-page project homepage (a self-contained index.html) — the kind of paper landing page researchers host on GitHub Pages.

    468 GitHub stars~3.3k tokensUpdated 2 mo ago
    Frontend & DesignAuto-check: notes

More from majiayu000/claude-skill-registry

All 1,273 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Pyzotero

    majiayu000/claude-skill-registry

    Interact with Zotero reference management libraries using the pyzotero Python client.

    666 GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check: notes
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed

Works with

Questions about PDF Processing

What does PDF Processing do?

Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity. PDF Processing is an agent skill from majiayu000/claude-skill-registry. Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity.

When should I use PDF Processing?

PDF Processing fits situations like: working with one; more .pdf files; converting documents to; extracting text.

How do I install PDF Processing in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill pdf-processing -a claude-code`. Or copy the skill folder (skills/bash/pdf-processing in majiayu000/claude-skill-registry) into .claude/skills/pdf-processing in your project. Claude Code loads it when a task matches its description.

How do I install PDF Processing in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill pdf-processing -a codex`. Or copy the skill folder (skills/bash/pdf-processing in majiayu000/claude-skill-registry) into .agents/skills/pdf-processing in your project. Codex loads it when a task matches its description.

Can I use PDF Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill pdf-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-processing, .gemini/skills/pdf-processing, .github/skills/pdf-processing and .opencode/skills/pdf-processing in your project.

What does PDF Processing need to run?

Going by SKILL.md and its folder, PDF Processing needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does PDF Processing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF Processing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Processing use?

PDF Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Processing use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Processing?

Skills that share tags, products or a category with PDF Processing: Paper2poster (QuZhan51496/paper2anything, 468 stars), Paper2wechat (QuZhan51496/paper2anything, 468 stars), Markdown Exporter (bowenliang123/markdown-exporter, 271 stars) and Paper2xhs (QuZhan51496/paper2anything, 468 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Processing?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 1,273 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.