Agent skill

PDF Toolkit

by borghei in borghei/Claude-Skills

Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version.

MITAuto-check passedDocuments & Office

Install PDF Toolkit

skills CLI
$ npx skills add borghei/Claude-Skills --skill pdf-toolkit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills pdf-toolkit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/documents/pdf-toolkit .claude/skills/pdf-toolkit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-toolkit
GitHub stars
874
Token cost
~1.4k tokens
SKILL.md length
630 words
Files
4 (incl. scripts, references, assets)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version.

  • Works in 3 steps: Run: python scripts/pdf_auditor.py… → Review metadata fields → If metadata leaks, re-export from source…
  • Tasks that involve PDF
  • SKILL.md covers Table of Contents, Keywords, Clarify First and Quick Start, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

PDF Toolkit is an agent skill from borghei/Claude-Skills. Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version. Use before sending a PDF externally, when redacting sensitive metadata, or running a PDF security review.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts, reference files and assets (for example `assets/pdf_handoff_checklist.md`, `references/pdf_handoff_guide.md` and `scripts/pdf_auditor.py`).

It sits in Documents & Office, covering PDF and Security review. It works with JavaScript. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Tasks that involve PDF
  • Tasks that involve Security review

Example prompts

  • “/pdf-toolkit”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Run: python scripts/pdf_auditor.py document.pdf
  2. Review metadata fields
  3. If metadata leaks, re-export from source with cleaned properties (or use a redaction tool)

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Toolkit loads about 1.4k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 55 tokens; SKILL.md has 630 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 630 words, ~1,364 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-toolkit/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
pdf-toolkit
description
Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version. Use before sending a PDF externally, when redacting sensitive metadata, or running a PDF security review.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
documents
metadata.domain
document-automation
metadata.updated
2026-05-04
metadata.python-tools
pdf_auditor.py
metadata.tech-stack
pdf

PDF Toolkit

Audit .pdf files for metadata, page count, encryption status, embedded JavaScript, embedded files, and PDF version — using the standard library only.


Table of Contents


Keywords

pdf, pdf audit, pdf metadata, pdf review, pdf leakage, pdf security, redaction, document handoff


Clarify First

Before running the audit, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Audit purpose (pre-handoff metadata scrub, inbound security triage, or bulk outbound check) — selects which of the 3 workflows and which fields you act on
  • Recipient / handling context (external party, managed laptop) — sets what counts as a leak or a threat worth quarantining
  • Expected legitimate metadata (who the author/title should be) — without it you can't distinguish a leak from expected data

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

bash
python scripts/pdf_auditor.py contract.pdf

Outputs: PDF version, page count, file size, metadata (Author, Title, Producer, Creator, dates), encryption status, embedded JavaScript indicators, embedded file indicators.


Core Workflows

Workflow 1: Pre-Handoff PDF Metadata Audit

Goal: Stop leaking author identity, prior client names, or document history when handing a PDF to an external party.

Steps:

  1. Run: python scripts/pdf_auditor.py document.pdf
  2. Review metadata fields:
    • Author matches the sender (not "Bob's intern" from a prior project)
    • Title matches the document, not a leftover working title
    • Producer doesn't reveal an internal-only PDF tool
    • CreationDate and ModDate are reasonable for the deal
  3. If metadata leaks, re-export from source with cleaned properties (or use a redaction tool)

Time Estimate: 2-3 minutes per document.

Workflow 2: PDF Security Triage

Goal: Decide whether a received PDF can be opened safely on a managed laptop.

Steps:

  1. Run audit
  2. JavaScript indicator present → quarantine; review in a sandbox
  3. Embedded files indicator present → list of file types; quarantine if unexpected
  4. Encrypted with non-empty owner password → request password from sender via separate channel
  5. Decision: open / quarantine / reject

Time Estimate: 1-2 minutes per inbound document.

Workflow 3: Bulk Audit of an Outbound Document Set

Goal: Audit every PDF in a folder before zipping for a customer or partner.

Steps:

  1. Loop: for f in *.pdf; do python scripts/pdf_auditor.py "$f" --json; done > audit.jsonl
  2. Parse the JSON Lines for any metadata leakage or anomalies
  3. Re-export problem files from source
  4. Re-run audit until clean

Time Estimate: 1-2 minutes per file.


Show full SKILL.md (220 more words)Show less

Tools

pdf_auditor.py

Reads a PDF using stdlib parsing — no pypdf or pdfplumber required. Detects:

  • PDF version (from header)
  • Page count (via /Type /Page object scan)
  • File size
  • Document Info / XMP metadata (Title, Author, Subject, Keywords, Producer, Creator, CreationDate, ModDate)
  • Encryption status (/Encrypt reference present)
  • JavaScript indicators (/JS, /JavaScript, /AA keys)
  • Embedded files indicator (/EmbeddedFiles)
bash
python scripts/pdf_auditor.py document.pdf
python scripts/pdf_auditor.py document.pdf --json

Limits:

  • Does not extract text content — pure PDF text extraction with stdlib is unreliable. For text extraction install pdfplumber or pypdf separately.
  • Cannot decrypt encrypted files.
  • Detects only the presence of JavaScript/embedded files, not their behavior.

Reference Guides

  • references/pdf_handoff_guide.md — What to scrub from PDFs before external send; PDF/A and PDF/UA basics; common leakage patterns

Templates

  • assets/pdf_handoff_checklist.md — Pre-send PDF sign-off checklist

Best Practices

  • Re-export rather than redact. Redaction tools that "remove" content can leave it recoverable. The safest path is regenerating the PDF from the source document with sensitive fields removed.
  • Scrub document properties at the source. In Word: File → Inspect Document → Document Inspector. In Pages: File → Properties. Then export to PDF.
  • Don't trust filenames. A file named Public-Report.pdf can carry private metadata indistinguishable to the human eye.
  • Use PDF/A for archival. PDF/A removes JavaScript and external dependencies, making documents safe for long-term archive.

Integration Points

  • Pairs with legal/ for redacted contract handoffs
  • Pairs with c-level-advisor/board-deck-builder for board pack handoff
  • Used by marketing/ for whitepaper / case-study handoff

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references, assets) in documents/pdf-toolkit of borghei/Claude-Skills.

  • SKILL.md
  • assets/pdf_handoff_checklist.md
  • references/pdf_handoff_guide.md
  • scripts/pdf_auditor.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

PDF Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Toolkit this skillborghei/Claude-Skills874—~1.4kAutomated safety check: PassMIT
Analyzing Malicious PDF With Peepdfmukul975/Anthropic-Cybersecurity-Skills34k—~799Automated safety check: PassApache-2.0
Jev SEOAgriciDaniel/jev-seo515—~2.5kAutomated safety check: NotesMIT
PDF Extraction Fallback 80956bHKUDS/OpenSpace7.7k—~1.8kAutomated safety check: PassMIT
PDF Extraction FallbacksHKUDS/OpenSpace7.7k—~1.8kAutomated safety check: PassMIT
PDF Extraction Fallbacks 7d54a9HKUDS/OpenSpace7.7k—~2kAutomated safety check: PassMIT

Similar skills

  • Analyzing Malicious PDF With Peepdf

    mukul975/Anthropic-Cybersecurity-Skills

    Perform static analysis of malicious PDF documents using peepdf, pdfid, and pdf-parser to extract embedded JavaScript, shellcode, and suspicious objects.

    34k GitHub stars~799 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Jev SEO

    AgriciDaniel/jev-seo

    Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).

    515 GitHub stars~2.5k tokensUpdated 15 days ago
    Documents & OfficeAuto-check: notes
  • Multi-fallback PDF download and text extraction with early failure detection

    7.7k GitHub stars~1.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks

    7.7k GitHub stars~1.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Multi-fallback PDF extraction with sequential approaches and early failure detection

    7.7k GitHub stars~2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Edit PDF

    SimplePDF/simplepdf-embed

    Edit and fill PDF documents. An agent skill from SimplePDF/simplepdf-embed.

    407 GitHub stars~1.5k tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Questions about PDF Toolkit

What does PDF Toolkit do?

Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version. PDF Toolkit is an agent skill from borghei/Claude-Skills. Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version.

When should I use PDF Toolkit?

PDF Toolkit fits situations like: tasks that involve PDF; tasks that involve Security review.

How do I install PDF Toolkit in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill pdf-toolkit -a claude-code`. Or copy the skill folder (documents/pdf-toolkit in borghei/Claude-Skills) into .claude/skills/pdf-toolkit in your project. Claude Code loads it when a task matches its description.

How do I install PDF Toolkit in Codex?

Run `npx skills add borghei/Claude-Skills --skill pdf-toolkit -a codex`. Or copy the skill folder (documents/pdf-toolkit in borghei/Claude-Skills) into .agents/skills/pdf-toolkit in your project. Codex loads it when a task matches its description.

Can I use PDF Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill pdf-toolkit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-toolkit, .gemini/skills/pdf-toolkit, .github/skills/pdf-toolkit and .opencode/skills/pdf-toolkit in your project.

What does PDF Toolkit need to run?

Going by SKILL.md and its folder, PDF Toolkit needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does PDF Toolkit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF Toolkit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF Toolkit use?

PDF Toolkit is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Toolkit use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to PDF Toolkit?

Skills that share tags, products or a category with PDF Toolkit: Analyzing Malicious PDF With Peepdf (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Jev SEO (AgriciDaniel/jev-seo, 515 stars), PDF Extraction Fallback 80956b (HKUDS/OpenSpace, 7.7k stars) and PDF Extraction Fallbacks (HKUDS/OpenSpace, 7.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Toolkit?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.