Agent skill

PDF

by holon-run in holon-run/holon

Inspect, extract, assemble, generate, render, OCR, and validate PDF files with local tools while treating active content and external services as explicit security boundaries.

Apache-2.0Auto-check passedDocuments & Office

Install PDF

skills CLI
$ npx skills add holon-run/holon --skill pdf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install holon-run/holon pdf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/holon-run/holon.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf .claude/skills/pdf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf
GitHub stars
153
Token cost
~1.5k tokens
SKILL.md length
776 words
Files
3 (incl. references, assets)
Skills in repo
13
Repo updated
First seen
Licence
Apache-2.0

At a glance

Inspect, extract, assemble, generate, render, OCR, and validate PDF files with local tools while treating active content and external services as explicit security boundaries.

  • Works in 9 steps: Confirm the requested pages, operation,… → Identify the file type and inspect… → Decide whether the task is extraction,… → …
  • Tasks that involve PDF
  • SKILL.md covers Summary, When To Use, Safety Boundaries and Backend Selection, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

PDF is an agent skill from holon-run/holon. Inspect, extract, assemble, generate, render, OCR, and validate PDF files with local tools while treating active content and external services as explicit security boundaries.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files and assets (for example `references/generation.md`).

It sits in Documents & Office, covering PDF. It works with pypdf. The repository describes itself as: An agent workbench for ongoing work: preserve goals and progress, connect events and schedules, and resume when the next condition is met. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the requested pages, operation, output path, preservation needs,
  2. Identify the file type and inspect encryption, signatures, actions,
  3. Decide whether the task is extraction, page transformation, annotation,
  4. For generation, establish the document type, audience, visual tone, page
  5. Select the narrowest local backend that supports the operation.
  6. Produce a new output without activating links or embedded content.
  7. Reopen the result and check page count, dimensions, metadata, bookmarks,
  8. Render every changed or generated page to images. Inspect the actual pages,
  9. For extraction or OCR, sample the output against rendered pages and record

What it can do on your machine

Read from SKILL.md and the folder at commit 671bd7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF loads about 1.5k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 45 tokens; SKILL.md has 776 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from holon-run/holon at commit 671bd7a, republished under its Apache-2.0 licence (© holon-run). 776 words, ~1,517 tokens.

Download SKILL.mdSave it as .claude/skills/pdf/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
pdf
description
Inspect, extract, assemble, generate, render, OCR, and validate PDF files with local tools while treating active content and external services as explicit security boundaries.

PDF

Summary

Use this skill for PDF reading, extraction, generation, page operations, forms, rendering, and optional OCR. Match the tool to the operation instead of treating PDF as a simple editable text format.

Common neutral backends include pypdf for page and metadata operations, pdfplumber for text and table inspection, ReportLab or pdf-lib for generation, HTML/CSS plus a local browser for typography-heavy documents, Poppler for rendering, and qpdf for structural checks. OCR is a separate, explicitly selected path such as local Tesseract.

When To Use

  • Extracting text, tables, metadata, links, bookmarks, or form information
  • Splitting, merging, rotating, cropping, stamping, or encrypting documents
  • Generating a PDF from structured content
  • Rendering pages for visual review
  • Running OCR on operator-approved scanned pages
  • Comparing page-level content before and after a transformation

Safety Boundaries

  • Preserve the original and write transformed output to a new path.
  • Treat JavaScript, actions, attachments, forms, links, signatures, and embedded files as untrusted.
  • Never execute document JavaScript, launch actions, embedded programs, or external links.
  • Do not remove passwords, permissions, signatures, or protection as a way to bypass access controls.
  • External OCR or conversion requires explicit approval before upload.
  • Warn that any content change can invalidate a digital signature.

If a parser reports corruption, suspicious object expansion, extreme page dimensions, excessive object counts, or unsupported encryption, stop the operation and report the limitation.

Backend Selection

  • Use pypdf for page assembly, rotation, cropping, metadata, common forms, and supported encryption operations.
  • Use pdfplumber for positioned text and rule-based table extraction.
  • Prefer HTML/CSS plus a local browser or WeasyPrint for reports, proposals, manuals, and other prose-heavy documents where typography and page design matter.
  • Use ReportLab or pdf-lib for precise drawing, overlays, labels, forms, and documents whose layout is naturally coordinate-driven.
  • Use Poppler tools for local text extraction or rendering when installed.
  • Use qpdf for structural validation and supported transformations.
  • Use Tesseract only for explicitly requested local OCR and retain the original page images.

Do not describe text replacement as general PDF editing. Existing visual content usually requires reconstruction, redaction, annotation, or a source document change rather than in-place prose editing.

Workflow

  1. Confirm the requested pages, operation, output path, preservation needs, forms or bookmarks requirements, and whether OCR is allowed.
  2. Identify the file type and inspect encryption, signatures, actions, attachments, links, page boxes, fonts, and page count.
  3. Decide whether the task is extraction, page transformation, annotation, redaction, generation, or reconstruction.
  4. For generation, establish the document type, audience, visual tone, page size, language coverage, font files, hierarchy, color tokens, and recurring components before writing rendering code. Read references/generation.md for the required typography and layout workflow. The reusable assets/chinese-report.css is a neutral starting point for polished Chinese or mixed Chinese/Latin reports, not a substitute for matching the operator's requested brand or style.
  5. Select the narrowest local backend that supports the operation.
  6. Produce a new output without activating links or embedded content.
  7. Reopen the result and check page count, dimensions, metadata, bookmarks, forms, attachments, encryption state, and expected text.
  8. Render every changed or generated page to images. Inspect the actual pages, not only the source HTML or drawing commands, and revise visible defects before delivery.
  9. For extraction or OCR, sample the output against rendered pages and record ambiguous tables, reading order, missing glyphs, and low-confidence text.
Show full SKILL.md (231 more words)Show less

Quality Rules

  • Never rely on Helvetica, Times, or another Latin-only base font for Chinese, Japanese, or Korean text. Discover an appropriate local font, verify glyph coverage, and ensure fonts are embedded or otherwise reliably available in the produced PDF.
  • Use a deliberate type scale, spacing rhythm, margins, line length, and color palette. A technically valid PDF with backend-default styling is not a finished generated document.
  • For prose-heavy documents, prefer semantic flow layout over manually placing every line. Prevent headings, tables, figures, and callouts from splitting in visually confusing ways.
  • Check rendered pages for missing-glyph boxes, substituted fonts, clipping, overlap, isolated headings, sparse spill pages, inconsistent spacing, weak contrast, and unreadably dense tables.
  • Preserve intended page order, orientation, crop boxes, and dimensions.
  • Distinguish visual redaction from secure content removal; verify that redacted text and related objects are not extractable.
  • Do not infer table structure without checking coordinates and rendered pages.
  • Retain source page references in extracted content when useful.
  • Label OCR-derived text and do not silently replace uncertain characters.
  • Verify links, bookmarks, forms, and signatures if the requested operation can affect them.

Delivery

Report the output path, page range and operation, tools used, whether OCR or any network service was used, structural and visual checks, signature or form impact, font family and embedding status for generated documents, and extraction or OCR limitations. State explicitly when active content was present but not executed.

© holon-run, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references, assets) in skills/pdf of holon-run/holon.

  • SKILL.md
  • assets/chinese-report.css
  • references/generation.md

Open the folder on GitHubat commit 671bd7a

Compare with similar skills

PDF next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF this skillholon-run/holon153—~1.5kAutomated safety check: PassApache-2.0
PDFnuoyimanaituling/manus-x830—~985Automated safety check: PassNone
PDFeinverne/dotfiles12147 repos~1.8kAutomated safety check: PassProprietary
Reportlabjimmc414/Kosmos5941 repos~4.2kAutomated safety check: PassNone
PDFguyi-a/pi-ling106—~3.3kAutomated safety check: PassMIT
PDF ReadingWide-Moat/open-computer-use1261 repos~2.7kAutomated safety check: PassProprietary

Similar skills

  • PDF

    nuoyimanaituling/manus-x

    Process PDF files - extract text, read content, create PDFs, merge or split documents.

    830 GitHub stars~985 tokensUpdated 8 mo ago
    Documents & OfficeAuto-check passed
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 47 repos~1.8k tokens
    Documents & OfficeAuto-check passed
  • Reportlab

    jimmc414/Kosmos

    PDF generation toolkit. An agent skill from jimmc414/Kosmos.

    594 GitHub starsUsed in 1 repo~4.2k tokens
    Documents & OfficeAuto-check passed
  • PDF

    guyi-a/pi-ling

    PDF 相关的所有操作:从零生成(reportlab / pypdf)、格式转化(md/html → PDF)、修改(合并 / 拆分 / 旋转 / 加水印 / 提图片 / 元数据)、读内容(pdfplumber / extractdocumenttext)、OCR 扫描件、加密解密。触发场景:用户说"生成 PDF" / "做份 PDF 简历" / "合并这几份 PDF" / "给 PDF…

    106 GitHub stars~3.3k tokensUpdated 22 days ago
    Documents & OfficeAuto-check passed
  • PDF Reading

    Wide-Moat/open-computer-use

    A skill your agent uses when you need to read, inspect, or extract content from PDF files — especially when file content is NOT in your context and you need to read it from disk.

    126 GitHub starsUsed in 1 repo~2.7k tokens
    Documents & OfficeAuto-check passed
  • PDF

    LeastBit/Claude_skills_zh-CN

    全面的 PDF 操作工具包,用于提取文本和表格、创建新 PDF、合并/拆分文档以及处理表单。当 Claude 需要填写 PDF 表单或以编程方式大规模处理、生成或分析 PDF 文档时使用。

    588 GitHub stars~1.5k tokensUpdated 8 mo ago
    Documents & OfficeAuto-check passed

More from holon-run/holon

All 13 skills in this repo
  • Video Production

    holon-run/holon

    Assemble existing local images, videos, audio and subtitles into preview/final videos with FFmpeg, technical QC and provenance.

    153 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • GitHub Issue Solve

    holon-run/holon

    Solve a GitHub issue by collecting context, implementing a fix, and opening or updating a pull request.

    153 GitHub stars~912 tokensUpdated today
    Auto-check passed
  • GitHub PR Fix

    holon-run/holon

    Fix a GitHub pull request by addressing feedback or CI failures, pushing changes, and publishing replies.

    153 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Code Health Audit

    holon-run/holon

    Audit code-health and technical-debt signals with evidence, rank proportionate interventions from focused cleanup to broad coordinated refactors, and draft implementation-ready plans without…

    153 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Code Review

    holon-run/holon

    Review a set of code changes using evidence-backed findings, explicit confidence, and a clear coverage summary.

    153 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Ghx

    holon-run/holon

    Guidance for safe, reliable GitHub CLI workflows across issues, pull requests, and reviews.

    153 GitHub stars~1k tokensUpdated today
    Auto-check passed

Works with

Questions about PDF

What does PDF do?

Inspect, extract, assemble, generate, render, OCR, and validate PDF files with local tools while treating active content and external services as explicit security boundaries. PDF is an agent skill from holon-run/holon. Inspect, extract, assemble, generate, render, OCR, and validate PDF files with local tools while treating active content and external services as explicit security boundaries.

When should I use PDF?

PDF fits situations like: tasks that involve PDF.

How do I install PDF in Claude Code?

Run `npx skills add holon-run/holon --skill pdf -a claude-code`. Or copy the skill folder (skills/pdf in holon-run/holon) into .claude/skills/pdf in your project. Claude Code loads it when a task matches its description.

How do I install PDF in Codex?

Run `npx skills add holon-run/holon --skill pdf -a codex`. Or copy the skill folder (skills/pdf in holon-run/holon) into .agents/skills/pdf in your project. Codex loads it when a task matches its description.

Can I use PDF in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add holon-run/holon --skill pdf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf, .gemini/skills/pdf, .github/skills/pdf and .opencode/skills/pdf in your project.

What does PDF need to run?

SKILL.md names no scripts, command-line tools or credentials: PDF is instructions for the agent only.

Does PDF access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF use?

PDF is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to PDF?

Skills that share tags, products or a category with PDF: PDF (nuoyimanaituling/manus-x, 830 stars), PDF (einverne/dotfiles, 121 stars), Reportlab (jimmc414/Kosmos, 594 stars) and PDF (guyi-a/pi-ling, 106 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF?

holon-run (a GitHub organization) maintains it in holon-run/holon, which has 153 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 8, 2026.

Source: holon-run/holon on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.