Agent skill

PDF Editing

by benchflow-ai in benchflow-ai/skillsbench

Complete guide for reading and editing PDF documents with PyMuPDF.

Apache-2.0Auto-check passedDocuments & Office

Install PDF Editing

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill pdf-editing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench pdf-editing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/edit-pdf/environment/skills/pdf-editing .claude/skills/pdf-editing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-editing
GitHub stars
1.8k
Token cost
~2.5k tokens
SKILL.md length
562 words
Files
1
Skills in repo
178
Repo updated
First seen
Licence
Apache-2.0

At a glance

Complete guide for reading and editing PDF documents with PyMuPDF.

  • Works in 2 steps: For REPLACING text (e.g., updating name,… → For TRUE REDACTION of sensitive data…
  • Tasks that involve PDF
  • SKILL.md covers CRITICAL RULES - READ FIRST, Overview, Reading PDF Content and Finding Text Positions, plus 7 more sections
  • Calls python3

What it does

PDF Editing is an agent skill from benchflow-ai/skillsbench. Complete guide for reading and editing PDF documents with PyMuPDF.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-editing”

Requirements

  • Python 3
  • Node.js

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. For REPLACING text (e.g., updating name, email, DOB)
  2. For TRUE REDACTION of sensitive data (e.g., student ID)

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Editing loads about 2.5k tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 562 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 562 words, ~2,484 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-editing/SKILL.md (or your agent's skills folder).
name
pdf-editing
description
Complete guide for reading and editing PDF documents with PyMuPDF.

PDF Editing Skill

CRITICAL RULES - READ FIRST

NEVER DO THESE:

  • NEVER use strikethrough lines to cross out text
  • NEVER rasterize or flatten the PDF to images
  • NEVER convert PDF pages to PNG/JPG and draw on them
  • NEVER use pdf-to-image-to-pdf workflows
  • NEVER use add_redact_annot() with BLACK fill (use WHITE fill instead)
  • NEVER add text NEXT TO old values - REPLACE them at the SAME position

TWO APPROACHES - CHOOSE THE RIGHT ONE:

  1. For REPLACING text (e.g., updating name, email, DOB):

    • Use draw_rect() with white fill to cover old text
    • Use insert_text() at the SAME position
    • Text layer is preserved
  2. For TRUE REDACTION of sensitive data (e.g., student ID):

    • Use add_redact_annot(rect, fill=(1,1,1)) with WHITE fill
    • Call apply_redactions() to REMOVE text from PDF structure
    • Then insert_text() to add masked value (e.g., "****5678")
    • Original text is completely removed, not just covered

Overview

USE PYTHON WITH PyMuPDF (fitz) - it is pre-installed and produces the best results.

PyMuPDF preserves the text layer properly, making text extractable after editing. JavaScript libraries like pdf-lib may create text that tools like pypdf cannot extract.

bash
# PyMuPDF is already installed - just use it
python3 -c "import fitz; print('PyMuPDF ready')"

Reading PDF Content

python
import fitz

doc = fitz.open("input.pdf")
page = doc[0]

# Extract all text to understand the document
text = page.get_text()
print(text)

Finding Text Positions

python
# search_for() returns list of rectangles where text is found
rects = page.search_for("Label Text")
if rects:
    rect = rects[0]
    # rect.x0, rect.y0 = top-left corner
    # rect.x1, rect.y1 = bottom-right corner
    print(f"Found at: ({rect.x0}, {rect.y0}) to ({rect.x1}, {rect.y1})")

Inserting Text

python
# Insert text at a specific position
page.insert_text(
    (x_position, y_position),  # coordinates
    "text to insert",
    fontsize=11,
    color=(0, 0, 0)  # black
)

doc.save("output.pdf")

Common Pattern: Fill Empty Form Fields

When a form field is empty (no existing value), insert text next to the label:

python
import fitz

doc = fitz.open("input.pdf")
page = doc[0]

# Find label, insert value to the right of it
label_rect = page.search_for("FIELD LABEL:")[0]
page.insert_text((label_rect.x1 + 5, label_rect.y1), "value", fontsize=11)

# For today's date
from datetime import datetime
date_rect = page.search_for("Date")[0]
today = datetime.now().strftime("%Y/%m/%d")
page.insert_text((date_rect.x1 + 5, date_rect.y1), today, fontsize=11)

# For signatures - insert the person's name
sig_rect = page.search_for("signature")[0]
page.insert_text((sig_rect.x1 + 5, sig_rect.y1), "Full Name", fontsize=12)

doc.save("output.pdf")

Replacing Existing Text (CRITICAL - MUST FOLLOW)

When you need to replace text that's already in the PDF with new/correct values:

  1. First extract the PDF text to see what's currently there
  2. Compare with the correct values from your input source
  3. For any value that needs to change: COVER with white rectangle, then insert new text AT THE SAME POSITION
WRONG vs RIGHT approaches:

WRONG - DO NOT DO THIS:

python
# WRONG: Adding text next to old value
page.insert_text((old_rect.x1 + 10, old_rect.y1), new_value)  # NO!

# WRONG: Using strikethrough
page.draw_line(start, end, color=(0,0,0))  # NO!

# WRONG: Rasterizing to image
pix = page.get_pixmap()  # NO! Destroys text layer

RIGHT - DO THIS:

python
import fitz

doc = fitz.open("input.pdf")
page = doc[0]

# First, extract text to see what's in the PDF
current_text = page.get_text()
print(current_text)  # Examine what values exist

# To replace a value you found that needs changing:
old_value = "..."  # The value you found in the PDF that's wrong
new_value = "..."  # The correct value from your input source

rects = page.search_for(old_value)
if rects:
    rect = rects[0]

    # STEP 1: Draw WHITE rectangle to COMPLETELY COVER old text
    page.draw_rect(rect, color=(1, 1, 1), fill=(1, 1, 1), width=0)

    # STEP 2: Insert new text at the SAME position (not offset!)
    page.insert_text((rect.x0, rect.y1), new_value, fontsize=11, color=(0, 0, 0))

doc.save("output.pdf")

Key points:

  • draw_rect(rect, fill=(1,1,1), width=0) draws a WHITE filled rectangle
  • (1, 1, 1) is white in RGB (0-1 scale)
  • Insert new text at rect.x0 (same X position), NOT rect.x1 + offset
  • The old text becomes invisible under the white rectangle
  • The new text appears in the same location
  • Text layer is preserved (text remains extractable)
Show full SKILL.md (231 more words)Show less

TRUE REDACTION for Sensitive Data (IMPORTANT!)

For sensitive data like student IDs, you must use TRUE REDACTION that removes the original text from the PDF structure. A white rectangle only VISUALLY covers text - tools like pypdf can still extract the hidden text!

CRITICAL DISTINCTION:

  • draw_rect() = Visual cover only (text still extractable by machines)
  • add_redact_annot() + apply_redactions() = TRUE redaction (text removed from PDF)

Example: "A12345678" should become "****5678" (show only last 4 digits)

python
import fitz

doc = fitz.open("input.pdf")
page = doc[0]

# 1. First find what value is in the PDF by reading the text
pdf_text = page.get_text()
# Find the ID in the text (e.g., "A88888888")

original_in_pdf = "A88888888"  # Value you found in the PDF
# Extract just the digits for masking
digits = ''.join(c for c in original_in_pdf if c.isdigit())
masked = "****" + digits[-4:]  # Result: "****8888"

rects = page.search_for(original_in_pdf)
if rects:
    rect = rects[0]

    # STEP 1: Create TIGHT bounding box to avoid covering nearby labels
    tight_rect = fitz.Rect(rect.x0, rect.y0 + 8, rect.x1, rect.y1 - 2)

    # STEP 2: Add redaction annotation with WHITE fill (not black!)
    page.add_redact_annot(tight_rect, fill=(1, 1, 1))  # WHITE fill

    # STEP 2: Apply redactions - this REMOVES the text from PDF structure
    page.apply_redactions()

    # STEP 3: Insert masked value at the same position
    page.insert_text((rect.x0, rect.y1), masked, fontsize=11, color=(0, 0, 0))

doc.save("output.pdf")

Why this works:

  • add_redact_annot(rect, fill=(1, 1, 1)) - marks area with WHITE rectangle
  • apply_redactions() - REMOVES the underlying text from PDF structure
  • insert_text() - adds the masked value

WRONG approaches:

python
# WRONG: Black box redaction (ugly and suspicious)
page.add_redact_annot(rect, fill=(0, 0, 0))  # NO! Use white fill

# WRONG: Only using draw_rect (text still extractable!)
page.draw_rect(rect, fill=(1, 1, 1))  # NO! This only covers visually
page.insert_text(...)  # Original text is still in PDF structure!

# WRONG: Adding masked value NEXT TO original
page.insert_text((rect.x1 + 10, rect.y1), masked)  # NO!

Workflow: Compare and Update

python
import fitz

doc = fitz.open("input.pdf")
page = doc[0]

# Step 1: Read PDF to see current content
pdf_text = page.get_text()
print(pdf_text)

# Step 2: Read your input source to get correct values
# (parse input file, extract the values you need)

# Step 3: For each value that differs, replace it
# Build a dict of {old_value_in_pdf: correct_value}
replacements = {}  # Populate by comparing PDF content with correct values

for old_val, new_val in replacements.items():
    if old_val == new_val:
        continue  # Skip if already correct
    rects = page.search_for(old_val)
    if rects:
        rect = rects[0]
        # Cover old text with white rectangle
        page.draw_rect(rect, color=(1, 1, 1), fill=(1, 1, 1), width=0)
        # Insert new text
        page.insert_text((rect.x0, rect.y1), new_val, fontsize=11, color=(0, 0, 0))

doc.save("output.pdf")

Important Guidelines

  1. COVER then REPLACE - Use white rectangle to cover old text, insert new text at SAME position
  2. Never use strikethrough - No lines through text, no crossing out
  3. Never rasterize - No converting PDF to images, no get_pixmap() workflows
  4. Never add text NEXT TO old values - Replace AT the same position
  5. Never use add_redact_annot() with black - It creates black boxes
  6. Preserve all labels - Form labels should remain visible
  7. Preserve text layer - Text must remain extractable after editing
  8. Match font size - Typically 10-12pt for forms

WARNING: pdf-lib may create text that cannot be extracted by pypdf. Use Python with PyMuPDF instead whenever possible.

If you must use Node.js/JavaScript, use pdf-lib with the same approach:

javascript
const { PDFDocument, rgb } = require('pdf-lib');
const fs = require('fs');

async function editPdf() {
  const pdfBytes = fs.readFileSync('input.pdf');
  const pdfDoc = await PDFDocument.load(pdfBytes);
  const page = pdfDoc.getPages()[0];
  const { height } = page.getSize();

  // To replace text at known coordinates:
  // STEP 1: Draw WHITE rectangle to cover old text
  page.drawRectangle({
    x: oldTextX,
    y: oldTextY,
    width: oldTextWidth,
    height: oldTextHeight,
    color: rgb(1, 1, 1),  // WHITE
  });

  // STEP 2: Draw new text at the SAME position
  page.drawText('New Value', {
    x: oldTextX,  // SAME X position, not offset!
    y: oldTextY,
    size: 11,
    color: rgb(0, 0, 0),
  });

  fs.writeFileSync('output.pdf', await pdfDoc.save());
}

WRONG with pdf-lib:

javascript
// WRONG: Adding text to the right of old value
page.drawText('New Value', {
  x: oldTextX + oldTextWidth + 10,  // NO! Don't offset
  ...
});

// WRONG: Drawing strikethrough line
page.drawLine({
  start: { x: x1, y: y },
  end: { x: x2, y: y },
  color: rgb(0, 0, 0),  // NO strikethrough!
});

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks/edit-pdf/environment/skills/pdf-editing of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

PDF Editing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Editing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Editing this skillbenchflow-ai/skillsbench1.8k—~2.5kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice8.9k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3141 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    314 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from benchflow-ai/skillsbench

All 178 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about PDF Editing

What does PDF Editing do?

Complete guide for reading and editing PDF documents with PyMuPDF. PDF Editing is an agent skill from benchflow-ai/skillsbench. Complete guide for reading and editing PDF documents with PyMuPDF.

When should I use PDF Editing?

PDF Editing fits situations like: tasks that involve PDF.

How do I install PDF Editing in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill pdf-editing -a claude-code`. Or copy the skill folder (tasks/edit-pdf/environment/skills/pdf-editing in benchflow-ai/skillsbench) into .claude/skills/pdf-editing in your project. Claude Code loads it when a task matches its description.

How do I install PDF Editing in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill pdf-editing -a codex`. Or copy the skill folder (tasks/edit-pdf/environment/skills/pdf-editing in benchflow-ai/skillsbench) into .agents/skills/pdf-editing in your project. Codex loads it when a task matches its description.

Can I use PDF Editing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill pdf-editing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-editing, .gemini/skills/pdf-editing, .github/skills/pdf-editing and .opencode/skills/pdf-editing in your project.

What does PDF Editing need to run?

Going by SKILL.md and its folder, PDF Editing needs the command-line tools its instructions call (python3). Our summary lists: Python 3; Node.js.

Does PDF Editing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PDF Editing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Editing use?

PDF Editing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Editing use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Editing?

Skills that share tags, products or a category with PDF Editing: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.9k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Editing?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.