Agent skill

Academic PDF Redaction

by benchflow-ai in benchflow-ai/skillsbench

Redact text from PDF documents for blind review anonymization

Apache-2.0Auto-check passedDocuments & Office

Install Academic PDF Redaction

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill academic-pdf-redaction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench academic-pdf-redaction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/paper-anonymizer/environment/skills/academic-pdf-redaction .claude/skills/academic-pdf-redaction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
academic-pdf-redaction
GitHub stars
1.8k
Token cost
~933 tokens
SKILL.md length
153 words
Files
1
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

Redact text from PDF documents for blind review anonymization

  • Works in 3 steps: PRESERVE References section -… → ONLY redact specific text matches -… → VERIFY output - Check that 80%+ of…
  • Tasks that involve PDF
  • SKILL.md covers CRITICAL RULES, Common Pitfalls to AVOID, Patterns to Redact (Before… and PyMuPDF (fitz) - Recommended…, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Academic PDF Redaction is an agent skill from benchflow-ai/skillsbench. Redact text from PDF documents for blind review anonymization

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/academic-pdf-redaction”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. PRESERVE References section - Self-citations MUST remain intact
  2. ONLY redact specific text matches - Never redact entire pages/regions
  3. VERIFY output - Check that 80%+ of original text remains

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Academic PDF Redaction loads about 933 tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 153 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~933

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 153 words, ~933 tokens.

Download SKILL.mdSave it as .claude/skills/academic-pdf-redaction/SKILL.md (or your agent's skills folder).
name
academic-pdf-redaction
description
Redact text from PDF documents for blind review anonymization

PDF Redaction for Blind Review

Redact identifying information from academic papers for blind review.

CRITICAL RULES

  1. PRESERVE References section - Self-citations MUST remain intact
  2. ONLY redact specific text matches - Never redact entire pages/regions
  3. VERIFY output - Check that 80%+ of original text remains

Common Pitfalls to AVOID

python
# ❌ WRONG - This removes ALL text from the page:
for block in page.get_text("blocks"):
    page.add_redact_annot(fitz.Rect(block[:4]))

# ❌ WRONG - Drawing rectangles over text:
page.draw_rect(fitz.Rect(0, 0, 600, 100), fill=(0,0,0))

# ✅ CORRECT - Only redact specific search matches:
for rect in page.search_for("John Smith"):
    page.add_redact_annot(rect)

Patterns to Redact (Before References Only)

IMPORTANT: Use FULL names/phrases, not partial matches!

  • ✅ "John Smith" (full name)
  • ❌ "Smith" (partial - would incorrectly match "Smith et al." citations in References)
  1. Author names - FULL names only (e.g., "John Smith", not just "Smith")
  2. Affiliations - Universities, companies (e.g., "Duke University")
  3. Email addresses - Pattern: *@*.edu, *@*.com
  4. Venue names - Conference/workshop names (e.g., "ICML 2024", "ICML Workshop")
  5. arXiv identifiers - Pattern: arXiv:XXXX.XXXXX
  6. DOIs - Pattern: 10.XXXX/...
  7. Acknowledgement names - Names in "Acknowledgements" section
  8. Equal contribution footnotes - e.g., "Equal contribution", "* Equal contribution"
python
import fitz
import os

def redact_with_pymupdf(input_path: str, output_path: str, patterns: list[str]):
    """Redact specific patterns from PDF using PyMuPDF."""
    doc = fitz.open(input_path)
    original_len = sum(len(p.get_text()) for p in doc)

    # Find References page - stop redacting there
    references_page = None
    for i, page in enumerate(doc):
        if "references" in page.get_text().lower():
            references_page = i
            break

    for page_num, page in enumerate(doc):
        if references_page is not None and page_num >= references_page:
            continue  # Skip References section

        for pattern in patterns:
            # ONLY redact exact search matches
            for rect in page.search_for(pattern):
                page.add_redact_annot(rect, fill=(0, 0, 0))
        page.apply_redactions()

    os.makedirs(os.path.dirname(output_path), exist_ok=True)
    doc.save(output_path)
    doc.close()

    # MUST verify after saving
    verify_redaction(input_path, output_path)

REQUIRED: Verification Function

Always run this after ANY redaction to catch errors early:

python
import fitz

def verify_redaction(original_path, output_path):
    """Verify redaction didn't corrupt the PDF."""
    orig = fitz.open(original_path)
    redc = fitz.open(output_path)

    orig_len = sum(len(p.get_text()) for p in orig)
    redc_len = sum(len(p.get_text()) for p in redc)

    print(f"Original: {len(orig)} pages, {orig_len} chars")
    print(f"Redacted: {len(redc)} pages, {redc_len} chars")
    print(f"Retained: {redc_len/orig_len:.1%}")

    # DEFENSIVE CHECKS - fail fast if something went wrong
    if len(redc) != len(orig):
        raise ValueError(f"Page count changed: {len(orig)} -> {len(redc)}")
    if redc_len < 1000:
        raise ValueError(f"PDF corrupted: only {redc_len} chars remain!")
    if redc_len < orig_len * 0.7:
        raise ValueError(f"Too much removed: kept only {redc_len/orig_len:.0%}")

    orig.close()
    redc.close()
    print("✓ Verification passed")

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks/paper-anonymizer/environment/skills/academic-pdf-redaction of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Academic PDF Redaction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Academic PDF Redaction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Academic PDF Redaction this skillbenchflow-ai/skillsbench1.8k—~933Automated safety check: PassApache-2.0
Split PDFscunning1975/MixtapeTools4742 repos~2.9kAutomated safety check: PassNone
Paper Interpretationdigoal/blog8.6k—~1.5kAutomated safety check: PassGPL-2.0
Paper2slidesQuZhan51496/paper2anything450—~3.8kAutomated safety check: NotesApache-2.0
Paper LensYSQ-boop/paper-lens101—~1.3kAutomated safety check: PassApache-2.0
Geng Academic Fraud Detectorwooly99/geng-academic-fraud-detector278—~970Automated safety check: PassNone

Similar skills

  • Split PDF

    scunning1975/MixtapeTools

    Download, split, and deeply read academic PDFs. An agent skill from scunning1975/MixtapeTools.

    474 GitHub starsUsed in 2 repos~2.9k tokens
    Documents & OfficeAuto-check passed
  • 从论文 PDF 文件或论文 PDF URL 生成通俗易懂、图文并茂、带批判性评估的中文 Markdown 解读,并保存到当前项目的 markdown 目录。Use when the user asks to interpret,精读,解读,summarize,explain,analyze, or write an article from an academic paper PDF…

    8.6k GitHub stars~1.5k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Paper2slides

    QuZhan51496/paper2anything

    Turn an academic paper PDF into a presentation deck (.pptx) end-to-end.

    450 GitHub stars~3.8k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes
  • Paper Lens

    YSQ-boop/paper-lens

    Read and critically analyze one academic paper from an arXiv URL/ID or a local PDF, producing a source-grounded Markdown report that can grow from a quick read into a reviewer-level deep review.

    101 GitHub stars~1.3k tokensUpdated 11 days ago
    Documents & OfficeAuto-check passed
  • Geng Academic Fraud Detector

    wooly99/geng-academic-fraud-detector

    学术论文打假检测器,致敬耿同学。分析学术论文 PDF,检测数据造假、图片复用/拼接、Western blot 操纵、统计异常等学术不端行为。当用户提供论文 PDF 要求"查重"、"打假"、"检测造假"、"论文分析"、"学术打假"时使用。

    278 GitHub stars~970 tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Paper2poster

    QuZhan51496/paper2anything

    Convert academic papers (PDF) into conference posters (HTML/PNG).

    450 GitHub stars~9.2k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check: notes

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Academic PDF Redaction

What does Academic PDF Redaction do?

Redact text from PDF documents for blind review anonymization. Academic PDF Redaction is an agent skill from benchflow-ai/skillsbench.

When should I use Academic PDF Redaction?

Academic PDF Redaction fits situations like: tasks that involve PDF.

How do I install Academic PDF Redaction in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill academic-pdf-redaction -a claude-code`. Or copy the skill folder (tasks/paper-anonymizer/environment/skills/academic-pdf-redaction in benchflow-ai/skillsbench) into .claude/skills/academic-pdf-redaction in your project. Claude Code loads it when a task matches its description.

How do I install Academic PDF Redaction in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill academic-pdf-redaction -a codex`. Or copy the skill folder (tasks/paper-anonymizer/environment/skills/academic-pdf-redaction in benchflow-ai/skillsbench) into .agents/skills/academic-pdf-redaction in your project. Codex loads it when a task matches its description.

Can I use Academic PDF Redaction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill academic-pdf-redaction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/academic-pdf-redaction, .gemini/skills/academic-pdf-redaction, .github/skills/academic-pdf-redaction and .opencode/skills/academic-pdf-redaction in your project.

What does Academic PDF Redaction need to run?

SKILL.md names no scripts, command-line tools or credentials: Academic PDF Redaction is instructions for the agent only. Our summary lists: Python 3.

Does Academic PDF Redaction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Academic PDF Redaction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Academic PDF Redaction use?

Academic PDF Redaction is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Academic PDF Redaction use?

About 933 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Academic PDF Redaction?

Skills that share tags, products or a category with Academic PDF Redaction: Split PDF (scunning1975/MixtapeTools, 474 stars), Paper Interpretation (digoal/blog, 8.6k stars), Paper2slides (QuZhan51496/paper2anything, 450 stars) and Paper Lens (YSQ-boop/paper-lens, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Academic PDF Redaction?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,835 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.