Agent skill

PDF Extraction Fallbacks

by HKUDS in HKUDS/OpenSpace

Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks

MITAuto-check passedDocuments & Office

Install PDF Extraction Fallbacks

skills CLI
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks .claude/skills/pdf-extraction-fallbacks && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-extraction-fallbacks
GitHub stars
7.7k
Token cost
~1.8k tokens
SKILL.md length
315 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks

  • Works in 6 steps: Download and Validate → Primary Extraction (pdftotext) → Fallback 1 (PyMuPDF/fitz) → …
  • Tasks that involve PDF
  • SKILL.md covers Purpose, Core Pattern, Step-by-Step Instructions and Decision Tree, plus 4 more sections
  • Calls curl, pip and pdftotext

What it does

PDF Extraction Fallbacks is an agent skill from HKUDS/OpenSpace. Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering PDF. It works with JavaScript. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-extraction-fallbacks”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Download and Validate
  2. Primary Extraction (pdftotext)
  3. Fallback 1 (PyMuPDF/fitz)
  4. Fallback 2 (pdfplumber)
  5. Handle JavaScript-Protected Pages
  6. Alternative Sources

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • pip
    • pdftotext
    • python3
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Extraction Fallbacks loads about 1.8k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 315 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 315 words, ~1,830 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-extraction-fallbacks/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-extraction-fallbacks
description
Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks

PDF Extraction with Multi-Fallback Strategy

Purpose

When extracting text from PDFs (especially regulatory documents, handbooks, or protected content), single-method approaches often fail due to JavaScript protection, CORS restrictions, encoding issues, or corrupted downloads. This skill provides a robust multi-fallback workflow that detects failures early and tries sequential extraction methods.

Core Pattern

  1. Download with validation - Check file size and content sanity immediately
  2. Sequential extraction attempts - Try multiple tools in order of reliability
  3. Early failure detection - Don't proceed with obviously corrupt files
  4. Document fallback path - Log which method succeeded for future reference

Step-by-Step Instructions

Step 1: Download and Validate

Before attempting extraction, validate the downloaded file:

bash
# Download the PDF
curl -L -o document.pdf "$URL"

# Check file size (reject if < 1KB - likely error page)
FILE_SIZE=$(stat -f%z document.pdf 2>/dev/null || stat -c%s document.pdf 2>/dev/null)
if [ "$FILE_SIZE" -lt 1024 ]; then
    echo "ERROR: File too small ($FILE_SIZE bytes) - likely not a valid PDF"
    # Check if it's an HTML error page
    head -c 200 document.pdf | grep -i "<html\|<!doctype\|error\|access denied" && \
        echo "Detected HTML error page instead of PDF"
    exit 1
fi

# Check PDF magic bytes
HEAD_BYTES=$(head -c 4 document.pdf)
if [ "$HEAD_BYTES" != "%PDF" ]; then
    echo "ERROR: File does not start with PDF magic bytes"
    head -c 100 document.pdf
    exit 1
fi
Step 2: Primary Extraction (pdftotext)
bash
# Try pdftotext first (fastest, most reliable for simple PDFs)
if command -v pdftotext &> /dev/null; then
    pdftotext -layout document.pdf output.txt 2>/dev/null
    if [ -s output.txt ]; then
        WORD_COUNT=$(wc -w < output.txt)
        if [ "$WORD_COUNT" -gt 50 ]; then
            echo "SUCCESS: pdftotext extracted $WORD_COUNT words"
            exit 0
        fi
    fi
fi
Step 3: Fallback 1 (PyMuPDF/fitz)
python
# Try PyMuPDF - handles more complex PDFs
import fitz  # pymupdf

try:
    doc = fitz.open("document.pdf")
    text = ""
    for page in doc:
        text += page.get_text()
    
    if len(text.strip()) > 500:  # Sanity check
        with open("output.txt", "w") as f:
            f.write(text)
        print(f"SUCCESS: PyMuPDF extracted {len(text)} characters")
    else:
        print("WARNING: PyMuPDF extraction too short, trying next method")
except Exception as e:
    print(f"PyMuPDF failed: {e}")
Step 4: Fallback 2 (pdfplumber)
python
# Try pdfplumber - better for tables and structured content
import pdfplumber

try:
    text = ""
    with pdfplumber.open("document.pdf") as pdf:
        for page in pdf.pages:
            page_text = page.extract_text()
            if page_text:
                text += page_text + "\n"
    
    if len(text.strip()) > 500:
        with open("output.txt", "w") as f:
            f.write(text)
        print(f"SUCCESS: pdfplumber extracted {len(text)} characters")
    else:
        print("WARNING: pdfplumber extraction too short")
except Exception as e:
    print(f"pdfplumber failed: {e}")
Step 5: Handle JavaScript-Protected Pages

If all methods fail, the PDF may be JavaScript-protected:

python
# Check for JavaScript in PDF
import fitz

doc = fitz.open("document.pdf")
has_js = False
for page in doc:
    if page.get_java_script():
        has_js = True
        break

if has_js:
    print("WARNING: PDF contains JavaScript - may be protected")
    # Try rendering pages as images and OCR (requires additional tools)
    # Or try alternative download source
Step 6: Alternative Sources

If the primary URL fails:

  • Try alternative domains (e.g., .gov mirrors, archive.org)
  • Check if the document is available via API
  • Look for HTML version of the same content
  • Search for the document title + "pdf" to find mirrors

Decision Tree

Download PDF
    │
    ├─→ File < 1KB? → REJECT (likely error page)
    ├─→ No %PDF header? → REJECT (not a PDF)
    │
    └─→ Valid PDF
         │
         ├─→ pdftotext → >50 words? → SUCCESS
         │              └─→ Try next
         │
         ├─→ PyMuPDF → >500 chars? → SUCCESS
         │             └─→ Try next
         │
         ├─→ pdfplumber → >500 chars? → SUCCESS
         │                 └─→ Try next
         │
         └─→ All failed → Check for JS protection, try alternative sources

Common Failure Modes

SymptomLikely CauseSolution
File < 100 bytesJavaScript error pageCheck CORS, try different user-agent
File ~1-5KBHTML error/warning pageParse HTML for actual PDF link
pdftotext returns emptyEncrypted/protected PDFTry PyMuPDF with password handling
Garbled text outputEncoding issueTry pdfplumber, specify encoding
Extraction very shortImages-only PDFNeed OCR (tesseract)

Example Complete Workflow Script

Save as extract_pdf_robust.sh:

bash
#!/bin/bash
set -e

URL="$1"
OUTPUT="${2:-output.txt}"
TEMP_PDF="temp_download.pdf"

echo "Downloading from: $URL"
curl -L -A "Mozilla/5.0" -o "$TEMP_PDF" "$URL"

# Validate
SIZE=$(stat -c%s "$TEMP_PDF" 2>/dev/null || stat -f%z "$TEMP_PDF")
echo "Downloaded: $SIZE bytes"

if [ "$SIZE" -lt 1024 ]; then
    echo "ERROR: File too small - checking content..."
    head -200 "$TEMP_PDF"
    exit 1
fi

if ! head -c 4 "$TEMP_PDF" | grep -q "%PDF"; then
    echo "ERROR: Not a valid PDF file"
    head -100 "$TEMP_PDF"
    exit 1
fi

# Try extraction methods
python3 << 'PYTHON'
import sys
import fitz
import pdfplumber

pdf_path = "temp_download.pdf"
output_path = "output.txt"

# Method 1: PyMuPDF
try:
    doc = fitz.open(pdf_path)
    text = "".join(page.get_text() for page in doc)
    if len(text.strip()) > 500:
        with open(output_path, "w") as f:
            f.write(text)
        print(f"PyMuPDF: {len(text)} chars")
        sys.exit(0)
except Exception as e:
    print(f"PyMuPDF failed: {e}")

# Method 2: pdfplumber
try:
    text = ""
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            txt = page.extract_text()
            if txt:
                text += txt + "\n"
    if len(text.strip()) > 500:
        with open(output_path, "w") as f:
            f.write(text)
        print(f"pdfplumber: {len(text)} chars")
        sys.exit(0)
except Exception as e:
    print(f"pdfplumber failed: {e}")

print("All extraction methods failed")
sys.exit(1)
PYTHON

Best Practices

  1. Always validate - Never assume a download succeeded
  2. Log the path taken - Record which method worked for debugging
  3. Set reasonable thresholds - 50 words / 500 chars minimum for "success"
  4. Keep raw PDF - Don't delete the original until extraction is confirmed
  5. Retry with variations - Different user-agents, referer headers, or mirrors

Dependencies

  • curl - For downloading
  • poppler-utils (pdftotext) - Optional, fast extraction
  • PyMuPDF (fitz) - pip install pymupdf
  • pdfplumber - pip install pdfplumber

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/pdf-extraction-fallbacks of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

PDF Extraction Fallbacks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Extraction Fallbacks compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Extraction Fallbacks this skillHKUDS/OpenSpace7.7k—~1.8kAutomated safety check: PassMIT
Jev SEOAgriciDaniel/jev-seo527—~2.5kAutomated safety check: NotesMIT
Analyzing Malicious PDF With Peepdfmukul975/Anthropic-Cybersecurity-Skills34k—~799Automated safety check: PassApache-2.0
PDF Toolkitborghei/Claude-Skills881—~1.4kAutomated safety check: PassMIT
Edit PDFSimplePDF/simplepdf-embed407—~1.5kAutomated safety check: PassMIT
Build With SimplepdfSimplePDF/simplepdf-embed407—~7kAutomated safety check: PassMIT

Similar skills

  • Jev SEO

    AgriciDaniel/jev-seo

    Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).

    527 GitHub stars~2.5k tokensUpdated 16 days ago
    Documents & OfficeAuto-check: notes
  • Analyzing Malicious PDF With Peepdf

    mukul975/Anthropic-Cybersecurity-Skills

    Perform static analysis of malicious PDF documents using peepdf, pdfid, and pdf-parser to extract embedded JavaScript, shellcode, and suspicious objects.

    34k GitHub stars~799 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • PDF Toolkit

    borghei/Claude-Skills

    Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Edit PDF

    SimplePDF/simplepdf-embed

    Edit and fill PDF documents. An agent skill from SimplePDF/simplepdf-embed.

    407 GitHub stars~1.5k tokensUpdated 8 days ago
    Documents & OfficeAuto-check passed
  • Build With Simplepdf

    SimplePDF/simplepdf-embed

    Integrate SimplePDF into a web application for PDF viewing, editing, filling, signing, programmatic control, AI-agent interaction, human-in-the-loop form prefilling, submissions, webhooks, or…

    407 GitHub stars~7k tokensUpdated 8 days ago
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about PDF Extraction Fallbacks

What does PDF Extraction Fallbacks do?

Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks. PDF Extraction Fallbacks is an agent skill from HKUDS/OpenSpace.

When should I use PDF Extraction Fallbacks?

PDF Extraction Fallbacks fits situations like: tasks that involve PDF.

How do I install PDF Extraction Fallbacks in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-extraction-fallbacks in HKUDS/OpenSpace) into .claude/skills/pdf-extraction-fallbacks in your project. Claude Code loads it when a task matches its description.

How do I install PDF Extraction Fallbacks in Codex?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-extraction-fallbacks in HKUDS/OpenSpace) into .agents/skills/pdf-extraction-fallbacks in your project. Codex loads it when a task matches its description.

Can I use PDF Extraction Fallbacks in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-extraction-fallbacks, .gemini/skills/pdf-extraction-fallbacks, .github/skills/pdf-extraction-fallbacks and .opencode/skills/pdf-extraction-fallbacks in your project.

What does PDF Extraction Fallbacks need to run?

Going by SKILL.md and its folder, PDF Extraction Fallbacks needs the command-line tools its instructions call (curl, pip, pdftotext, python3 and python). Our summary lists: Python 3.

Does PDF Extraction Fallbacks access the network?

SKILL.md contains no URLs. Its commands use curl and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Extraction Fallbacks safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Extraction Fallbacks use?

PDF Extraction Fallbacks is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Extraction Fallbacks use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Extraction Fallbacks?

Skills that share tags, products or a category with PDF Extraction Fallbacks: Jev SEO (AgriciDaniel/jev-seo, 527 stars), Analyzing Malicious PDF With Peepdf (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), PDF Toolkit (borghei/Claude-Skills, 881 stars) and Edit PDF (SimplePDF/simplepdf-embed, 407 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Extraction Fallbacks?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,749 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.