Agent skill

PDF Download Extract Fallback

by HKUDS in HKUDS/OpenSpace

Multi-step PDF download and text extraction with progressive fallback strategies

MITAuto-check passedDocuments & Office

Install PDF Download Extract Fallback

skills CLI
$ npx skills add HKUDS/OpenSpace --skill pdf-download-extract-fallback -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace pdf-download-extract-fallback --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-download-extract-fallback .claude/skills/pdf-download-extract-fallback && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-download-extract-fallback
GitHub stars
7.8k
Token cost
~2.4k tokens
SKILL.md length
800 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Multi-step PDF download and text extraction with progressive fallback strategies

  • Works in 6 steps: Download PDF with Browser User-Agent → Verify File Type Before Parsing → Primary Extraction with pdftotext → …
  • Tasks that involve PDF
  • SKILL.md covers Overview, Step-by-Step Instructions, Complete Workflow Script and Best Practices, plus 9 more sections
  • Calls apt-get, curl and pip

What it does

PDF Download Extract Fallback is an agent skill from HKUDS/OpenSpace. Multi-step PDF download and text extraction with progressive fallback strategies

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering PDF. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-download-extract-fallback”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Download PDF with Browser User-Agent
  2. Verify File Type Before Parsing
  3. Primary Extraction with pdftotext
  4. Fallback to PyMuPDF (fitz)
  5. Graceful Degradation to Domain Knowledge
  6. Download PDF with Browser User-Agent

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • apt-get
    • curl
    • pip
    • brew
    • yum
    • python3
    • pdftotext

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Download Extract Fallback loads about 2.4k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 800 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 800 words, ~2,401 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-download-extract-fallback/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-download-extract-fallback
description
Multi-step PDF download and text extraction with progressive fallback strategies

PDF Download and Extract with Fallback

This skill provides a robust workflow for acquiring PDF documents from web sources and extracting their text content, with multiple fallback mechanisms to handle various failure modes.

Overview

When working with PDFs from web sources, encounters with JavaScript redirects, corrupted files, missing tools, or inaccessible content are common. This workflow ensures maximum success rate through progressive fallback strategies.

Step-by-Step Instructions

Step 1: Download PDF with Browser User-Agent

Many PDF hosting sites use JavaScript-based redirects or block automated requests. Use curl with a realistic browser user-agent:

bash
curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" -o output.pdf "URL_HERE"

Key flags:

  • -L: Follow redirects
  • -A: Set user-agent header to mimic a real browser
  • -o: Specify output filename
Step 2: Verify File Type Before Parsing

Always validate the downloaded file is actually a PDF before attempting extraction:

bash
file output.pdf

Expected output should contain "PDF document". If not:

  • The URL may have redirected to an HTML error page
  • The file may be corrupted
  • Access may be blocked
Step 3: Primary Extraction with pdftotext

First attempt extraction using the standard pdftotext utility (part of poppler-utils):

bash
pdftotext output.pdf output.txt

If pdftotext is not available, install it:

bash
# Debian/Ubuntu
apt-get update && apt-get install -y poppler-utils

# macOS
brew install poppler

# RHEL/CentOS
yum install -y poppler-utils
Step 4: Fallback to PyMuPDF (fitz)

If pdftotext fails or produces poor results, use Python's PyMuPDF library:

python
import fitz  # PyMuPDF

doc = fitz.open("output.pdf")
text = ""
for page in doc:
    text += page.get_text()
doc.close()

with open("output.txt", "w") as f:
    f.write(text)

Install if needed:

bash
pip install pymupdf
Step 5: Graceful Degradation to Domain Knowledge

If the PDF cannot be accessed or extracted after all attempts:

  1. Document the failure mode (network issue, corrupted file, access denied, etc.)
  2. Extract any partial content that was successfully retrieved
  3. Supplement missing content from established domain knowledge
  4. Clearly mark which portions are from source vs. generated from knowledge
  5. Provide citations for any claimed requirements or specifications

Example degradation note:

NOTE: Source document [URL] was inaccessible due to [reason]. 
Content below combines partial extraction with established domain knowledge 
for [topic]. Verify against official sources when available.

Complete Workflow Script

bash
#!/bin/bash
# pdf-extract-workflow.sh
# pdf-extract-workflow.sh - Handles both URL downloads and local files

INPUT="$1"
OUTPUT_PDF="downloaded.pdf"
OUTPUT_TXT="extracted.txt"

if [[ "$INPUT" =~ ^https?:// ]]; then
    # Mode A: URL download
    PDF_URL="$INPUT"
    echo "Downloading PDF from URL..."
    curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" -o "$OUTPUT_PDF" "$PDF_URL"
else
    # Mode B: Local file
    if [ ! -f "$INPUT" ]; then
        echo "ERROR: Local file not found: $INPUT"
        exit 1
    fi
    OUTPUT_PDF="$INPUT"
    echo "Using local file: $INPUT"
fi

# Step 2: Verify file type
echo "Verifying file type..."
if ! file "$OUTPUT_PDF" | grep -q "PDF document"; then
    echo "WARNING: Downloaded file is not a valid PDF"
    echo "Attempting fallback extraction anyway..."
fi

# Step 3: Try pdftotext
echo "Attempting pdftotext extraction..."
if command -v pdftotext &> /dev/null; then
    if pdftotext "$OUTPUT_PDF" "$OUTPUT_TXT" 2>/dev/null; then
        echo "Extraction successful with pdftotext"
        exit 0
    fi
fi

# Step 4: Fallback to PyMuPDF
echo "Falling back to PyMuPDF..."
python3 << 'PYTHON_SCRIPT'
import fitz
import sys

try:
    doc = fitz.open("downloaded.pdf")
    text = ""
    for page in doc:
        text += page.get_text()
    doc.close()
    with open("extracted.txt", "w") as f:
        f.write(text)
    print("Extraction successful with PyMuPDF")
    sys.exit(0)
except Exception as e:
    print(f"PyMuPDF failed: {e}")
    sys.exit(1)
PYTHON_SCRIPT

# Step 5: Handle complete failure
# Step 5: Handle complete failure (domain knowledge fallback)
if [ $? -ne 0 ]; then
    echo "ERROR: All extraction methods failed."
    echo "ACTION: Generate content from domain knowledge and clearly mark source limitations."
    echo "Document the failure and proceed with knowledge-based content generation."
fi

Best Practices

  1. Always verify before parsing: Never assume a downloaded file is valid
  2. Pre-check tools before extraction: Verify pdftotext or PyMuPDF availability before starting
  3. Local files skip download: When PDF is already on disk, begin at file validation step
  4. Preserve original PDF: Keep the downloaded file for debugging if needed
  5. Log each step: Document which method succeeded for future reference
  6. Check extraction quality: Verify extracted text is readable and complete
  7. Cite source limitations: When using fallback knowledge, clearly indicate source gaps

Common Failure Modes

SymptomCauseSolution
HTML content in fileURL redirected to error page or wrong file typeCheck HTTP status, verify file with file command
Empty extractionPassword-protected or scanned PDFTry OCR tools or request accessible version
Garbled textEncoding issuesTry PyMuPDF with different extraction mode
Curl blockedAnti-bot measuresAdd more headers, use delay between requests
pdftotext not foundTool not installedRun apt-get install poppler-utils or use PyMuPDF fallback
PyMuPDF import failedPackage not installedRun pip install pymupdf
File not found (local)Incorrect path or file not accessibleVerify file path, check permissions, confirm file was uploaded
Show full SKILL.md (335 more words)Show less

When to Use This Skill

  • Mode A (URL download): Downloading documents from web sources
  • Mode B (Local file): Processing PDFs already on disk or uploaded as reference files
  • Extracting content from technical manuals or handbooks
  • Processing PDFs in automated pipelines where reliability matters
  • Any situation where PDF access may be unreliable or restricted

Local File Processing Workflow

When you already have the PDF file locally (not from a URL):

Step L1: Verify File Exists
bash
if [ ! -f "your_file.pdf" ]; then
    echo "ERROR: File not found"
    echo "ACTION: Verify the file path and that the file was successfully uploaded"
    exit 1
fi
Step L2: Validate File Type
bash
file your_file.pdf

Expected output should contain "PDF document". If not, the file may be corrupted or mislabeled.

Step L3: Proceed to Extraction

After validation, skip directly to Step 3: Primary Extraction with pdftotext in the main workflow.

Tool Availability Pre-Check

Before attempting any PDF extraction, verify your environment has the necessary tools:

bash
# Check pdftotext availability
command -v pdftotext && echo "pdftotext: AVAILABLE" || echo "pdftotext: NOT FOUND - install poppler-utils"

# Check PyMuPDF availability  
python3 -c "import fitz; print('PyMuPDF: AVAILABLE')" 2>/dev/null || echo "PyMuPDF: NOT FOUND - run: pip install pymupdf"

Installation commands if tools are missing:

bash
# Install pdftotext (poppler-utils)
apt-get update && apt-get install -y poppler-utils  # Debian/Ubuntu
yum install -y poppler-utils                        # RHEL/CentOS
brew install poppler                                # macOS

# Install PyMuPDF
pip install pymupdf

name: pdf-download-extract-fallback description: Multi-step PDF download and text extraction with progressive fallback strategies

This skill provides a robust workflow for acquiring PDF documents from web sources or processing locally-available files and extracting their text content, with multiple fallback mechanisms to handle various failure modes.

Entry Point: Determine Your Starting Point

Before beginning, identify your scenario:

ScenarioStart HereSkip
PDF already on local diskStep 2 (Verify File Type)Step 1 (Download)
PDF at a web URLStep 1 (Download)None

Overview

When working with PDFs from web sources or local files, encounters with corrupted files, missing tools, or inaccessible content are common. This workflow ensures maximum success rate through progressive fallback strategies.

Mode A: Web URL Download

Step 1: Download PDF with Browser User-Agent

Many PDF hosting sites use JavaScript-based redirects or block automated requests. Use curl with a realistic browser user-agent:

bash
curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" -o output.pdf "URL_HERE"

Key flags:

  • -L: Follow redirects
  • -A: Set user-agent header to mimic a real browser
  • -o: Specify output filename

Mode B: Local File Processing

If you already have the PDF file locally, skip Step 1 and begin here:

Always validate the file is actually a PDF before attempting extraction:

Complete Workflow Script (Handles Both URL and Local File)

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/pdf-download-extract-fallback of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

PDF Download Extract Fallback next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Download Extract Fallback compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Download Extract Fallback this skillHKUDS/OpenSpace7.8k—~2.4kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill4k—~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice9.2k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3161 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    4k GitHub stars~2.2k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    9.2k GitHub stars~19k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    316 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.8k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.8k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.8k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.8k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about PDF Download Extract Fallback

What does PDF Download Extract Fallback do?

Multi-step PDF download and text extraction with progressive fallback strategies. PDF Download Extract Fallback is an agent skill from HKUDS/OpenSpace.

When should I use PDF Download Extract Fallback?

PDF Download Extract Fallback fits situations like: tasks that involve PDF.

How do I install PDF Download Extract Fallback in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill pdf-download-extract-fallback -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback in HKUDS/OpenSpace) into .claude/skills/pdf-download-extract-fallback in your project. Claude Code loads it when a task matches its description.

How do I install PDF Download Extract Fallback in Codex?

Run `npx skills add HKUDS/OpenSpace --skill pdf-download-extract-fallback -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback in HKUDS/OpenSpace) into .agents/skills/pdf-download-extract-fallback in your project. Codex loads it when a task matches its description.

Can I use PDF Download Extract Fallback in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill pdf-download-extract-fallback -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-download-extract-fallback, .gemini/skills/pdf-download-extract-fallback, .github/skills/pdf-download-extract-fallback and .opencode/skills/pdf-download-extract-fallback in your project.

What does PDF Download Extract Fallback need to run?

Going by SKILL.md and its folder, PDF Download Extract Fallback needs the command-line tools its instructions call (apt-get, curl, pip, brew, yum and python3). Our summary lists: Python 3.

Does PDF Download Extract Fallback access the network?

SKILL.md contains no URLs. Its commands use curl and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Download Extract Fallback safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Download Extract Fallback use?

PDF Download Extract Fallback is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Download Extract Fallback use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Download Extract Fallback?

Skills that share tags, products or a category with PDF Download Extract Fallback: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 4k stars), GenOffice Document CLI (genspark-ai/genoffice, 9.2k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Download Extract Fallback?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,754 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.