Agent skill

PDF Extract Ordered Fallback

by HKUDS in HKUDS/OpenSpace

PDF extraction with ordered tool chain: readfile, then runshell/pdftotext, then executecodesandbox/PyMuPDF

MITAuto-check passedDocuments & Office

Install PDF Extract Ordered Fallback

skills CLI
$ npx skills add HKUDS/OpenSpace --skill pdf-extract-ordered-fallback -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace pdf-extract-ordered-fallback --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced .claude/skills/pdf-extract-ordered-fallback && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-extract-ordered-fallback
GitHub stars
7.7k
Token cost
~2.6k tokens
SKILL.md length
934 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

PDF extraction with ordered tool chain: readfile, then runshell/pdftotext, then executecodesandbox/PyMuPDF

  • Works in 6 steps: Initial Extraction Attempt with… → Download PDF with Browser User-Agent → Verify File Type Before Parsing → …
  • Tasks that involve PDF
  • SKILL.md covers Overview, Ordered Tool Chain Summary, Step-by-Step Instructions and Complete Workflow Script, plus 5 more sections
  • Calls curl, apt-get and pdftotext

What it does

PDF Extract Ordered Fallback is an agent skill from HKUDS/OpenSpace. PDF extraction with ordered tool chain: readfile, then runshell/pdftotext, then executecodesandbox/PyMuPDF

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering PDF. It works with Python. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-extract-ordered-fallback”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Initial Extraction Attempt with read_file Tool
  2. Download PDF with Browser User-Agent
  3. Verify File Type Before Parsing
  4. Primary Shell Extraction with pdftotext via run_shell
  5. Secondary Python Fallback with PyMuPDF via execute_code_sandbox
  6. Graceful Degradation to Domain Knowledge

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • apt-get
    • pdftotext
    • brew
    • yum
    • pip
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Extract Ordered Fallback loads about 2.6k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 934 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 934 words, ~2,579 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-extract-ordered-fallback/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-extract-ordered-fallback
description
PDF extraction with ordered tool chain: read_file, then run_shell/pdftotext, then execute_code_sandbox/PyMuPDF

PDF Download and Extract with Ordered Fallback

This skill provides a robust workflow for acquiring PDF documents from web sources and extracting their text content, with a clearly ordered sequence of tool invocations to maximize success rate.

Overview

When working with PDFs from web sources, encounters with JavaScript redirects, corrupted files, missing tools, or inaccessible content are common. This workflow ensures maximum success rate through a严格 ordered fallback sequence that prioritizes shell-based tools over Python sandbox execution.

Ordered Tool Chain Summary

StepToolMethodPriority
0read_fileDirect PDF text extractionFirst attempt
1run_shellpdftotext commandPrimary fallback (if Step 0 returns binary/fails)
2execute_code_sandboxPyMuPDF Python librarySecondary fallback (if Step 1 fails)
3Domain knowledgeManual content generationLast resort

Key principle: Always try shell tools (run_shell) before Python sandbox (execute_code_sandbox) when both are viable options. Shell execution is more reliable in constrained environments.

Step-by-Step Instructions

Step 0: Initial Extraction Attempt with read_file Tool

First, attempt to extract PDF text using the read_file tool. This is the simplest approach and handles many PDFs correctly:

read_file filetype="pdf" file_path="path/to/document.pdf"

Expected outcomes:

  • Success: Returns extracted text content - proceed to use this directly
  • Binary/Image data returned: The tool failed to extract text; file content is raw binary or image data
    • Immediate action: Proceed to Step 1 (run_shell with pdftotext)
  • Error returned: Tool failed entirely; proceed to Step 1 (run_shell with pdftotext)

Critical: If read_file returns binary data (PNG/JPEG headers, raw PDF bytes), do NOT attempt to parse it manually. Immediately switch to shell-based pdftotext.

Step 1: Download PDF with Browser User-Agent

Many PDF hosting sites use JavaScript-based redirects or block automated requests. Use curl with a realistic browser user-agent:

bash
curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" -o output.pdf "URL_HERE"

Key flags:

  • -L: Follow redirects
  • -A: Set user-agent header to mimic a real browser
  • -o: Specify output filename
Step 2: Verify File Type Before Parsing

Always validate the downloaded file is actually a PDF before attempting extraction:

bash
file output.pdf

Expected output should contain "PDF document". If not:

  • The URL may have redirected to an HTML error page
  • The file may be corrupted
  • Access may be blocked
Step 3: Primary Shell Extraction with pdftotext via run_shell

This step takes priority over Python-based extraction. If Step 0 failed or if you're working with a newly downloaded PDF, use run_shell with pdftotext before attempting any Python libraries:

bash
pdftotext downloaded.pdf extracted.txt

Execute via run_shell:

run_shell command="pdftotext downloaded.pdf extracted.txt"

If pdftotext is not available, install it first:

bash
# Debian/Ubuntu
apt-get update && apt-get install -y poppler-utils

# macOS  
brew install poppler

# RHEL/CentOS
yum install -y poppler-utils

Why shell-first? Shell-based pdftotext is more reliable, faster, and avoids sandbox execution issues that can affect Python code execution in constrained environments.

Step 4: Secondary Python Fallback with PyMuPDF via execute_code_sandbox

Only if run_shell with pdftotext fails or is unavailable, fall back to Python's PyMuPDF library via execute_code_sandbox:

python
import fitz  # PyMuPDF

doc = fitz.open("downloaded.pdf")
text = ""
for page in doc:
    text += page.get_text()
doc.close()

with open("extracted.txt", "w") as f:
    f.write(text)

Execute within execute_code_sandbox:

execute_code_sandbox code="<Python code above>"

Install if needed:

bash
pip install pymupdf

Note: Some environments may experience execute_code_sandbox failures (unknown errors). This is why shell-based extraction (Step 3) must be attempted first.

Step 5: Graceful Degradation to Domain Knowledge

If the PDF cannot be accessed or extracted after all attempts:

  1. Document the failure mode (network issue, corrupted file, access denied, etc.)
  2. Extract any partial content that was successfully retrieved
  3. Supplement missing content from established domain knowledge
  4. Clearly mark which portions are from source vs. generated from knowledge
  5. Provide citations for any claimed requirements or specifications

Example degradation note:

NOTE: Source document [URL] was inaccessible due to [reason]. 
Content below combines partial extraction with established domain knowledge 
for [topic]. Verify against official sources when available.

Complete Workflow Script

bash
#!/bin/bash
# pdf-extract-workflow.sh

PDF_URL="$1"
OUTPUT_PDF="downloaded.pdf"
OUTPUT_TXT="extracted.txt"

# Step 0/1: Download with browser user-agent
echo "Downloading PDF..."
curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" -o "$OUTPUT_PDF" "$PDF_URL"

# Step 1: Verify file type
echo "Verifying file type..."
if ! file "$OUTPUT_PDF" | grep -q "PDF document"; then
    echo "WARNING: Downloaded file is not a valid PDF"
    echo "Attempting fallback extraction anyway..."
fi

# Step 2: Try pdftotext via shell (PRIMARY EXTRACTION)
echo "Attempting pdftotext extraction..."
if command -v pdftotext &> /dev/null; then
    if pdftotext "$OUTPUT_PDF" "$OUTPUT_TXT" 2>/dev/null; then
        echo "Extraction successful with pdftotext"
        exit 0
    fi
fi

# Step 3: Fallback to PyMuPDF via Python sandbox (SECONDARY)
echo "Falling back to PyMuPDF..."
python3 << 'PYTHON_SCRIPT'
import fitz
import sys

try:
    doc = fitz.open("downloaded.pdf")
    text = ""
    for page in doc:
        text += page.get_text()
    doc.close()
    with open("extracted.txt", "w") as f:
        f.write(text)
    print("Extraction successful with PyMuPDF")
    sys.exit(0)
except Exception as e:
    print(f"PyMuPDF failed: {e}")
    sys.exit(1)
PYTHON_SCRIPT

# Step 4: Handle complete failure
if [ $? -ne 0 ]; then
    echo "All extraction methods failed. Generate content from domain knowledge."
    echo "Document the failure and proceed with knowledge-based content generation."
fi

Agent-Specific Tool Invocation Pattern

For AI agents with access to specialized tools, follow this exact sequence:

# ITERATION 1: Try read_file first
read_file filetype="pdf" file_path="document.pdf"

# If read_file returns binary data or fails:
# ITERATION 2: Use run_shell with pdftotext
run_shell command="pdftotext document.pdf extracted.txt"

# If run_shell fails:
# ITERATION 3: Use execute_code_sandbox with PyMuPDF
execute_code_sandbox code="import fitz; doc = fitz.open('document.pdf'); ..."

# If all automated extraction fails:
# ITERATION 4+: Document failure mode and generate from domain knowledge

Critical anti-pattern to avoid: Do NOT attempt execute_code_sandbox before run_shell for PDF extraction. Shell tools are more reliable and should be prioritized.

Show full SKILL.md (360 more words)Show less

Failure Recovery from Observed Patterns

Based on execution analysis, here are common failure cascades and recovery strategies:

Pattern: read_file Returns Binary Data

Symptom: read_file returns PNG/JPEG image data or raw PDF bytes instead of text Cause: Tool cannot extract text from scanned PDFs or certain PDF structures
Recovery: Immediately switch to run_shell with pdftotext - do not attempt to parse binary data

Pattern: execute_code_sandbox Returns Unknown Error

Symptom: Multiple execute_code_sandbox calls fail with "unknown error" Cause: Sandbox execution environment issues or resource constraints Recovery: This is why run_shell must be attempted first - shell execution bypasses sandbox limitations

Pattern: read_webpage Fails on All URLs

Symptom: All read_webpage calls to domain URLs return errors Cause: Anti-bot measures, network issues, or site blocking Recovery: Focus on PDF extraction from locally downloaded files; supplement missing context from domain knowledge with clear citations

Best Practices

  1. Always verify before parsing: Never assume a downloaded file is valid
  2. Preserve original PDF: Keep the downloaded file for debugging if needed
  3. Log each step: Document which method succeeded for future reference
  4. Check extraction quality: Verify extracted text is readable and complete
  5. Cite source limitations: When using fallback knowledge, clearly indicate source gaps
  6. Follow tool ordering: read_file → run_shell → execute_code_sandbox → domain knowledge
  7. Shell before Python: Prioritize run_shell over execute_code_sandbox when both are viable
  8. Detect binary early: If read_file returns non-text data, immediately switch to shell tools

Common Failure Modes

SymptomCauseSolution
HTML content in PDFURL redirected to error pageCheck HTTP status, try alternate URL
Empty extractionPassword-protected or scanned PDFTry OCR tools or request accessible version
Garbled textEncoding issuesTry PyMuPDF with different extraction mode
read_file returns binaryScanned PDF or tool limitationImmediately use run_shell with pdftotext
execute_code_sandbox unknown errorSandbox execution failureThis is why run_shell should be tried first
Curl blockedAnti-bot measuresAdd more headers, use delay between requests

When to Use This Skill

  • Downloading regulatory documents from government websites
  • Extracting content from technical manuals or handbooks
  • Processing PDFs in automated pipelines where reliability matters
  • Situations where tool execution constraints may limit Python sandbox availability
  • Any situation where PDF access may be unreliable or restricted

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

PDF Extract Ordered Fallback next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Extract Ordered Fallback compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Extract Ordered Fallback this skillHKUDS/OpenSpace7.7k—~2.6kAutomated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Software Certificate SkillIvanCodesDev/software-certificate-skill156—~1.6kAutomated safety check: PassMIT
Doc Cleanernotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Office To Mdshuyu-labs/WebCode278—~1kAutomated safety check: NotesCustom licence

Similar skills

  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Software Certificate Skill

    IvanCodesDev/software-certificate-skill

    面向普通用户,从真实软件项目全自动生成中国软件著作权申请资料:一次收集登记事实,自动分析业务、选择可追溯源码、取得真实界面证据,生成申请表信息、规范黑白灰操作手册、代码前后30页或全部材料及真实 DOCX/PDF;内部验证、渲染、哈希与备份只进入系统临时运行区,项目最终仅保留正式资料。适配 Codex、Claude…

    156 GitHub stars~1.6k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Doc Cleaner

    notoriouslab/doc-cleaner

    Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

    309 GitHub stars~712 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Office To Md

    shuyu-labs/WebCode

    Convert Office documents (Word, Excel, PowerPoint, PDF) to Markdown format.

    278 GitHub stars~1k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check: notes
  • Cc Streaming Export Safety

    doccker/cc-use-exp

    当实现用户驱动的大文件导出或批量序列化(Excel/CSV/JSON/JSONL/PDF,数据量未知或超过 1 万行/10 MB)时触发;普通小文件下载、静态资源下载、非导出 Writer/Report 类不触发。防止 OOM、临时文件残留、同步导出阻塞 HTTP 线程和表格公式注入。

    1.1k GitHub stars~2.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about PDF Extract Ordered Fallback

What does PDF Extract Ordered Fallback do?

PDF extraction with ordered tool chain: readfile, then runshell/pdftotext, then executecodesandbox/PyMuPDF. PDF Extract Ordered Fallback is an agent skill from HKUDS/OpenSpace.

When should I use PDF Extract Ordered Fallback?

PDF Extract Ordered Fallback fits situations like: tasks that involve PDF.

How do I install PDF Extract Ordered Fallback in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extract-ordered-fallback -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced in HKUDS/OpenSpace) into .claude/skills/pdf-extract-ordered-fallback in your project. Claude Code loads it when a task matches its description.

How do I install PDF Extract Ordered Fallback in Codex?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extract-ordered-fallback -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced in HKUDS/OpenSpace) into .agents/skills/pdf-extract-ordered-fallback in your project. Codex loads it when a task matches its description.

Can I use PDF Extract Ordered Fallback in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill pdf-extract-ordered-fallback -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-extract-ordered-fallback, .gemini/skills/pdf-extract-ordered-fallback, .github/skills/pdf-extract-ordered-fallback and .opencode/skills/pdf-extract-ordered-fallback in your project.

What does PDF Extract Ordered Fallback need to run?

Going by SKILL.md and its folder, PDF Extract Ordered Fallback needs the command-line tools its instructions call (curl, apt-get, pdftotext, brew, yum and pip). Our summary lists: Python 3.

Does PDF Extract Ordered Fallback access the network?

SKILL.md contains no URLs. Its commands use curl and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Extract Ordered Fallback safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Extract Ordered Fallback use?

PDF Extract Ordered Fallback is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Extract Ordered Fallback use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Extract Ordered Fallback?

Skills that share tags, products or a category with PDF Extract Ordered Fallback: Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Software Certificate Skill (IvanCodesDev/software-certificate-skill, 156 stars), Doc Cleaner (notoriouslab/doc-cleaner, 309 stars) and Mineru (Nebutra/MinerU-Skill, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Extract Ordered Fallback?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,743 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.