Agent skill

PDF Extract Shell First

by HKUDS in HKUDS/OpenSpace

PDF text extraction with tool cascade prioritizing shell pdftotext before Python fallback

MITAuto-check passedDocuments & Office

Install PDF Extract Shell First

skills CLI
$ npx skills add HKUDS/OpenSpace --skill pdf-extract-shell-first -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace pdf-extract-shell-first --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-e27e0c .claude/skills/pdf-extract-shell-first && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-extract-shell-first
GitHub stars
7.7k
Token cost
~2.6k tokens
SKILL.md length
739 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

PDF text extraction with tool cascade prioritizing shell pdftotext before Python fallback

  • Works in 5 steps: Download PDF (URL Only) → Try read_file (Primary Attempt) → Use run_shell with pdftotext (Preferred… → …
  • Tasks that involve PDF
  • SKILL.md covers Why Shell-First?, Entry Point: Determine Your…, Complete Workflow and Tool Selection Decision Tree, plus 5 more sections
  • Calls curl, apt-get and python3

What it does

PDF Extract Shell First is an agent skill from HKUDS/OpenSpace. PDF text extraction with tool cascade prioritizing shell pdftotext before Python fallback

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering PDF. It works with Python. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-extract-shell-first”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Download PDF (URL Only)
  2. Try read_file (Primary Attempt)
  3. Use run_shell with pdftotext (Preferred Fallback)
  4. Use execute_code_sandbox with PyMuPDF (Last Resort)
  5. Graceful Degradation to Domain Knowledge

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • apt-get
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Extract Shell First loads about 2.6k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 739 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 739 words, ~2,634 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-extract-shell-first/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-extract-shell-first
description
PDF text extraction with tool cascade prioritizing shell pdftotext before Python fallback

PDF Extract with Shell-First Tool Cascade

This skill provides an optimized workflow for extracting text content from PDF documents (local files or downloaded URLs) using a prioritized tool cascade that favors shell-based extraction before falling back to Python libraries.

Why Shell-First?

Analysis of execution patterns shows:

  • read_file on PDFs sometimes returns binary/image data instead of text
  • run_shell with pdftotext has higher success rate and fewer sandbox errors
  • execute_code_sandbox can fail with "unknown error" in constrained environments
  • Shell tools are more reliable for PDF text extraction when available

Entry Point: Determine Your Starting Point

Before beginning, identify your scenario:

ScenarioStart HereSkip
PDF already on local diskStep 1 (Try read_file)Shell download steps
PDF at a web URLShell download, then Step 1None
Need maximum reliabilityFull cascade (all 3 tools)None

Complete Workflow

Step 0: Download PDF (URL Only)

If your PDF is at a web URL, download it first using browser user-agent:

bash
curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" -o target.pdf "URL_HERE"

Key flags:

  • -L: Follow redirects
  • -A: Set user-agent header to mimic a real browser
  • -o: Specify output filename

If you already have the PDF locally, skip to Step 1.

Step 1: Try read_file (Primary Attempt)

First, attempt to extract text using the read_file tool:

read_file(filetype="pdf", file_path="target.pdf")

Evaluate the response:

Response TypeInterpretationNext Action
Clean readable textSuccessProceed to content analysis
Binary data / PNG image / garbledread_file returned raw dataGo to Step 2 immediately
Error / timeoutTool failureGo to Step 2 immediately

Critical: If read_file returns binary image data or garbled content, do not retry read_file. Immediately proceed to Step 2.

Step 2: Use run_shell with pdftotext (Preferred Fallback)

When read_file fails or returns binary data, use run_shell with pdftotext:

bash
run_shell(command="pdftotext target.pdf output.txt")

Then read the extracted text:

read_file(filetype="txt", file_path="output.txt")

If pdftotext is not found, install it first:

bash
run_shell(command="apt-get update && apt-get install -y poppler-utils")
# Or for macOS:
run_shell(command="brew install poppler")

Then retry:

bash
run_shell(command="pdftotext target.pdf output.txt")

Verify extraction quality:

  • Check that output.txt exists and has content
  • Sample the text to ensure it's readable (not garbled)
  • If extraction looks corrupted, proceed to Step 3
Step 3: Use execute_code_sandbox with PyMuPDF (Last Resort)

If pdftotext is unavailable or produces poor results, use Python's PyMuPDF via execute_code_sandbox:

python
import fitz  # PyMuPDF

doc = fitz.open("target.pdf")
text = ""
for page in doc:
    text += page.get_text()
doc.close()

with open("output.txt", "w") as f:
    f.write(text)
print(f"Extracted {len(text)} characters from {len(doc)} pages")

Execute via:

execute_code_sandbox(code="<python code above>")

Then read the result:

read_file(filetype="txt", file_path="output.txt")

Note: execute_code_sandbox may fail with "unknown error" in some environments. If this occurs, document the failure and proceed to Step 4.

Step 4: Graceful Degradation to Domain Knowledge

If all extraction methods fail:

  1. Document the specific failure mode for each tool attempted
  2. Extract any partial content that was successfully retrieved
  3. Supplement missing content from established domain knowledge
  4. Clearly mark which portions are from source vs. generated from knowledge
  5. Provide citations for any claimed requirements or specifications

Example degradation documentation:

EXTRACTION FAILURE REPORT:
- Source: [URL or file path]
- read_file: Returned binary/image data (no text extraction)
- run_shell/pdftotext: [Tool not available / produced garbled output / succeeded]
- execute_code_sandbox/PyMuPDF: [Failed with unknown error / succeeded]

NOTE: Content below combines partial extraction with established domain 
knowledge for [topic]. Verify against official sources when available.

Tool Selection Decision Tree

                    PDF to Extract
                          │
                          ▼
                  ┌───────────────┐
                  │  read_file    │
                  │  (primary)    │
                  └───────┬───────┘
                          │
            ┌─────────────┼─────────────┐
            │             │             │
     Returns text   Returns binary   Error/timeout
        (✓)         / image data         │
            │             │             │
            ▼             ▼             ▼
       SUCCESS    ┌───────────────┐
                  │ run_shell     │
                  │ pdftotext     │
                  └───────┬───────┘
                          │
                  ┌───────┼───────┐
                  │       │       │
             Succeeds  Not      Garbled
                (✓)   avail.    output
                  │       │       │
                  ▼       ▼       ▼
             SUCCESS ┌───────────────┐
                     │ execute_code  │
                     │ _sandbox      │
                     │ PyMuPDF       │
                     └───────┬───────┘
                             │
                     ┌───────┼───────┐
                     │       │       │
                Succeeds   Fails   Error
                   (✓)      │       │
                     │      ▼       │
                     ▼   Domain     │
                 SUCCESS  Knowledge │
                             │      │
                             └──────┘
                              FAILURE
                              DOCUMENTED

Complete Automated Script

bash
#!/bin/bash
# pdf-extract-cascade.sh
# Implements the full tool cascade for PDF extraction

INPUT="$1"
OUTPUT_PDF="target.pdf"
OUTPUT_TXT="output.txt"

# Step 0: Handle URL vs local file
if [[ "$INPUT" =~ ^https?:// ]]; then
    echo "Downloading PDF from URL..."
    curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" -o "$OUTPUT_PDF" "$INPUT"
else
    if [ ! -f "$INPUT" ]; then
        echo "ERROR: Local file not found: $INPUT"
        exit 1
    fi
    OUTPUT_PDF="$INPUT"
fi

# Step 1: Verify file type
echo "Verifying file type..."
if ! file "$OUTPUT_PDF" | grep -q "PDF document"; then
    echo "WARNING: File is not a valid PDF"
    file "$OUTPUT_PDF"
fi

# Step 2: Try pdftotext (shell-first approach)
echo "Attempting pdftotext extraction..."
if command -v pdftotext &> /dev/null; then
    if pdftotext "$OUTPUT_PDF" "$OUTPUT_TXT" 2>/dev/null; then
        if [ -s "$OUTPUT_TXT" ]; then
            echo "SUCCESS: Extraction completed with pdftotext"
            wc -l "$OUTPUT_TXT"
            exit 0
        fi
    fi
fi

# Step 3: Fallback to PyMuPDF
echo "Falling back to PyMuPDF..."
python3 << 'PYTHON_SCRIPT'
import fitz
import sys

try:
    doc = fitz.open("target.pdf")
    text = ""
    for page in doc:
        text += page.get_text()
    doc.close()
    with open("output.txt", "w") as f:
        f.write(text)
    print(f"SUCCESS: Extracted {len(text)} characters from {len(doc)} pages")
    sys.exit(0)
except Exception as e:
    print(f"PyMuPDF failed: {e}")
    sys.exit(1)
PYTHON_SCRIPT

# Step 4: Handle complete failure
if [ $? -ne 0 ]; then
    echo "FAILURE: All extraction methods failed"
    echo "ACTION: Generate content from domain knowledge"
    echo "Document each tool's failure mode for future reference"
    exit 1
fi
Show full SKILL.md (302 more words)Show less

Best Practices

  1. Never retry read_file on binary response: If read_file returns image/binary data, immediately switch to run_shell
  2. Prefer run_shell over execute_code_sandbox: Shell tools have higher reliability and fewer sandbox-related errors
  3. Verify before trusting: Always check extracted text is readable, not just that the command succeeded
  4. Document failures: Record which tools failed and how, to inform future extraction attempts
  5. Preserve originals: Keep the source PDF for debugging and re-extraction if needed
  6. Check tool availability early: Test for pdftotext before starting complex workflows

Common Failure Modes and Responses

SymptomLikely CauseRecommended Action
read_file returns PNG/binaryPDF rendered as image, not parsedImmediately use run_shell with pdftotext
pdftotext: command not foundpoppler-utils not installedRun apt-get install poppler-utils first
pdftotext produces empty filePassword-protected or scanned PDFTry PyMuPDF, or use OCR tools
execute_code_sandbox "unknown error"Sandbox execution issueDocument failure, use domain knowledge fallback
Garbled text outputEncoding issuesTry PyMuPDF with page.get_text("text")
All tools failSeverely corrupted or encrypted PDFDocument limitation, use knowledge-based content

When to Use This Skill

  • Extracting text from local PDF files where read_file may return binary data
  • Processing PDFs in automated pipelines requiring high reliability
  • Situations where execute_code_sandbox has shown instability
  • Working with PDFs from sources that may deliver rendered images instead of parseable text
  • Any workflow where shell tool availability can be assumed or easily installed

Migration Notes from pdf-download-extract-fallback

This skill (pdf-extract-shell-first) differs from the parent in these key ways:

  1. Explicit tool sequencing: Clearly prioritizes read_file → run_shell → execute_code_sandbox
  2. No retry on binary read_file: Instructs immediate fallback when binary data detected
  3. Shell-first philosophy: Emphasizes pdftotext via run_shell as preferred over Python
  4. Reduced download focus: Assumes PDF is available or downloads in pre-step; focuses on extraction cascade
  5. Decision tree visualization: Provides clear flowchart for tool selection

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-e27e0c of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

PDF Extract Shell First next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Extract Shell First compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Extract Shell First this skillHKUDS/OpenSpace7.7k—~2.6kAutomated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Software Certificate SkillIvanCodesDev/software-certificate-skill156—~1.6kAutomated safety check: PassMIT
Doc Cleanernotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Office To Mdshuyu-labs/WebCode278—~1kAutomated safety check: NotesCustom licence

Similar skills

  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Software Certificate Skill

    IvanCodesDev/software-certificate-skill

    面向普通用户,从真实软件项目全自动生成中国软件著作权申请资料:一次收集登记事实,自动分析业务、选择可追溯源码、取得真实界面证据,生成申请表信息、规范黑白灰操作手册、代码前后30页或全部材料及真实 DOCX/PDF;内部验证、渲染、哈希与备份只进入系统临时运行区,项目最终仅保留正式资料。适配 Codex、Claude…

    156 GitHub stars~1.6k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Doc Cleaner

    notoriouslab/doc-cleaner

    Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

    309 GitHub stars~712 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Office To Md

    shuyu-labs/WebCode

    Convert Office documents (Word, Excel, PowerPoint, PDF) to Markdown format.

    278 GitHub stars~1k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check: notes
  • Cc Streaming Export Safety

    doccker/cc-use-exp

    当实现用户驱动的大文件导出或批量序列化(Excel/CSV/JSON/JSONL/PDF,数据量未知或超过 1 万行/10 MB)时触发;普通小文件下载、静态资源下载、非导出 Writer/Report 类不触发。防止 OOM、临时文件残留、同步导出阻塞 HTTP 线程和表格公式注入。

    1.1k GitHub stars~2.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about PDF Extract Shell First

What does PDF Extract Shell First do?

PDF text extraction with tool cascade prioritizing shell pdftotext before Python fallback. PDF Extract Shell First is an agent skill from HKUDS/OpenSpace.

When should I use PDF Extract Shell First?

PDF Extract Shell First fits situations like: tasks that involve PDF.

How do I install PDF Extract Shell First in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extract-shell-first -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-e27e0c in HKUDS/OpenSpace) into .claude/skills/pdf-extract-shell-first in your project. Claude Code loads it when a task matches its description.

How do I install PDF Extract Shell First in Codex?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extract-shell-first -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-e27e0c in HKUDS/OpenSpace) into .agents/skills/pdf-extract-shell-first in your project. Codex loads it when a task matches its description.

Can I use PDF Extract Shell First in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill pdf-extract-shell-first -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-extract-shell-first, .gemini/skills/pdf-extract-shell-first, .github/skills/pdf-extract-shell-first and .opencode/skills/pdf-extract-shell-first in your project.

What does PDF Extract Shell First need to run?

Going by SKILL.md and its folder, PDF Extract Shell First needs the command-line tools its instructions call (curl, apt-get and python3). Our summary lists: Python 3.

Does PDF Extract Shell First access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Extract Shell First safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Extract Shell First use?

PDF Extract Shell First is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Extract Shell First use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Extract Shell First?

Skills that share tags, products or a category with PDF Extract Shell First: Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Software Certificate Skill (IvanCodesDev/software-certificate-skill, 156 stars), Doc Cleaner (notoriouslab/doc-cleaner, 309 stars) and Mineru (Nebutra/MinerU-Skill, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Extract Shell First?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,743 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.