Jev SEO
AgriciDaniel/jev-seo
Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).
Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks .claude/skills/pdf-extraction-fallbacks && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pdf-extraction-fallbacks" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks into .claude/skills/pdf-extraction-fallbacks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacksType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks .agents/skills/pdf-extraction-fallbacks && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pdf-extraction-fallbacks" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks into .agents/skills/pdf-extraction-fallbacks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks .cursor/skills/pdf-extraction-fallbacks && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pdf-extraction-fallbacks" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks into .cursor/skills/pdf-extraction-fallbacks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/HKUDS/OpenSpace.git --path benchmarks/gdpval/skills/pdf-extraction-fallbacks--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks .gemini/skills/pdf-extraction-fallbacks && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pdf-extraction-fallbacks" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks into .gemini/skills/pdf-extraction-fallbacks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacksInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .github/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks .github/skills/pdf-extraction-fallbacks && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pdf-extraction-fallbacks" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks into .github/skills/pdf-extraction-fallbacks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks .opencode/skills/pdf-extraction-fallbacks && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pdf-extraction-fallbacks" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks into .opencode/skills/pdf-extraction-fallbacks/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pdf-extraction-fallbacksMulti-fallback PDF/text extraction with early failure detection and sequential tool fallbacks
PDF Extraction Fallbacks is an agent skill from HKUDS/OpenSpace. Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.
It sits in Documents & Office, covering PDF. It works with JavaScript. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlpippdftotextpython3pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF Extraction Fallbacks loads about 1.8k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 315 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 315 words, ~1,830 tokens.
.claude/skills/pdf-extraction-fallbacks/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.When extracting text from PDFs (especially regulatory documents, handbooks, or protected content), single-method approaches often fail due to JavaScript protection, CORS restrictions, encoding issues, or corrupted downloads. This skill provides a robust multi-fallback workflow that detects failures early and tries sequential extraction methods.
Before attempting extraction, validate the downloaded file:
# Download the PDF
curl -L -o document.pdf "$URL"
# Check file size (reject if < 1KB - likely error page)
FILE_SIZE=$(stat -f%z document.pdf 2>/dev/null || stat -c%s document.pdf 2>/dev/null)
if [ "$FILE_SIZE" -lt 1024 ]; then
echo "ERROR: File too small ($FILE_SIZE bytes) - likely not a valid PDF"
# Check if it's an HTML error page
head -c 200 document.pdf | grep -i "<html\|<!doctype\|error\|access denied" && \
echo "Detected HTML error page instead of PDF"
exit 1
fi
# Check PDF magic bytes
HEAD_BYTES=$(head -c 4 document.pdf)
if [ "$HEAD_BYTES" != "%PDF" ]; then
echo "ERROR: File does not start with PDF magic bytes"
head -c 100 document.pdf
exit 1
fi# Try pdftotext first (fastest, most reliable for simple PDFs)
if command -v pdftotext &> /dev/null; then
pdftotext -layout document.pdf output.txt 2>/dev/null
if [ -s output.txt ]; then
WORD_COUNT=$(wc -w < output.txt)
if [ "$WORD_COUNT" -gt 50 ]; then
echo "SUCCESS: pdftotext extracted $WORD_COUNT words"
exit 0
fi
fi
fi# Try PyMuPDF - handles more complex PDFs
import fitz # pymupdf
try:
doc = fitz.open("document.pdf")
text = ""
for page in doc:
text += page.get_text()
if len(text.strip()) > 500: # Sanity check
with open("output.txt", "w") as f:
f.write(text)
print(f"SUCCESS: PyMuPDF extracted {len(text)} characters")
else:
print("WARNING: PyMuPDF extraction too short, trying next method")
except Exception as e:
print(f"PyMuPDF failed: {e}")# Try pdfplumber - better for tables and structured content
import pdfplumber
try:
text = ""
with pdfplumber.open("document.pdf") as pdf:
for page in pdf.pages:
page_text = page.extract_text()
if page_text:
text += page_text + "\n"
if len(text.strip()) > 500:
with open("output.txt", "w") as f:
f.write(text)
print(f"SUCCESS: pdfplumber extracted {len(text)} characters")
else:
print("WARNING: pdfplumber extraction too short")
except Exception as e:
print(f"pdfplumber failed: {e}")If all methods fail, the PDF may be JavaScript-protected:
# Check for JavaScript in PDF
import fitz
doc = fitz.open("document.pdf")
has_js = False
for page in doc:
if page.get_java_script():
has_js = True
break
if has_js:
print("WARNING: PDF contains JavaScript - may be protected")
# Try rendering pages as images and OCR (requires additional tools)
# Or try alternative download sourceIf the primary URL fails:
.gov mirrors, archive.org)Download PDF
│
├─→ File < 1KB? → REJECT (likely error page)
├─→ No %PDF header? → REJECT (not a PDF)
│
└─→ Valid PDF
│
├─→ pdftotext → >50 words? → SUCCESS
│ └─→ Try next
│
├─→ PyMuPDF → >500 chars? → SUCCESS
│ └─→ Try next
│
├─→ pdfplumber → >500 chars? → SUCCESS
│ └─→ Try next
│
└─→ All failed → Check for JS protection, try alternative sources| Symptom | Likely Cause | Solution |
|---|---|---|
| File < 100 bytes | JavaScript error page | Check CORS, try different user-agent |
| File ~1-5KB | HTML error/warning page | Parse HTML for actual PDF link |
| pdftotext returns empty | Encrypted/protected PDF | Try PyMuPDF with password handling |
| Garbled text output | Encoding issue | Try pdfplumber, specify encoding |
| Extraction very short | Images-only PDF | Need OCR (tesseract) |
Save as extract_pdf_robust.sh:
#!/bin/bash
set -e
URL="$1"
OUTPUT="${2:-output.txt}"
TEMP_PDF="temp_download.pdf"
echo "Downloading from: $URL"
curl -L -A "Mozilla/5.0" -o "$TEMP_PDF" "$URL"
# Validate
SIZE=$(stat -c%s "$TEMP_PDF" 2>/dev/null || stat -f%z "$TEMP_PDF")
echo "Downloaded: $SIZE bytes"
if [ "$SIZE" -lt 1024 ]; then
echo "ERROR: File too small - checking content..."
head -200 "$TEMP_PDF"
exit 1
fi
if ! head -c 4 "$TEMP_PDF" | grep -q "%PDF"; then
echo "ERROR: Not a valid PDF file"
head -100 "$TEMP_PDF"
exit 1
fi
# Try extraction methods
python3 << 'PYTHON'
import sys
import fitz
import pdfplumber
pdf_path = "temp_download.pdf"
output_path = "output.txt"
# Method 1: PyMuPDF
try:
doc = fitz.open(pdf_path)
text = "".join(page.get_text() for page in doc)
if len(text.strip()) > 500:
with open(output_path, "w") as f:
f.write(text)
print(f"PyMuPDF: {len(text)} chars")
sys.exit(0)
except Exception as e:
print(f"PyMuPDF failed: {e}")
# Method 2: pdfplumber
try:
text = ""
with pdfplumber.open(pdf_path) as pdf:
for page in pdf.pages:
txt = page.extract_text()
if txt:
text += txt + "\n"
if len(text.strip()) > 500:
with open(output_path, "w") as f:
f.write(text)
print(f"pdfplumber: {len(text)} chars")
sys.exit(0)
except Exception as e:
print(f"pdfplumber failed: {e}")
print("All extraction methods failed")
sys.exit(1)
PYTHONcurl - For downloadingpoppler-utils (pdftotext) - Optional, fast extractionPyMuPDF (fitz) - pip install pymupdfpdfplumber - pip install pdfplumber© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in benchmarks/gdpval/skills/pdf-extraction-fallbacks of HKUDS/OpenSpace.
Open the folder on GitHubat commit 3827781
PDF Extraction Fallbacks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF Extraction Fallbacks this skillHKUDS/OpenSpace | 7.7k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Jev SEOAgriciDaniel/jev-seo | 527 | — | ~2.5k | Automated safety check: Notes | MIT | |
| Analyzing Malicious PDF With Peepdfmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~799 | Automated safety check: Pass | Apache-2.0 | |
| PDF Toolkitborghei/Claude-Skills | 881 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Edit PDFSimplePDF/simplepdf-embed | 407 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Build With SimplepdfSimplePDF/simplepdf-embed | 407 | — | ~7k | Automated safety check: Pass | MIT |
AgriciDaniel/jev-seo
Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).
mukul975/Anthropic-Cybersecurity-Skills
Perform static analysis of malicious PDF documents using peepdf, pdfid, and pdf-parser to extract embedded JavaScript, shellcode, and suspicious objects.
borghei/Claude-Skills
Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version.
SimplePDF/simplepdf-embed
Edit and fill PDF documents. An agent skill from SimplePDF/simplepdf-embed.
SimplePDF/simplepdf-embed
Integrate SimplePDF into a web application for PDF viewing, editing, filling, signing, programmatic control, AI-agent interaction, human-in-the-loop form prefilling, submissions, webhooks, or…
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
HKUDS/OpenSpace
Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.
HKUDS/OpenSpace
Handle cascading data retrieval tool failures by falling back to embedded knowledge generation
HKUDS/OpenSpace
Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.
HKUDS/OpenSpace
A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.
HKUDS/OpenSpace
Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.
HKUDS/OpenSpace
Fallback workflow for executing Python code when executecodesandbox fails repeatedly
Works with
Categories
Multi-fallback PDF/text extraction with early failure detection and sequential tool fallbacks. PDF Extraction Fallbacks is an agent skill from HKUDS/OpenSpace.
PDF Extraction Fallbacks fits situations like: tasks that involve PDF.
Run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-extraction-fallbacks in HKUDS/OpenSpace) into .claude/skills/pdf-extraction-fallbacks in your project. Claude Code loads it when a task matches its description.
Run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-extraction-fallbacks in HKUDS/OpenSpace) into .agents/skills/pdf-extraction-fallbacks in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-extraction-fallbacks, .gemini/skills/pdf-extraction-fallbacks, .github/skills/pdf-extraction-fallbacks and .opencode/skills/pdf-extraction-fallbacks in your project.
Going by SKILL.md and its folder, PDF Extraction Fallbacks needs the command-line tools its instructions call (curl, pip, pdftotext, python3 and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use curl and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
PDF Extraction Fallbacks is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with PDF Extraction Fallbacks: Jev SEO (AgriciDaniel/jev-seo, 527 stars), Analyzing Malicious PDF With Peepdf (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), PDF Toolkit (borghei/Claude-Skills, 881 stars) and Edit PDF (SimplePDF/simplepdf-embed, 407 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,749 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.
Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.