Jev SEO
AgriciDaniel/jev-seo
Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).
Multi-fallback PDF extraction with sequential approaches and early failure detection
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks-7d54a9 --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 .claude/skills/pdf-extraction-fallbacks-7d54a9 && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pdf-extraction-fallbacks-7d54a9" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 into .claude/skills/pdf-extraction-fallbacks-7d54a9/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks-7d54a9", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks-7d54a9 --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 .agents/skills/pdf-extraction-fallbacks-7d54a9 && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pdf-extraction-fallbacks-7d54a9" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 into .agents/skills/pdf-extraction-fallbacks-7d54a9/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks-7d54a9", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks-7d54a9 --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 .cursor/skills/pdf-extraction-fallbacks-7d54a9 && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pdf-extraction-fallbacks-7d54a9" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 into .cursor/skills/pdf-extraction-fallbacks-7d54a9/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks-7d54a9", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/HKUDS/OpenSpace.git --path benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks-7d54a9 --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 .gemini/skills/pdf-extraction-fallbacks-7d54a9 && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pdf-extraction-fallbacks-7d54a9" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 into .gemini/skills/pdf-extraction-fallbacks-7d54a9/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks-7d54a9", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks-7d54a9Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .github/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 .github/skills/pdf-extraction-fallbacks-7d54a9 && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pdf-extraction-fallbacks-7d54a9" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 into .github/skills/pdf-extraction-fallbacks-7d54a9/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks-7d54a9", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install HKUDS/OpenSpace pdf-extraction-fallbacks-7d54a9 --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 .opencode/skills/pdf-extraction-fallbacks-7d54a9 && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pdf-extraction-fallbacks-7d54a9" agent skill from https://github.com/HKUDS/OpenSpace/tree/main/benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 into .opencode/skills/pdf-extraction-fallbacks-7d54a9/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pdf-extraction-fallbacks-7d54a9", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pdf-extraction-fallbacks-7d54a9Multi-fallback PDF extraction with sequential approaches and early failure detection
PDF Extraction Fallbacks 7d54a9 is an agent skill from HKUDS/OpenSpace. Multi-fallback PDF extraction with sequential approaches and early failure detection
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.
It sits in Documents & Office, covering PDF. It works with JavaScript. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlpdftotextFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
PDF Extraction Fallbacks 7d54a9 loads about 2k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 271 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 271 words, ~1,956 tokens.
.claude/skills/pdf-extraction-fallbacks-7d54a9/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.This skill provides a robust workflow for extracting text from PDFs when source documents may fail to download or extract due to JavaScript protection, CORS restrictions, or encoding issues.
# Download PDF and immediately validate
curl -L -o document.pdf "URL_HERE"
# Check file size (reject if < 1KB - likely error page)
FILE_SIZE=$(stat -f%z document.pdf 2>/dev/null || stat -c%s document.pdf 2>/dev/null)
if [ "$FILE_SIZE" -lt 1024 ]; then
echo "FAIL: File too small ($FILE_SIZE bytes) - likely error page"
# Log the actual content to diagnose
head -c 500 document.pdf
exit 1
fi
# Check for HTML/error content instead of PDF
if head -c 500 document.pdf | grep -qi "<!DOCTYPE html\|<html\|error\|access denied"; then
echo "FAIL: Downloaded HTML/error page instead of PDF"
exit 1
fiTry extraction methods in order of reliability:
Fallback 1: pdftotext (poppler-utils)
if command -v pdftotext &> /dev/null; then
pdftotext -layout document.pdf output.txt 2>/dev/null
if [ -s output.txt ] && [ $(wc -c < output.txt) -gt 100 ]; then
echo "SUCCESS: pdftotext extraction"
exit 0
fi
fiFallback 2: PyMuPDF (fitz)
import fitz # PyMuPDF
def extract_with_pymupdf(pdf_path):
try:
doc = fitz.open(pdf_path)
text = ""
for page in doc:
text += page.get_text()
doc.close()
if len(text.strip()) > 100:
return text
return None
except Exception as e:
print(f"PyMuPDF failed: {e}")
return NoneFallback 3: pdfplumber
import pdfplumber
def extract_with_pdfplumber(pdf_path):
try:
text = ""
with pdfplumber.open(pdf_path) as pdf:
for page in pdf.pages:
page_text = page.extract_text()
if page_text:
text += page_text + "\n"
if len(text.strip()) > 100:
return text
return None
except Exception as e:
print(f"pdfplumber failed: {e}")
return NoneAfter any extraction, validate the output:
def validate_extraction(text, min_chars=100, min_words=20):
"""Check if extracted text is meaningful content."""
if not text:
return False, "Empty extraction"
text = text.strip()
if len(text) < min_chars:
return False, f"Too short: {len(text)} chars"
words = text.split()
if len(words) < min_words:
return False, f"Too few words: {len(words)}"
# Check for error patterns
error_patterns = [
"access denied", "permission denied", "javascript required",
"failed to load", "cannot display", "corrupted"
]
text_lower = text.lower()
for pattern in error_patterns:
if pattern in text_lower[:500]: # Check beginning
return False, f"Error pattern detected: {pattern}"
return True, "Valid extraction"import requests
import subprocess
import os
from pathlib import Path
def robust_pdf_extraction(url, output_path="extracted.txt", temp_pdf="temp.pdf"):
"""
Multi-fallback PDF extraction with validation at each step.
Returns (success, text_or_error)
"""
# Step 1: Download with validation
try:
response = requests.get(url, timeout=30, headers={
'User-Agent': 'Mozilla/5.0 (compatible; DocumentExtractor/1.0)'
})
response.raise_for_status()
except Exception as e:
return False, f"Download failed: {e}"
# Check response size
if len(response.content) < 1024:
return False, f"Downloaded content too small: {len(response.content)} bytes"
# Check for HTML error pages
if response.content[:500].lower().find(b'<html') != -1:
return False, "Downloaded HTML page instead of PDF"
# Save PDF
Path(temp_pdf).write_bytes(response.content)
# Step 2: Try extraction methods in order
extraction_methods = [
("pdftotext", extract_pdftotext),
("PyMuPDF", extract_pymupdf),
("pdfplumber", extract_pdfplumber),
]
for method_name, extract_func in extraction_methods:
try:
text = extract_func(temp_pdf)
valid, msg = validate_extraction(text)
if valid:
Path(output_path).write_text(text)
return True, text
print(f"{method_name}: {msg}")
except Exception as e:
print(f"{method_name} exception: {e}")
# Cleanup
os.remove(temp_pdf)
return False, "All extraction methods failed"
def extract_pdftotext(pdf_path):
result = subprocess.run(
["pdftotext", "-layout", pdf_path, "-"],
capture_output=True, text=True, timeout=60
)
return result.stdout if result.returncode == 0 else None
def extract_pymupdf(pdf_path):
import fitz
doc = fitz.open(pdf_path)
text = "".join(page.get_text() for page in doc)
doc.close()
return text
def extract_pdfplumber(pdf_path):
import pdfplumber
text = ""
with pdfplumber.open(pdf_path) as pdf:
for page in pdf.pages:
page_text = page.extract_text()
if page_text:
text += page_text + "\n"
return text| Check | Threshold | Action |
|---|---|---|
| File size | < 1KB | Reject - likely error page |
| Content type | HTML detected | Reject - not a PDF |
| Extracted text | < 100 chars | Try next fallback |
| Word count | < 20 words | Try next fallback |
| Error patterns | Found in first 500 chars | Reject extraction |
| Symptom | Likely Cause | Solution |
|---|---|---|
| 92-byte "PDF" | JavaScript error page | Use headless browser (Playwright/Selenium) |
| HTML content | Redirect to login/error | Check authentication requirements |
| Empty extraction | Scan-only PDF | Use OCR (pytesseract) as additional fallback |
| Garbled text | Encoding issues | Try different PDF libraries |
# For regulatory document retrieval workflows
def retrieve_regulatory_doc(doc_url, output_dir="docs"):
success, result = robust_pdf_extraction(
doc_url,
output_path=f"{output_dir}/content.txt",
temp_pdf=f"{output_dir}/temp.pdf"
)
if success:
print(f"✓ Extracted {len(result)} characters")
return result
else:
print(f"✗ Failed: {result}")
# Log URL for manual review
with open("failed_urls.log", "a") as f:
f.write(f"{doc_url}: {result}\n")
return None© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 of HKUDS/OpenSpace.
Open the folder on GitHubat commit 3827781
PDF Extraction Fallbacks 7d54a9 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| PDF Extraction Fallbacks 7d54a9 this skillHKUDS/OpenSpace | 7.8k | — | ~2k | Automated safety check: Pass | MIT | |
| Jev SEOAgriciDaniel/jev-seo | 543 | — | ~2.5k | Automated safety check: Notes | MIT | |
| Analyzing Malicious PDF With Peepdfmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~799 | Automated safety check: Pass | Apache-2.0 | |
| PDF Toolkitborghei/Claude-Skills | 891 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Edit PDFSimplePDF/simplepdf-embed | 407 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Build With SimplepdfSimplePDF/simplepdf-embed | 407 | — | ~7k | Automated safety check: Pass | MIT |
AgriciDaniel/jev-seo
Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).
mukul975/Anthropic-Cybersecurity-Skills
Perform static analysis of malicious PDF documents using peepdf, pdfid, and pdf-parser to extract embedded JavaScript, shellcode, and suspicious objects.
borghei/Claude-Skills
Audit PDF files for metadata leakage, page count, encryption, JavaScript, embedded files, and version.
SimplePDF/simplepdf-embed
Edit and fill PDF documents. An agent skill from SimplePDF/simplepdf-embed.
SimplePDF/simplepdf-embed
Integrate SimplePDF into a web application for PDF viewing, editing, filling, signing, programmatic control, AI-agent interaction, human-in-the-loop form prefilling, submissions, webhooks, or…
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
HKUDS/OpenSpace
Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.
HKUDS/OpenSpace
Handle cascading data retrieval tool failures by falling back to embedded knowledge generation
HKUDS/OpenSpace
Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.
HKUDS/OpenSpace
A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.
HKUDS/OpenSpace
Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.
HKUDS/OpenSpace
Fallback workflow for executing Python code when executecodesandbox fails repeatedly
Works with
Categories
Multi-fallback PDF extraction with sequential approaches and early failure detection. PDF Extraction Fallbacks 7d54a9 is an agent skill from HKUDS/OpenSpace.
PDF Extraction Fallbacks 7d54a9 fits situations like: tasks that involve PDF.
Run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 in HKUDS/OpenSpace) into .claude/skills/pdf-extraction-fallbacks-7d54a9 in your project. Claude Code loads it when a task matches its description.
Run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-extraction-fallbacks-7d54a9 in HKUDS/OpenSpace) into .agents/skills/pdf-extraction-fallbacks-7d54a9 in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill pdf-extraction-fallbacks-7d54a9 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-extraction-fallbacks-7d54a9, .gemini/skills/pdf-extraction-fallbacks-7d54a9, .github/skills/pdf-extraction-fallbacks-7d54a9 and .opencode/skills/pdf-extraction-fallbacks-7d54a9 in your project.
Going by SKILL.md and its folder, PDF Extraction Fallbacks 7d54a9 needs the command-line tools its instructions call (curl and pdftotext). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
PDF Extraction Fallbacks 7d54a9 is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with PDF Extraction Fallbacks 7d54a9: Jev SEO (AgriciDaniel/jev-seo, 543 stars), Analyzing Malicious PDF With Peepdf (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), PDF Toolkit (borghei/Claude-Skills, 891 stars) and Edit PDF (SimplePDF/simplepdf-embed, 407 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,754 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.
Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.