Agent skill

PDF Extract Create Workflow

by HKUDS in HKUDS/OpenSpace

Complete PDF lifecycle: download, extract, and generate structured documents with reportlab

MITAuto-check passedDocuments & Office

Install PDF Extract Create Workflow

skills CLI
$ npx skills add HKUDS/OpenSpace --skill pdf-extract-create-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace pdf-extract-create-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-29b2f4 .claude/skills/pdf-extract-create-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-extract-create-workflow
GitHub stars
7.7k
Token cost
~4.4k tokens
SKILL.md length
761 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Complete PDF lifecycle: download, extract, and generate structured documents with reportlab

  • Works in 5 steps: Download PDF with Browser User-Agent → Verify File Type Before Parsing → Primary Extraction with pdftotext → …
  • Tasks that involve PDF
  • SKILL.md covers Overview, Entry Point: Determine Your…, Mode A: Web URL Download and Mode B: Local File Processing…, plus 5 more sections
  • Calls python3, curl and pip

What it does

PDF Extract Create Workflow is an agent skill from HKUDS/OpenSpace. Complete PDF lifecycle: download, extract, and generate structured documents with reportlab

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering PDF. It works with pypdf. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-extract-create-workflow”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Download PDF with Browser User-Agent
  2. Verify File Type Before Parsing
  3. Primary Extraction with pdftotext
  4. Fallback to PyMuPDF (fitz)
  5. Graceful Degradation to Domain Knowledge

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • curl
    • pip
    • apt-get
    • pdftotext
    • brew
    • yum

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Extract Create Workflow loads about 4.4k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 761 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 761 words, ~4,405 tokens.

Download SKILL.mdSave it as .claude/skills/pdf-extract-create-workflow/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
pdf-extract-create-workflow
description
Complete PDF lifecycle: download, extract, and generate structured documents with reportlab

PDF Extract and Create Workflow

This skill provides a complete PDF lifecycle workflow for acquiring PDF documents from web sources or local files, extracting their text content, AND generating new structured PDFs from processed data—with multiple fallback mechanisms throughout.

Overview

When working with PDFs, you may need to:

  1. Download PDFs from web sources (with anti-bot protection)
  2. Extract text content from PDFs (with fallback strategies)
  3. Generate new PDFs from processed data (with professional formatting)

This workflow ensures maximum success rate through progressive fallback strategies for extraction and templated approaches for generation.

Entry Point: Determine Your Starting Point

Before beginning, identify your scenario:

ScenarioStart HereSkip
PDF already on local diskStep 2 (Verify File Type)Step 1 (Download)
PDF at a web URLStep 1 (Download)None
Need to CREATE a PDF from dataMode C (Generate)Modes A & B
Need to extract AND createMode A/B → Mode CNone

Mode A: Web URL Download

Step 1: Download PDF with Browser User-Agent

Many PDF hosting sites use JavaScript-based redirects or block automated requests. Use curl with a realistic browser user-agent:

bash
curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" -o output.pdf "URL_HERE"

Key flags:

  • -L: Follow redirects
  • -A: Set user-agent header to mimic a real browser
  • -o: Specify output filename

Additional headers for difficult sites:

bash
curl -L \
  -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" \
  -H "Accept: application/pdf,*/*" \
  -H "Accept-Language: en-US,en;q=0.9" \
  -H "Connection: keep-alive" \
  -o output.pdf "URL_HERE"

Mode B: Local File Processing & Extraction

If you already have the PDF file locally, skip Step 1 and begin here:

Step 2: Verify File Type Before Parsing

Always validate the downloaded file is actually a PDF before attempting extraction:

bash
file output.pdf

Expected output should contain "PDF document". If not:

  • The URL may have redirected to an HTML error page
  • The file may be corrupted
  • Access may be blocked
Step 3: Primary Extraction with pdftotext

First attempt extraction using the standard pdftotext utility (part of poppler-utils):

bash
pdftotext output.pdf output.txt

If pdftotext is not available, install it:

bash
# Debian/Ubuntu
apt-get update && apt-get install -y poppler-utils

# macOS
brew install poppler

# RHEL/CentOS
yum install -y poppler-utils
Step 4: Fallback to PyMuPDF (fitz)

If pdftotext fails or produces poor results, use Python's PyMuPDF library:

python
import fitz  # PyMuPDF

doc = fitz.open("output.pdf")
text = ""
for page in doc:
    text += page.get_text()
doc.close()

with open("output.txt", "w") as f:
    f.write(text)

Install if needed:

bash
pip install pymupdf
Step 5: Graceful Degradation to Domain Knowledge

If the PDF cannot be accessed or extracted after all attempts:

  1. Document the failure mode (network issue, corrupted file, access denied, etc.)
  2. Extract any partial content that was successfully retrieved
  3. Supplement missing content from established domain knowledge
  4. Clearly mark which portions are from source vs. generated from knowledge
  5. Provide citations for any claimed requirements or specifications

Example degradation note:

NOTE: Source document [URL] was inaccessible due to [reason]. 
Content below combines partial extraction with established domain knowledge 
for [topic]. Verify against official sources when available.

Mode C: PDF Generation with ReportLab

After extracting or processing data, generate professional PDFs using Python's reportlab library.

Installation
bash
pip install reportlab
Step C1: Basic Document Structure

Create a multi-page PDF with title page, sections, and proper formatting:

python
from reportlab.lib.pagesizes import letter
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreak
from reportlab.lib.styles import getSampleStyleSheet, ParagraphStyle
from reportlab.lib.units import inch
from reportlab.lib.enums import TA_CENTER, TA_LEFT

def create_structured_pdf(output_path, title, sections):
    """
    Create a structured PDF with title page and sections.
    
    Args:
        output_path: Path for output PDF
        title: Document title
        sections: List of dicts with 'heading' and 'content' keys
    """
    doc = SimpleDocTemplate(
        output_path,
        pagesize=letter,
        rightMargin=72,
        leftMargin=72,
        topMargin=72,
        bottomMargin=72
    )
    
    styles = getSampleStyleSheet()
    story = []
    
    # Title Page
    title_style = ParagraphStyle(
        'CustomTitle',
        parent=styles['Heading1'],
        fontSize=24,
        alignment=TA_CENTER,
        spaceAfter=30
    )
    story.append(Paragraph(title, title_style))
    story.append(Spacer(1, 2*inch))
    story.append(PageBreak())
    
    # Content Sections
    heading_style = ParagraphStyle(
        'CustomHeading',
        parent=styles['Heading2'],
        fontSize=16,
        spaceBefore=12,
        spaceAfter=6
    )
    body_style = ParagraphStyle(
        'CustomBody',
        parent=styles['Normal'],
        fontSize=11,
        leading=14,
        spaceAfter=12
    )
    
    for section in sections:
        story.append(Paragraph(section['heading'], heading_style))
        # Handle long text by splitting into paragraphs
        for paragraph in section['content'].split('\n\n'):
            if paragraph.strip():
                story.append(Paragraph(paragraph, body_style))
        story.append(Spacer(1, 0.2*inch))
    
    doc.build(story)
    print(f"PDF created: {output_path}")
Step C2: Adding Tables

For structured data, include tables with proper formatting:

python
from reportlab.platypus import Table, TableStyle
from reportlab.lib import colors

def create_table(data, col_widths=None):
    """
    Create a formatted table for PDF.
    
    Args:
        data: List of lists (rows x columns)
        col_widths: Optional list of column widths
    """
    table = Table(data, colWidths=col_widths)
    table.setStyle(TableStyle([
        # Header row
        ('BACKGROUND', (0, 0), (-1, 0), colors.grey),
        ('TEXTCOLOR', (0, 0), (-1, 0), colors.whitesmoke),
        ('ALIGN', (0, 0), (-1, -1), 'LEFT'),
        ('FONTNAME', (0, 0), (-1, 0), 'Helvetica-Bold'),
        ('FONTSIZE', (0, 0), (-1, 0), 12),
        ('BOTTOMPADDING', (0, 0), (-1, 0), 12),
        # Data rows
        ('BACKGROUND', (0, 1), (-1, -1), colors.beige),
        ('TEXTCOLOR', (0, 1), (-1, -1), colors.black),
        ('FONTNAME', (0, 1), (-1, -1), 'Helvetica'),
        ('FONTSIZE', (0, 1), (-1, -1), 10),
        # Grid
        ('GRID', (0, 0), (-1, -1), 1, colors.black),
        ('ROWBACKGROUNDS', (0, 1), (-1, -1), [colors.white, colors.lightgrey]),
    ]))
    return table
Step C3: Multi-Section Document Template

Complete example creating an organized document with multiple sections:

python
from reportlab.lib.pagesizes import letter
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreak, Table, TableStyle
from reportlab.lib.styles import getSampleStyleSheet, ParagraphStyle
from reportlab.lib.enums import TA_CENTER
from reportlab.lib import colors

def create_report_pdf(output_path, title, subtitle, sections, table_data=None):
    """
    Create a complete report PDF with title, sections, and optional tables.
    
    Args:
        output_path: Output PDF path
        title: Main title
        subtitle: Subtitle or date
        sections: List of {'heading': str, 'content': str} dicts
        table_data: Optional list of lists for tables
    """
    doc = SimpleDocTemplate(
        output_path,
        pagesize=letter,
        rightMargin=50,
        leftMargin=50,
        topMargin=50,
        bottomMargin=50
    )
    
    styles = getSampleStyleSheet()
    story = []
    
    # Title Page
    title_style = ParagraphStyle(
        'Title',
        parent=styles['Heading1'],
        fontSize=28,
        alignment=TA_CENTER,
        spaceAfter=20,
        fontName='Helvetica-Bold'
    )
    subtitle_style = ParagraphStyle(
        'Subtitle',
        parent=styles['Normal'],
        fontSize=14,
        alignment=TA_CENTER,
        spaceAfter=50,
        textColor=colors.darkgrey
    )
    
    story.append(Paragraph(title, title_style))
    story.append(Paragraph(subtitle, subtitle_style))
    story.append(PageBreak())
    
    # Content
    heading_style = ParagraphStyle(
        'SectionHeading',
        parent=styles['Heading2'],
        fontSize=16,
        spaceBefore=20,
        spaceAfter=10,
        fontName='Helvetica-Bold',
        textColor=colors.darkblue
    )
    body_style = ParagraphStyle(
        'Body',
        parent=styles['Normal'],
        fontSize=11,
        leading=15,
        spaceAfter=12
    )
    
    for i, section in enumerate(sections):
        story.append(Paragraph(section['heading'], heading_style))
        
        # Split content into paragraphs
        for para in section['content'].split('\n\n'):
            if para.strip():
                # Handle very long paragraphs
                story.append(Paragraph(para, body_style))
        
        # Add table after specific section if provided
        if table_data and i == 0:
            story.append(Spacer(1, 0.3*inch))
            table = Table(table_data)
            table.setStyle(TableStyle([
                ('BACKGROUND', (0, 0), (-1, 0), colors.darkblue),
                ('TEXTCOLOR', (0, 0), (-1, 0), colors.white),
                ('ALIGN', (0, 0), (-1, -1), 'LEFT'),
                ('FONTNAME', (0, 0), (-1, 0), 'Helvetica-Bold'),
                ('GRID', (0, 0), (-1, -1), 0.5, colors.grey),
                ('ROWBACKGROUNDS', (0, 1), (-1, -1), [colors.white, colors.lightgrey]),
            ]))
            story.append(table)
            story.append(Spacer(1, 0.3*inch))
        
        if i < len(sections) - 1:
            story.append(PageBreak())
    
    doc.build(story)
    return output_path

# Example usage
if __name__ == "__main__":
    sections = [
        {
            'heading': 'Section 1: Overview',
            'content': 'This is the first section content...\n\nAdditional paragraph here.'
        },
        {
            'heading': 'Section 2: Details',
            'content': 'Detailed information goes here...'
        }
    ]
    
    table_data = [
        ['Header 1', 'Header 2', 'Header 3'],
        ['Row 1 Col 1', 'Row 1 Col 2', 'Row 1 Col 3'],
        ['Row 2 Col 1', 'Row 2 Col 2', 'Row 2 Col 3'],
    ]
    
    create_report_pdf(
        "output_report.pdf",
        "Report Title",
        "Generated: 2024",
        sections,
        table_data
    )
Step C4: Error Handling for PDF Generation

PDF generation can fail in multiple ways. Handle gracefully:

python
def safe_pdf_generation(output_path, title, sections, max_retries=3):
    """
    Generate PDF with retry logic and error handling.
    """
    import traceback
    from reportlab.lib.utils import ImageReader
    
    for attempt in range(max_retries):
        try:
            create_report_pdf(output_path, title, sections)
            # Verify file was created
            import os
            if os.path.exists(output_path) and os.path.getsize(output_path) > 0:
                print(f"✓ PDF generated successfully: {output_path}")
                return True
            else:
                raise Exception("PDF file empty or not created")
        except Exception as e:
            print(f"Attempt {attempt + 1}/{max_retries} failed: {e}")
            if attempt < max_retries - 1:
                import time
                time.sleep(1)  # Brief delay before retry
            else:
                print(f"PDF generation failed after {max_retries} attempts")
                print(traceback.format_exc())
                # Fallback: create minimal text file
                with open(output_path.replace('.pdf', '.txt'), 'w') as f:
                    f.write(f"Title: {title}\n\n")
                    for section in sections:
                        f.write(f"{section['heading']}\n{section['content']}\n\n")
                return False
Show full SKILL.md (297 more words)Show less
Step C5: Best Practices for PDF Generation
  1. Page Breaks: Insert PageBreak() between major sections for readability
  2. Consistent Styling: Define ParagraphStyle objects once and reuse
  3. Text Wrapping: ReportLab handles wrapping automatically; split long content with \n\n
  4. Margins: Use at least 50-72 point margins for standard letter size
  5. Font Selection: Stick to Helvetica, Times-Roman, or Courier for compatibility
  6. File Verification: Always check the PDF was created and has content > 0 bytes
  7. Error Recovery: Have a fallback (e.g., .txt output) if PDF generation fails
  8. Memory Management: For very large documents, build in chunks or use separate files

Complete Workflow Script (Handles Download, Extract, and Generate)

bash
#!/bin/bash
# pdf-lifecycle-workflow.sh
# Handles URL downloads, local files, and PDF generation

INPUT="$1"
MODE="${2:-extract}"  # extract, generate, or both
OUTPUT_PDF="downloaded.pdf"
OUTPUT_TXT="extracted.txt"
OUTPUT_REPORT="generated_report.pdf"

if [[ "$MODE" == "generate" ]]; then
    echo "Mode: PDF Generation"
    python3 generate_pdf.py
    exit $?
fi

if [[ "$INPUT" =~ ^https?:// ]]; then
    # Mode A: URL download
    PDF_URL="$INPUT"
    echo "Downloading PDF from URL..."
    curl -L -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" -o "$OUTPUT_PDF" "$PDF_URL"
else
    # Mode B: Local file
    if [ ! -f "$INPUT" ]; then
        echo "ERROR: Local file not found: $INPUT"
        exit 1
    fi
    OUTPUT_PDF="$INPUT"
    echo "Using local file: $INPUT"
fi

# Step 2: Verify file type
echo "Verifying file type..."
if ! file "$OUTPUT_PDF" | grep -q "PDF document"; then
    echo "WARNING: File is not a valid PDF"
    echo "Attempting fallback extraction anyway..."
fi

# Step 3: Try pdftotext
echo "Attempting pdftotext extraction..."
if command -v pdftotext &> /dev/null; then
    if pdftotext "$OUTPUT_PDF" "$OUTPUT_TXT" 2>/dev/null; then
        echo "Extraction successful with pdftotext"
        if [[ "$MODE" == "both" ]]; then
            echo "Proceeding to PDF generation..."
            python3 generate_pdf.py
        fi
        exit 0
    fi
fi

# Step 4: Fallback to PyMuPDF
echo "Falling back to PyMuPDF..."
python3 << 'PYTHON_SCRIPT'
import fitz
import sys

try:
    doc = fitz.open("downloaded.pdf")
    text = ""
    for page in doc:
        text += page.get_text()
    doc.close()
    with open("extracted.txt", "w") as f:
        f.write(text)
    print("Extraction successful with PyMuPDF")
except Exception as e:
    print(f"PyMuPDF failed: {e}")
    sys.exit(1)
PYTHON_SCRIPT

# Step 5: Handle complete failure
if [ $? -ne 0 ]; then
    echo "ERROR: All extraction methods failed."
    echo "ACTION: Generate content from domain knowledge and clearly mark source limitations."
fi

if [[ "$MODE" == "both" ]]; then
    echo "Proceeding to PDF generation with extracted/fallback content..."
    python3 generate_pdf.py
fi

Common Failure Modes & Solutions

SymptomCauseSolution
HTML content in PDFURL redirected to error pageCheck HTTP status, try alternate URL
Empty extractionPassword-protected or scanned PDFTry OCR tools or request accessible version
Garbled textEncoding issuesTry PyMuPDF with different extraction mode
Curl blockedAnti-bot measuresAdd more headers, use delay between requests
PDF generation failsMissing fonts or memoryUse standard fonts, build in chunks
ReportLab errorsVersion incompatibilityUse pip install --upgrade reportlab
Unknown shell_agent errorTimeout on complex operationsUse direct Python execution instead

When to Use This Skill

ModeUse Case
Mode A (URL download)Downloading regulatory documents from government websites
Mode B (Local file)Processing PDFs already saved to disk
Mode C (Generate)Creating reports from extracted/processed data
Both (extract + generate)Full pipeline: acquire → process → report
Specific Scenarios
  • Extracting content from technical manuals or handbooks
  • Processing PDFs in automated pipelines where reliability matters
  • Creating structured reports from multiple data sources
  • Generating documentation with consistent formatting
  • Any situation where PDF access may be unreliable or restricted
  • Producing professional PDFs from text data with tables and sections

Quick Reference: Mode Selection

Need to get PDF from web?          → Mode A (Download)
Have PDF file already?             → Mode B (Extract)
Need to CREATE a PDF?              → Mode C (Generate)
Need full pipeline?                → Mode A/B → Mode C
Extraction failed?                 → Step 5 (Domain knowledge fallback)
Generation failed?                 → Step C4 (Error handling + txt fallback)

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-29b2f4 of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

PDF Extract Create Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Extract Create Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Extract Create Workflow this skillHKUDS/OpenSpace7.7k—~4.4kAutomated safety check: PassMIT
PDFnuoyimanaituling/manus-x830—~985Automated safety check: PassNone
PDFeinverne/dotfiles12147 repos~1.8kAutomated safety check: PassProprietary
Reportlabjimmc414/Kosmos5941 repos~4.2kAutomated safety check: PassNone
PDFguyi-a/pi-ling106—~3.3kAutomated safety check: PassMIT
PDF ReadingWide-Moat/open-computer-use1261 repos~2.7kAutomated safety check: PassProprietary

Similar skills

  • PDF

    nuoyimanaituling/manus-x

    Process PDF files - extract text, read content, create PDFs, merge or split documents.

    830 GitHub stars~985 tokensUpdated 8 mo ago
    Documents & OfficeAuto-check passed
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 47 repos~1.8k tokens
    Documents & OfficeAuto-check passed
  • Reportlab

    jimmc414/Kosmos

    PDF generation toolkit. An agent skill from jimmc414/Kosmos.

    594 GitHub starsUsed in 1 repo~4.2k tokens
    Documents & OfficeAuto-check passed
  • PDF

    guyi-a/pi-ling

    PDF 相关的所有操作:从零生成(reportlab / pypdf)、格式转化(md/html → PDF)、修改(合并 / 拆分 / 旋转 / 加水印 / 提图片 / 元数据)、读内容(pdfplumber / extractdocumenttext)、OCR 扫描件、加密解密。触发场景:用户说"生成 PDF" / "做份 PDF 简历" / "合并这几份 PDF" / "给 PDF…

    106 GitHub stars~3.3k tokensUpdated 22 days ago
    Documents & OfficeAuto-check passed
  • PDF Reading

    Wide-Moat/open-computer-use

    A skill your agent uses when you need to read, inspect, or extract content from PDF files — especially when file content is NOT in your context and you need to read it from disk.

    126 GitHub starsUsed in 1 repo~2.7k tokens
    Documents & OfficeAuto-check passed
  • PDF

    LeastBit/Claude_skills_zh-CN

    全面的 PDF 操作工具包,用于提取文本和表格、创建新 PDF、合并/拆分文档以及处理表单。当 Claude 需要填写 PDF 表单或以编程方式大规模处理、生成或分析 PDF 文档时使用。

    588 GitHub stars~1.5k tokensUpdated 8 mo ago
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about PDF Extract Create Workflow

What does PDF Extract Create Workflow do?

Complete PDF lifecycle: download, extract, and generate structured documents with reportlab. PDF Extract Create Workflow is an agent skill from HKUDS/OpenSpace.

When should I use PDF Extract Create Workflow?

PDF Extract Create Workflow fits situations like: tasks that involve PDF.

How do I install PDF Extract Create Workflow in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extract-create-workflow -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-29b2f4 in HKUDS/OpenSpace) into .claude/skills/pdf-extract-create-workflow in your project. Claude Code loads it when a task matches its description.

How do I install PDF Extract Create Workflow in Codex?

Run `npx skills add HKUDS/OpenSpace --skill pdf-extract-create-workflow -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-29b2f4 in HKUDS/OpenSpace) into .agents/skills/pdf-extract-create-workflow in your project. Codex loads it when a task matches its description.

Can I use PDF Extract Create Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill pdf-extract-create-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-extract-create-workflow, .gemini/skills/pdf-extract-create-workflow, .github/skills/pdf-extract-create-workflow and .opencode/skills/pdf-extract-create-workflow in your project.

What does PDF Extract Create Workflow need to run?

Going by SKILL.md and its folder, PDF Extract Create Workflow needs the command-line tools its instructions call (python3, curl, pip, apt-get, pdftotext and brew). Our summary lists: Python 3.

Does PDF Extract Create Workflow access the network?

SKILL.md contains no URLs. Its commands use curl and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is PDF Extract Create Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does PDF Extract Create Workflow use?

PDF Extract Create Workflow is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PDF Extract Create Workflow use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Extract Create Workflow?

Skills that share tags, products or a category with PDF Extract Create Workflow: PDF (nuoyimanaituling/manus-x, 830 stars), PDF (einverne/dotfiles, 121 stars), Reportlab (jimmc414/Kosmos, 594 stars) and PDF (guyi-a/pi-ling, 106 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Extract Create Workflow?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,749 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.