Agent skill

Document Gen Resilient Workflow

by HKUDS in HKUDS/OpenSpace

Multi-engine document generation with cascading PDF fallbacks and robust Unicode handling

MITAuto-check passedDocuments & Office

Install Document Gen Resilient Workflow

skills CLI
$ npx skills add HKUDS/OpenSpace --skill document-gen-resilient-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace document-gen-resilient-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/document-gen-fallback-enhanced-enhanced-2794b4 .claude/skills/document-gen-resilient-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
document-gen-resilient-workflow
GitHub stars
7.8k
Token cost
~3k tokens
SKILL.md length
679 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Multi-engine document generation with cascading PDF fallbacks and robust Unicode handling

  • Works in 4 steps: Create Source Content with write_file → Sanitize Unicode with Python Script → Convert to Target Formats with Cascading… → …
  • Tasks that involve PDF
  • SKILL.md covers When to Use, Core Technique, ⚠️ Unicode & PDF Engine Guide and Step-by-Step Workflow, plus 6 more sections
  • Calls apt-get, pip and pdftotext

What it does

Document Gen Resilient Workflow is an agent skill from HKUDS/OpenSpace. Multi-engine document generation with cascading PDF fallbacks and robust Unicode handling

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering PDF. It works with LaTeX and Python. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/document-gen-resilient-workflow”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Create Source Content with write_file
  2. Sanitize Unicode with Python Script
  3. Convert to Target Formats with Cascading Fallbacks
  4. Verify Outputs with Multiple Checks

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • apt-get
    • pip
    • pdftotext
    • pandoc

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Document Gen Resilient Workflow loads about 3k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 679 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 679 words, ~2,963 tokens.

Download SKILL.mdSave it as .claude/skills/document-gen-resilient-workflow/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
document-gen-resilient-workflow
description
Multi-engine document generation with cascading PDF fallbacks and robust Unicode handling

Resilient Document Generation Workflow (Multi-Engine Fallback)

When to Use

Use this skill when document generation tasks fail or when shell_agent returns unknown errors, especially for:

  • Generating documents in multiple formats (.docx, .pdf, .html)
  • PDF generation fails due to missing LaTeX, encoding issues, or tool errors
  • Documents contain special characters, symbols, or non-ASCII text
  • You need maximum reliability with automatic fallback options

Core Technique

Split document generation into discrete, observable steps with cascading fallbacks for PDF generation:

  1. Content creation → Use write_file to create source Markdown
  2. Unicode sanitization → Use Python script for reliable character replacement
  3. Format conversion → Try multiple PDF engines in sequence until one succeeds
  4. Verification → Check file exists, has non-zero size, and is valid

⚠️ Unicode & PDF Engine Guide

Different PDF engines have different Unicode support:

EngineUnicode SupportBest ForFallback Position
pdflatexLimited (ASCII-focused)Simple documents1st (fastest)
xelatexFull UnicodeDocuments with non-ASCII2nd
wkhtmltopdfGood UnicodeWeb-style documents3rd
reportlab (Python)Full controlProgrammatic PDFs4th
fpdf2 (Python)Full controlSimple text PDFs5th (last resort)
Character Replacement Table
CharacterIssueSafe Replacement
— (em dash)LaTeX incompatibility--
– (en dash)LaTeX incompatibility-
" " (curly quotes)Encoding errors" " (straight)
' ' (curly apostrophe)Encoding errors' (straight)
… (ellipsis)May not render...
→ ← ↑ ↓LaTeX incompatibility-> <- ^ v
✓ ✗May not render[x] [ ]
★ ●May not render* -
© ® ™May require packages(c) (r) (tm)
é ñ üFont-dependentKeep for xelatex, replace for pdflatex

Step-by-Step Workflow

Step 1: Create Source Content with write_file

Write your document content as Markdown to a source file:

write_file
path: /tmp/document_source.md
content: |
  # Document Title
  
  ## Section 1
  Content here...
  
  ## Section 2
  More content...
Step 2: Sanitize Unicode with Python Script

Create a reusable Python sanitizer for reliable character replacement:

write_file
path: /tmp/sanitize_unicode.py
content: |
  #!/usr/bin/env python3
  import sys
  import re
  
  if len(sys.argv) < 2:
      print("Usage: sanitize_unicode.py <input.md> [output.md]")
      sys.exit(1)
  
  input_file = sys.argv[1]
  output_file = sys.argv[2] if len(sys.argv) > 2 else input_file.replace('.md', '_sanitized.md')
  
  replacements = {
      '—': '--',    # em dash
      '–': '-',     # en dash
      '"': '"',     # left curly quote
      '"': '"',     # right curly quote
      "'": "'",     # left curly apostrophe
      "'": "'",     # right curly apostrophe
      '…': '...',   # ellipsis
      '→': '->',    # right arrow
      '←': '<-',    # left arrow
      '↑': '^',     # up arrow
      '↓': 'v',     # down arrow
      '✓': '[x]',   # checkmark
      '✗': '[ ]',   # cross
      '★': '*',     # star
      '●': '-',     # bullet
      '©': '(c)',   # copyright
      '®': '(r)',   # registered
      '™': '(tm)',  # trademark
  }
  
  with open(input_file, 'r', encoding='utf-8') as f:
      content = f.read()
  
  for old, new in replacements.items():
      content = content.replace(old, new)
  
  with open(output_file, 'w', encoding='utf-8') as f:
      f.write(content)
  
  print(f"Sanitized: {input_file} -> {output_file}")

Apply sanitization for PDF:

run_shell
command: python3 /tmp/sanitize_unicode.py /tmp/document_source.md /tmp/document_source_sanitized.md

Note: Keep original for DOCX/HTML (these formats handle Unicode well).

Step 3: Convert to Target Formats with Cascading Fallbacks
For DOCX (from original, no sanitization needed):
run_shell
command: pandoc /tmp/document_source.md -o output.docx
For PDF (try multiple engines in sequence):

Attempt 1: pdflatex (fastest, limited Unicode)

run_shell
command: pandoc /tmp/document_source_sanitized.md -o output.pdf --pdf-engine=pdflatex

If pdflatex fails, Attempt 2: xelatex (full Unicode)

run_shell
command: pandoc /tmp/document_source.md -o output.pdf --pdf-engine=xelatex

If xelatex fails, Attempt 3: wkhtmltopdf (web-based)

run_shell
command: pandoc /tmp/document_source.md -o output.pdf --pdf-engine=wkhtmltopdf

If all pandoc engines fail, Attempt 4: Python reportlab

write_file
path: /tmp/generate_pdf_reportlab.py
content: |
  #!/usr/bin/env python3
  from reportlab.lib.pagesizes import letter
  from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer
  from reportlab.lib.styles import getSampleStyleSheet
  import markdown
  
  input_md = '/tmp/document_source.md'
  output_pdf = 'output.pdf'
  
  with open(input_md, 'r', encoding='utf-8') as f:
      md_content = f.read()
  
  html_content = markdown.markdown(md_content)
  
  doc = SimpleDocTemplate(output_pdf, pagesize=letter)
  styles = getSampleStyleSheet()
  story = []
  
  # Simple HTML to flowables (basic implementation)
  for line in html_content.split('\n'):
      if line.strip():
          story.append(Paragraph(line, styles['Normal']))
          story.append(Spacer(1, 6))
  
  doc.build(story)
  print(f"PDF created: {output_pdf}")
run_shell
command: python3 /tmp/generate_pdf_reportlab.py

If reportlab fails, Attempt 5: Python fpdf2 (simpler)

write_file
path: /tmp/generate_pdf_fpdf.py
content: |
  #!/usr/bin/env python3
  from fpdf import FPDF
  
  input_md = '/tmp/document_source.md'
  output_pdf = 'output.pdf'
  
  pdf = FPDF()
  pdf.add_page()
  pdf.set_font('Helvetica', '', 12)
  
  with open(input_md, 'r', encoding='utf-8') as f:
      for line in f:
          # Simple line-by-line, escape special chars
          safe_line = line.encode('latin-1', 'replace').decode('latin-1')
          pdf.cell(0, 10, safe_line[:180], ln=True)
  
  pdf.output(output_pdf)
  print(f"PDF created: {output_pdf}")
run_shell
command: python3 /tmp/generate_pdf_fpdf.py
For HTML (from original):
run_shell
command: pandoc /tmp/document_source.md -o output.html
Step 4: Verify Outputs with Multiple Checks

Check 1: Files exist and have size

run_shell
command: ls -lh output.docx output.pdf output.html 2>/dev/null && echo "FILES_OK" || echo "FILES_MISSING"

Check 2: Validate PDF is not corrupted

run_shell
command: python3 -c "import fitz; doc=fitz.open('output.pdf'); print(f'PDF_VALID: {doc.page_count} pages')" 2>/dev/null || echo "PDF_CHECK_SKIPPED"

Check 3: Validate DOCX

run_shell
command: python3 -c "from docx import Document; d=Document('output.docx'); print(f'DOCX_VALID: {len(d.paragraphs)} paragraphs')" 2>/dev/null || echo "DOCX_CHECK_SKIPPED"

Check 4: Read back content for manual verification

read_file
filetype: md
file_path: output.html

Complete Example

markdown
# Generate Negotiation Strategy Document (Resilient Workflow)

## Step 1: Write Markdown source
write_file
path: /tmp/negotiation_strategy.md
content: |
  # Negotiation Strategy
  
  ## Executive Summary
  Content with original unicode characters...
  
  ## Resolution Path
  More content...

## Step 2: Sanitize for PDF
write_file
path: /tmp/sanitize_unicode.py
content: |
  [Python sanitizer script from Step 2 above]

run_shell
command: python3 /tmp/sanitize_unicode.py /tmp/negotiation_strategy.md /tmp/negotiation_strategy_sanitized.md

## Step 3: Convert to DOCX (original, Unicode-safe format)
run_shell
command: pandoc /tmp/negotiation_strategy.md -o negotiation_strategy.docx

## Step 4: Convert to PDF (try pdflatex first)
run_shell
command: pandoc /tmp/negotiation_strategy_sanitized.md -o negotiation_strategy.pdf --pdf-engine=pdflatex

## Step 4b: If pdflatex failed, try xelatex
run_shell
command: pandoc /tmp/negotiation_strategy.md -o negotiation_strategy.pdf --pdf-engine=xelatex

## Step 4c: If xelatex failed, try wkhtmltopdf
run_shell
command: pandoc /tmp/negotiation_strategy.md -o negotiation_strategy.pdf --pdf-engine=wkhtmltopdf

## Step 5: Convert to HTML (original)
run_shell
command: pandoc /tmp/negotiation_strategy.md -o negotiation_strategy.html

## Step 6: Verify all outputs
run_shell
command: ls -lh negotiation_strategy.* && echo "ALL_FILES_CREATED"

run_shell
command: python3 -c "import fitz; d=fitz.open('negotiation_strategy.pdf'); print(f'PDF: {d.page_count} pages')"
Show full SKILL.md (302 more words)Show less

Advantages Over shell_agent

Aspectshell_agentResilient Manual Workflow
Error visibilityOpaque, may retry silentlyEach step shows explicit output
PDF fallbackMay give up after first failureCascading engine attempts
DebuggingHard to isolateClear which engine/step failed
Unicode controlAgent-dependentYou control sanitization
RecoveryAutomatic but may loopManual intervention at known points
Tool requirementsAssumes pandoc worksMultiple engine options

Troubleshooting by Error Type

"LaTeX not found" or "pdflatex: command not found"
  • Solution: Try --pdf-engine=xelatex or --pdf-engine=wkhtmltopdf
  • Install LaTeX: apt-get install texlive-latex-recommended texlive-fonts-recommended
  • Or fallback to Python: Use reportlab/fpdf2 approach
"Encoding error" or "UnicodeDecodeError"
  • Solution: Use sanitized markdown file (Step 2)
  • Or: Add -f markdown+utf8 to pandoc command
  • Or: Try xelatex engine (better Unicode support)
"wkhtmltopdf not found"
  • Solution: Install via apt-get install wkhtmltopdf or fallback to Python
"ModuleNotFoundError: No module named 'reportlab'"
  • Solution: pip install reportlab or fallback to fpdf2
"Invalid PDF" or corrupted output
  • Diagnosis: Run pdftotext output.pdf - to check if text extracts
  • Solution: Try different PDF engine from fallback chain
DOCX formatting issues
  • Solution: Add --reference-doc=template.docx for custom styles
  • Or: Fix markdown structure in source file
All pandoc commands fail
  • Check: pandoc --version to verify installation
  • Fallback: Use Python-based PDF generation (reportlab/fpdf2)
  • For DOCX: Try python-docx library directly

Before starting, verify required tools are available:

run_shell
command: pandoc --version && echo "PANDOC_OK" || echo "PANDOC_MISSING"
run_shell
command: pdflatex --version && echo "PDFLATEX_OK" || echo "PDFLATEX_MISSING"
run_shell
command: python3 -c "import reportlab" && echo "REPORTLAB_OK" || echo "REPORTLAB_MISSING"

This helps you know which fallback path will be needed.

  • document-gen-unicode-safe: Parent skill with basic Unicode guidance
  • document-gen-fallback: Original fallback without Unicode or multi-engine support
  • Use this skill when you need maximum reliability with automatic fallbacks

When to Return to shell_agent

After manually completing this workflow successfully once for a document type, you can attempt shell_agent for similar tasks with reduced risk (you now know the fallback path). However, for critical documents or those with heavy Unicode content, consider always using this resilient manual workflow.

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/document-gen-fallback-enhanced-enhanced-2794b4 of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Document Gen Resilient Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Document Gen Resilient Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Document Gen Resilient Workflow this skillHKUDS/OpenSpace7.8k—~3kAutomated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Tikzscunning1975/MixtapeTools473—~1.7kAutomated safety check: PassNone
MineruNebutra/MinerU-Skill122—~1.4kAutomated safety check: PassMIT
Kimi PDFthvroyal/kimi-skills238—~1.9kAutomated safety check: PassNone
Lexoid CLIoidlabs-com/Lexoid109—~2kAutomated safety check: NotesApache-2.0

Similar skills

  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed
  • Tikz

    scunning1975/MixtapeTools

    Quick visual-collision check for figures — TikZ inside .tex files OR rendered .png/.jpg/.pdf figures from R/Python.

    473 GitHub stars~1.7k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~1.4k tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed
  • Kimi PDF

    thvroyal/kimi-skills

    Professional PDF solution. An agent skill from thvroyal/kimi-skills.

    238 GitHub stars~1.9k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Lexoid CLI

    oidlabs-com/Lexoid

    Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) from the terminal using the lexoid CLI.

    109 GitHub stars~2k tokensUpdated yesterday
    Documents & OfficeAuto-check: notes
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.8k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.8k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.8k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.8k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Document Gen Resilient Workflow

What does Document Gen Resilient Workflow do?

Multi-engine document generation with cascading PDF fallbacks and robust Unicode handling. Document Gen Resilient Workflow is an agent skill from HKUDS/OpenSpace.

When should I use Document Gen Resilient Workflow?

Document Gen Resilient Workflow fits situations like: tasks that involve PDF.

How do I install Document Gen Resilient Workflow in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill document-gen-resilient-workflow -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/document-gen-fallback-enhanced-enhanced-2794b4 in HKUDS/OpenSpace) into .claude/skills/document-gen-resilient-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Document Gen Resilient Workflow in Codex?

Run `npx skills add HKUDS/OpenSpace --skill document-gen-resilient-workflow -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/document-gen-fallback-enhanced-enhanced-2794b4 in HKUDS/OpenSpace) into .agents/skills/document-gen-resilient-workflow in your project. Codex loads it when a task matches its description.

Can I use Document Gen Resilient Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill document-gen-resilient-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-gen-resilient-workflow, .gemini/skills/document-gen-resilient-workflow, .github/skills/document-gen-resilient-workflow and .opencode/skills/document-gen-resilient-workflow in your project.

What does Document Gen Resilient Workflow need to run?

Going by SKILL.md and its folder, Document Gen Resilient Workflow needs the command-line tools its instructions call (apt-get, pip, pdftotext and pandoc). Our summary lists: Python 3.

Does Document Gen Resilient Workflow access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Document Gen Resilient Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Document Gen Resilient Workflow use?

Document Gen Resilient Workflow is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Document Gen Resilient Workflow use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Document Gen Resilient Workflow?

Skills that share tags, products or a category with Document Gen Resilient Workflow: Mineru (Nebutra/MinerU-Skill, 122 stars), Tikz (scunning1975/MixtapeTools, 473 stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and Kimi PDF (thvroyal/kimi-skills, 238 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Document Gen Resilient Workflow?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,750 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.