Office Artifacts
Prismer-AI/PrismerCloud
Generate real DOCX, PPTX, XLSX, PDF, CSV files using python-docx / python-pptx / openpyxl / reportlab by writing them into the dispatch artifacts dir, then explicitly deliver each one with cloud…
Extract text from PDFs and scanned documents. An agent skill from RedWoodOG/Hermes-Desktop.
$ npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install RedWoodOG/Hermes-Desktop ocr-and-documents --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/productivity/ocr-and-documents .claude/skills/ocr-and-documents && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ocr-and-documents" agent skill from https://github.com/RedWoodOG/Hermes-Desktop/tree/main/skills/productivity/ocr-and-documents into .claude/skills/ocr-and-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ocr-and-documents", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/RedWoodOG/Hermes-Desktop/tree/main/skills/productivity/ocr-and-documentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install RedWoodOG/Hermes-Desktop ocr-and-documents --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/productivity/ocr-and-documents .agents/skills/ocr-and-documents && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ocr-and-documents" agent skill from https://github.com/RedWoodOG/Hermes-Desktop/tree/main/skills/productivity/ocr-and-documents into .agents/skills/ocr-and-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ocr-and-documents", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install RedWoodOG/Hermes-Desktop ocr-and-documents --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/productivity/ocr-and-documents .cursor/skills/ocr-and-documents && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ocr-and-documents" agent skill from https://github.com/RedWoodOG/Hermes-Desktop/tree/main/skills/productivity/ocr-and-documents into .cursor/skills/ocr-and-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ocr-and-documents", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/RedWoodOG/Hermes-Desktop.git --path skills/productivity/ocr-and-documents--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install RedWoodOG/Hermes-Desktop ocr-and-documents --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/productivity/ocr-and-documents .gemini/skills/ocr-and-documents && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ocr-and-documents" agent skill from https://github.com/RedWoodOG/Hermes-Desktop/tree/main/skills/productivity/ocr-and-documents into .gemini/skills/ocr-and-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ocr-and-documents", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install RedWoodOG/Hermes-Desktop ocr-and-documentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/productivity/ocr-and-documents .github/skills/ocr-and-documents && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ocr-and-documents" agent skill from https://github.com/RedWoodOG/Hermes-Desktop/tree/main/skills/productivity/ocr-and-documents into .github/skills/ocr-and-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ocr-and-documents", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install RedWoodOG/Hermes-Desktop ocr-and-documents --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/productivity/ocr-and-documents .opencode/skills/ocr-and-documents && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ocr-and-documents" agent skill from https://github.com/RedWoodOG/Hermes-Desktop/tree/main/skills/productivity/ocr-and-documents into .opencode/skills/ocr-and-documents/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ocr-and-documents", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ocr-and-documentsExtract text from PDFs and scanned documents. An agent skill from RedWoodOG/Hermes-Desktop.
OCR And Documents is an agent skill from RedWoodOG/Hermes-Desktop. Extract text from PDFs and scanned documents. Use webextract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `DESCRIPTION.md`, `scripts/extract_marker.py` and `scripts/extract_pymupdf.py`).
It sits in Documents & Office, covering PDF, PowerPoint presentations and Word documents. It works with Microsoft PowerPoint, python-docx and Microsoft Word. The licence is MIT.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit be46b39. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpippython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
arxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
OCR And Documents loads about 1.3k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 320 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from RedWoodOG/Hermes-Desktop at commit be46b39, republished under its MIT licence (© RedWoodOG). 320 words, ~1,335 tokens.
.claude/skills/ocr-and-documents/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.For DOCX: use python-docx (parses actual document structure, far better than OCR).
For PPTX: see the powerpoint skill (uses python-pptx with full slide/notes support).
This skill covers PDFs and scanned documents.
If the document has a URL, always try web_extract first:
web_extract(urls=["https://arxiv.org/pdf/2402.03300"])
web_extract(urls=["https://example.com/report.pdf"])This handles PDF-to-markdown conversion via Firecrawl with no local dependencies.
Only use local extraction when: the file is local, web_extract fails, or you need batch processing.
| Feature | pymupdf (~25MB) | marker-pdf (~3-5GB) |
|---|---|---|
| Text-based PDF | ✅ | ✅ |
| Scanned PDF (OCR) | ❌ | ✅ (90+ languages) |
| Tables | ✅ (basic) | ✅ (high accuracy) |
| Equations / LaTeX | ❌ | ✅ |
| Code blocks | ❌ | ✅ |
| Forms | ❌ | ✅ |
| Headers/footers removal | ❌ | ✅ |
| Reading order detection | ❌ | ✅ |
| Images extraction | ✅ (embedded) | ✅ (with context) |
| Images → text (OCR) | ❌ | ✅ |
| EPUB | ✅ | ✅ |
| Markdown output | ✅ (via pymupdf4llm) | ✅ (native, higher quality) |
| Install size | ~25MB | ~3-5GB (PyTorch + models) |
| Speed | Instant | ~1-14s/page (CPU), ~0.2s/page (GPU) |
Decision: Use pymupdf unless you need OCR, equations, forms, or complex layout analysis.
If the user needs marker capabilities but the system lacks ~5GB free disk:
"This document needs OCR/advanced extraction (marker-pdf), which requires ~5GB for PyTorch and models. Your system has [X]GB free. Options: free up space, provide a URL so I can use web_extract, or I can try pymupdf which works for text-based PDFs but not scanned documents or equations."
pip install pymupdf pymupdf4llmVia helper script:
python scripts/extract_pymupdf.py document.pdf # Plain text
python scripts/extract_pymupdf.py document.pdf --markdown # Markdown
python scripts/extract_pymupdf.py document.pdf --tables # Tables
python scripts/extract_pymupdf.py document.pdf --images out/ # Extract images
python scripts/extract_pymupdf.py document.pdf --metadata # Title, author, pages
python scripts/extract_pymupdf.py document.pdf --pages 0-4 # Specific pagesInline:
python3 -c "
import pymupdf
doc = pymupdf.open('document.pdf')
for page in doc:
print(page.get_text())
"# Check disk space first
python scripts/extract_marker.py --check
pip install marker-pdfVia helper script:
python scripts/extract_marker.py document.pdf # Markdown
python scripts/extract_marker.py document.pdf --json # JSON with metadata
python scripts/extract_marker.py document.pdf --output_dir out/ # Save images
python scripts/extract_marker.py scanned.pdf # Scanned PDF (OCR)
python scripts/extract_marker.py document.pdf --use_llm # LLM-boosted accuracyCLI (installed with marker-pdf):
marker_single document.pdf --output_dir ./output
marker /path/to/folder --workers 4 # Batch# Abstract only (fast)
web_extract(urls=["https://arxiv.org/abs/2402.03300"])
# Full paper
web_extract(urls=["https://arxiv.org/pdf/2402.03300"])
# Search
web_search(query="arxiv GRPO reinforcement learning 2026")pymupdf handles these natively — use execute_code or inline Python:
# Split: extract pages 1-5 to a new PDF
import pymupdf
doc = pymupdf.open("report.pdf")
new = pymupdf.open()
for i in range(5):
new.insert_pdf(doc, from_page=i, to_page=i)
new.save("pages_1-5.pdf")# Merge multiple PDFs
import pymupdf
result = pymupdf.open()
for path in ["a.pdf", "b.pdf", "c.pdf"]:
result.insert_pdf(pymupdf.open(path))
result.save("merged.pdf")# Search for text across all pages
import pymupdf
doc = pymupdf.open("report.pdf")
for i, page in enumerate(doc):
results = page.search_for("revenue")
if results:
print(f"Page {i+1}: {len(results)} match(es)")
print(page.get_text("text"))No extra dependencies needed — pymupdf covers split, merge, search, and text extraction in one package.
web_extract is always first choice for URLs--help for full usage~/.cache/huggingface/ on first usepip install python-docx (better than OCR — parses actual structure)powerpoint skill (uses python-pptx)© RedWoodOG, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in skills/productivity/ocr-and-documents of RedWoodOG/Hermes-Desktop.
Open the folder on GitHubat commit be46b39
We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in RedWoodOG/Hermes-Desktop, which our catalogue first saw on October 7, 2026.
OCR And Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| OCR And Documents this skillRedWoodOG/Hermes-Desktop | 177 | 3 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Office ArtifactsPrismer-AI/PrismerCloud | 1.6k | — | ~2.6k | Automated safety check: Pass | MIT | |
| Working With Documentsaiskillstore/marketplace | 433 | — | ~1.5k | Automated safety check: Pass | None | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| GenOffice Document CLIgenspark-ai/genoffice | 9.2k | — | ~19k | Automated safety check: Pass | Apache-2.0 | |
| Docsagentdocsagent/docsagent | 625 | — | ~834 | Automated safety check: Pass | None |
Prismer-AI/PrismerCloud
Generate real DOCX, PPTX, XLSX, PDF, CSV files using python-docx / python-pptx / openpyxl / reportlab by writing them into the dispatch artifacts dir, then explicitly deliver each one with cloud…
aiskillstore/marketplace
Creates and edits Office documents: Word (.docx), PDF, and PowerPoint (.pptx).
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
genspark-ai/genoffice
Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.
docsagent/docsagent
Search and manage private, local document collections (PDF, PPTX, DOCX) offline.
zai-org/ZCode
Professional PDF toolkit covering four production workflows: reports, creative visuals, academic LaTeX, and existing PDF processing.
RedWoodOG/Hermes-Desktop
Create hand-drawn style diagrams using Excalidraw JSON format.
RedWoodOG/Hermes-Desktop
Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…
RedWoodOG/Hermes-Desktop
Production pipeline for ASCII art video — any format. An agent skill from RedWoodOG/Hermes-Desktop.
RedWoodOG/Hermes-Desktop
A skill your agent uses when encountering any bug, test failure, or unexpected behavior.
RedWoodOG/Hermes-Desktop
A skill your agent uses when implementing any feature or bugfix, before writing implementation code.
RedWoodOG/Hermes-Desktop
Delegate coding tasks to Claude Code (Anthropic's CLI agent).
Categories
Extract text from PDFs and scanned documents. An agent skill from RedWoodOG/Hermes-Desktop. OCR And Documents is an agent skill from RedWoodOG/Hermes-Desktop. Extract text from PDFs and scanned documents.
OCR And Documents fits situations like: tasks that involve PDF; tasks that involve PowerPoint presentations; tasks that involve Word documents.
Run `npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a claude-code`. Or copy the skill folder (skills/productivity/ocr-and-documents in RedWoodOG/Hermes-Desktop) into .claude/skills/ocr-and-documents in your project. Claude Code loads it when a task matches its description.
Run `npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a codex`. Or copy the skill folder (skills/productivity/ocr-and-documents in RedWoodOG/Hermes-Desktop) into .agents/skills/ocr-and-documents in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ocr-and-documents, .gemini/skills/ocr-and-documents, .github/skills/ocr-and-documents and .opencode/skills/ocr-and-documents in your project.
Going by SKILL.md and its folder, OCR And Documents needs Python for the scripts in its folder and the command-line tools its instructions call (python, pip and python3). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: arxiv.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
OCR And Documents is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with OCR And Documents: Office Artifacts (Prismer-AI/PrismerCloud, 1.6k stars), Working With Documents (aiskillstore/marketplace, 433 stars), Markitdown (ImCa0/just-laws, 781 stars) and GenOffice Document CLI (genspark-ai/genoffice, 9.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
RedWoodOG (a GitHub user) maintains it in RedWoodOG/Hermes-Desktop, which has 177 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on May 30, 2026.
Source: RedWoodOG/Hermes-Desktop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.