Agent skill

OCR And Documents

by RedWoodOG in RedWoodOG/Hermes-Desktop

Extract text from PDFs and scanned documents. An agent skill from RedWoodOG/Hermes-Desktop.

MITAuto-check passedDocuments & Office

Install OCR And Documents

skills CLI
$ npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RedWoodOG/Hermes-Desktop ocr-and-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/productivity/ocr-and-documents .claude/skills/ocr-and-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ocr-and-documents
GitHub stars
177
Used in
3 other repos
Token cost
~1.3k tokens
SKILL.md length
320 words
Files
4 (incl. scripts)
Skills in repo
62
Repo updated
First seen
Licence
MIT

At a glance

Extract text from PDFs and scanned documents. An agent skill from RedWoodOG/Hermes-Desktop.

  • Works in 2 steps: Remote URL Available? → Choose Local Extractor
  • Tasks that involve PDF
  • SKILL.md covers Step 1: Remote URL Available?, Step 2: Choose Local Extractor, pymupdf (lightweight) and marker-pdf (high-quality OCR), plus 3 more sections
  • Runs Python scripts from its folder; calls python, pip and python3; reaches arxiv.org

What it does

OCR And Documents is an agent skill from RedWoodOG/Hermes-Desktop. Extract text from PDFs and scanned documents. Use webextract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `DESCRIPTION.md`, `scripts/extract_marker.py` and `scripts/extract_pymupdf.py`).

It sits in Documents & Office, covering PDF, PowerPoint presentations and Word documents. It works with Microsoft PowerPoint, python-docx and Microsoft Word. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF
  • Tasks that involve PowerPoint presentations
  • Tasks that involve Word documents

Example prompts

  • “/ocr-and-documents”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Remote URL Available?
  2. Choose Local Extractor

What it can do on your machine

Read from SKILL.md and the folder at commit be46b39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OCR And Documents loads about 1.3k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 320 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from RedWoodOG/Hermes-Desktop at commit be46b39, republished under its MIT licence (© RedWoodOG). 320 words, ~1,335 tokens.

Download SKILL.mdSave it as .claude/skills/ocr-and-documents/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ocr-and-documents
description
Extract text from PDFs and scanned documents. Use web_extract for remote URLs, pymupdf for local text-based PDFs, marker-pdf for OCR/scanned docs. For DOCX use python-docx, for PPTX see the powerpoint skill.
version
2.3.0
author
Hermes Agent
license
MIT

PDF & Document Extraction

For DOCX: use python-docx (parses actual document structure, far better than OCR). For PPTX: see the powerpoint skill (uses python-pptx with full slide/notes support). This skill covers PDFs and scanned documents.

Step 1: Remote URL Available?

If the document has a URL, always try web_extract first:

web_extract(urls=["https://arxiv.org/pdf/2402.03300"])
web_extract(urls=["https://example.com/report.pdf"])

This handles PDF-to-markdown conversion via Firecrawl with no local dependencies.

Only use local extraction when: the file is local, web_extract fails, or you need batch processing.

Step 2: Choose Local Extractor

Featurepymupdf (~25MB)marker-pdf (~3-5GB)
Text-based PDF✅✅
Scanned PDF (OCR)❌✅ (90+ languages)
Tables✅ (basic)✅ (high accuracy)
Equations / LaTeX❌✅
Code blocks❌✅
Forms❌✅
Headers/footers removal❌✅
Reading order detection❌✅
Images extraction✅ (embedded)✅ (with context)
Images → text (OCR)❌✅
EPUB✅✅
Markdown output✅ (via pymupdf4llm)✅ (native, higher quality)
Install size~25MB~3-5GB (PyTorch + models)
SpeedInstant~1-14s/page (CPU), ~0.2s/page (GPU)

Decision: Use pymupdf unless you need OCR, equations, forms, or complex layout analysis.

If the user needs marker capabilities but the system lacks ~5GB free disk:

"This document needs OCR/advanced extraction (marker-pdf), which requires ~5GB for PyTorch and models. Your system has [X]GB free. Options: free up space, provide a URL so I can use web_extract, or I can try pymupdf which works for text-based PDFs but not scanned documents or equations."


pymupdf (lightweight)

bash
pip install pymupdf pymupdf4llm

Via helper script:

bash
python scripts/extract_pymupdf.py document.pdf              # Plain text
python scripts/extract_pymupdf.py document.pdf --markdown    # Markdown
python scripts/extract_pymupdf.py document.pdf --tables      # Tables
python scripts/extract_pymupdf.py document.pdf --images out/ # Extract images
python scripts/extract_pymupdf.py document.pdf --metadata    # Title, author, pages
python scripts/extract_pymupdf.py document.pdf --pages 0-4   # Specific pages

Inline:

bash
python3 -c "
import pymupdf
doc = pymupdf.open('document.pdf')
for page in doc:
    print(page.get_text())
"

marker-pdf (high-quality OCR)

bash
# Check disk space first
python scripts/extract_marker.py --check

pip install marker-pdf

Via helper script:

bash
python scripts/extract_marker.py document.pdf                # Markdown
python scripts/extract_marker.py document.pdf --json         # JSON with metadata
python scripts/extract_marker.py document.pdf --output_dir out/  # Save images
python scripts/extract_marker.py scanned.pdf                 # Scanned PDF (OCR)
python scripts/extract_marker.py document.pdf --use_llm      # LLM-boosted accuracy

CLI (installed with marker-pdf):

bash
marker_single document.pdf --output_dir ./output
marker /path/to/folder --workers 4    # Batch

Arxiv Papers

# Abstract only (fast)
web_extract(urls=["https://arxiv.org/abs/2402.03300"])

# Full paper
web_extract(urls=["https://arxiv.org/pdf/2402.03300"])

# Search
web_search(query="arxiv GRPO reinforcement learning 2026")

pymupdf handles these natively — use execute_code or inline Python:

python
# Split: extract pages 1-5 to a new PDF
import pymupdf
doc = pymupdf.open("report.pdf")
new = pymupdf.open()
for i in range(5):
    new.insert_pdf(doc, from_page=i, to_page=i)
new.save("pages_1-5.pdf")
python
# Merge multiple PDFs
import pymupdf
result = pymupdf.open()
for path in ["a.pdf", "b.pdf", "c.pdf"]:
    result.insert_pdf(pymupdf.open(path))
result.save("merged.pdf")
python
# Search for text across all pages
import pymupdf
doc = pymupdf.open("report.pdf")
for i, page in enumerate(doc):
    results = page.search_for("revenue")
    if results:
        print(f"Page {i+1}: {len(results)} match(es)")
        print(page.get_text("text"))

No extra dependencies needed — pymupdf covers split, merge, search, and text extraction in one package.


Notes

  • web_extract is always first choice for URLs
  • pymupdf is the safe default — instant, no models, works everywhere
  • marker-pdf is for OCR, scanned docs, equations, complex layouts — install only when needed
  • Both helper scripts accept --help for full usage
  • marker-pdf downloads ~2.5GB of models to ~/.cache/huggingface/ on first use
  • For Word docs: pip install python-docx (better than OCR — parses actual structure)
  • For PowerPoint: see the powerpoint skill (uses python-pptx)

© RedWoodOG, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/productivity/ocr-and-documents of RedWoodOG/Hermes-Desktop.

  • SKILL.md
  • DESCRIPTION.md
  • scripts/extract_marker.py
  • scripts/extract_pymupdf.py

Open the folder on GitHubat commit be46b39

Used in 3 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in RedWoodOG/Hermes-Desktop, which our catalogue first saw on October 7, 2026.

Compare with similar skills

OCR And Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OCR And Documents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OCR And Documents this skillRedWoodOG/Hermes-Desktop1773 repos~1.3kAutomated safety check: PassMIT
Office ArtifactsPrismer-AI/PrismerCloud1.6k—~2.6kAutomated safety check: PassMIT
Working With Documentsaiskillstore/marketplace433—~1.5kAutomated safety check: PassNone
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
GenOffice Document CLIgenspark-ai/genoffice9.2k—~19kAutomated safety check: PassApache-2.0
Docsagentdocsagent/docsagent625—~834Automated safety check: PassNone

Similar skills

  • Office Artifacts

    Prismer-AI/PrismerCloud

    Generate real DOCX, PPTX, XLSX, PDF, CSV files using python-docx / python-pptx / openpyxl / reportlab by writing them into the dispatch artifacts dir, then explicitly deliver each one with cloud…

    1.6k GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Working With Documents

    aiskillstore/marketplace

    Creates and edits Office documents: Word (.docx), PDF, and PowerPoint (.pptx).

    433 GitHub stars~1.5k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    9.2k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Docsagent

    docsagent/docsagent

    Search and manage private, local document collections (PDF, PPTX, DOCX) offline.

    625 GitHub stars~834 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • PDF

    zai-org/ZCode

    Professional PDF toolkit covering four production workflows: reports, creative visuals, academic LaTeX, and existing PDF processing.

    7.7k GitHub stars~18k tokensUpdated today
    Documents & OfficeAuto-check: notes

More from RedWoodOG/Hermes-Desktop

All 62 skills in this repo
  • Excalidraw

    RedWoodOG/Hermes-Desktop

    Create hand-drawn style diagrams using Excalidraw JSON format.

    177 GitHub starsUsed in 5 repos~1.8k tokens
    Auto-check passed
  • Obliteratus

    RedWoodOG/Hermes-Desktop

    Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…

    177 GitHub starsUsed in 5 repos~3.8k tokens
    Auto-check passed
  • Ascii Video

    RedWoodOG/Hermes-Desktop

    Production pipeline for ASCII art video — any format. An agent skill from RedWoodOG/Hermes-Desktop.

    177 GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Systematic Debugging

    RedWoodOG/Hermes-Desktop

    A skill your agent uses when encountering any bug, test failure, or unexpected behavior.

    177 GitHub starsUsed in 6 repos~2.6k tokens
    Auto-check passed
  • Test Driven Development

    RedWoodOG/Hermes-Desktop

    A skill your agent uses when implementing any feature or bugfix, before writing implementation code.

    177 GitHub starsUsed in 6 repos~2.4k tokens
    Auto-check passed
  • Claude Code

    RedWoodOG/Hermes-Desktop

    Delegate coding tasks to Claude Code (Anthropic's CLI agent).

    177 GitHub starsUsed in 4 repos~784 tokens
    Auto-check passed

Questions about OCR And Documents

What does OCR And Documents do?

Extract text from PDFs and scanned documents. An agent skill from RedWoodOG/Hermes-Desktop. OCR And Documents is an agent skill from RedWoodOG/Hermes-Desktop. Extract text from PDFs and scanned documents.

When should I use OCR And Documents?

OCR And Documents fits situations like: tasks that involve PDF; tasks that involve PowerPoint presentations; tasks that involve Word documents.

How do I install OCR And Documents in Claude Code?

Run `npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a claude-code`. Or copy the skill folder (skills/productivity/ocr-and-documents in RedWoodOG/Hermes-Desktop) into .claude/skills/ocr-and-documents in your project. Claude Code loads it when a task matches its description.

How do I install OCR And Documents in Codex?

Run `npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a codex`. Or copy the skill folder (skills/productivity/ocr-and-documents in RedWoodOG/Hermes-Desktop) into .agents/skills/ocr-and-documents in your project. Codex loads it when a task matches its description.

Can I use OCR And Documents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RedWoodOG/Hermes-Desktop --skill ocr-and-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ocr-and-documents, .gemini/skills/ocr-and-documents, .github/skills/ocr-and-documents and .opencode/skills/ocr-and-documents in your project.

What does OCR And Documents need to run?

Going by SKILL.md and its folder, OCR And Documents needs Python for the scripts in its folder and the command-line tools its instructions call (python, pip and python3). Our summary lists: Python 3.

Does OCR And Documents access the network?

SKILL.md names 1 domain. In commands or code: arxiv.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is OCR And Documents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does OCR And Documents use?

OCR And Documents is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OCR And Documents use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to OCR And Documents?

Skills that share tags, products or a category with OCR And Documents: Office Artifacts (Prismer-AI/PrismerCloud, 1.6k stars), Working With Documents (aiskillstore/marketplace, 433 stars), Markitdown (ImCa0/just-laws, 781 stars) and GenOffice Document CLI (genspark-ai/genoffice, 9.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OCR And Documents?

RedWoodOG (a GitHub user) maintains it in RedWoodOG/Hermes-Desktop, which has 177 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on May 30, 2026.

Source: RedWoodOG/Hermes-Desktop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.