Agent skill

Document Processing

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of…

MITAuto-check passedDocuments & Office

Install Document Processing

skills CLI
$ npx skills add ericrisco/rsc-harness --skill document-processing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness document-processing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/document-processing .claude/skills/document-processing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
document-processing
GitHub stars
156
Token cost
~2.4k tokens
SKILL.md length
893 words
Files
5 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of…

  • The deliverable is a documents bytes
  • SKILL.md covers Step 0 — does the PDF have a…, Engine selection, Extraction recipes and Form filling (AcroForm), plus 4 more sections
  • Runs Shell scripts from its folder; needs MISTRAL_API_KEY
  • Its literal content — text/tables out of PDFs

What it does

Document Processing is an agent skill from ericrisco/rsc-harness. Use when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of image-only scans. NOT schema-typed fields pulled from text (that is structured-extraction), NOT signature routing (e-signature) or spreadsheet cells/formulas (spreadsheet-ops).

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/engines.md`).

It sits in Documents & Office, covering PDF, Excel spreadsheets and Document parsing. It works with Microsoft Word and pypdf. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • The deliverable is a documents bytes
  • Its literal content — text/tables out of PDFs
  • AcroForm fill and flatten
  • Page merge/split

Example prompts

  • “/document-processing”

Requirements

  • Python 3
  • A Bash shell
  • A credential in MISTRAL_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MISTRAL_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Document Processing loads about 2.4k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 893 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 893 words, ~2,442 tokens.

Download SKILL.mdSave it as .claude/skills/document-processing/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
document-processing
description
Use when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of image-only scans. NOT schema-typed fields pulled from text (that is structured-extraction), NOT signature routing (e-signature) or spreadsheet cells/formulas (spreadsheet-ops).
tags
pdf, ocr, docx, forms, extraction, document-ai
recommends
structured-extraction, e-signature, spreadsheet-ops, rag, data-scraper
origin
risco

Document processing

File in, content out — or data in, file out. You open a byte stream (PDF, DOCX, scan) and either pull the content out, or you build a new document from a template and a data dict. That is the whole job: the deliverable is bytes of a document or the literal content of one.

The boundary test, apply it first:

  • Deliverable is raw text / Markdown / table cells / a generated file → you are in the right place.
  • Deliverable is a typed object matching a schema ({parties: [...], total: 1234.50}) → that is structured-extraction. This skill stops at "clean Markdown out of the file"; the schema-constrained extraction runs on that Markdown.

Everything else routes too: signing with an audit trail → e-signature, spreadsheet grids/formulas/XLSX-as-data → spreadsheet-ops, indexing for cross-document Q&A → rag (this skill produces the text rag ingests, it does not index it), downloading the files off a site → data-scraper.

Step 0 — does the PDF have a text layer?

The most expensive mistake in this skill is OCR'ing a PDF that already has a text layer. A digital PDF (exported from Word, a browser, a report tool) carries selectable text — extracting it is free, instant, and lossless. OCR is slow, costs money or GPU, and introduces errors. Never OCR a PDF you can extract.

Check before you pick an engine:

python
import pdfplumber

with pdfplumber.open("doc.pdf") as pdf:
    txt = pdf.pages[0].extract_text() or ""

if len(txt.strip()) > 20:
    print("text layer present -> extract directly (pdfplumber / pypdf)")
else:
    print("image-only or empty -> this is an OCR job")

If extract_text() returns empty (or near-empty) across the first few pages, it is a scan or image-only PDF and you go to the OCR branch. Symptom from the user's side: "the text copies out as garbage / random symbols" usually means a broken/embedded font, not a missing text layer — try pypdf extraction too before assuming OCR.

Engine selection

GoalUseWhy
Extract text + tables with layoutpdfplumberLayout-aware; extract_tables() returns rows/cols as Python lists → pandas/CSV.
Raw text, merge, split, rotate, page opspypdf (6.12.2)Pure-Python, no C deps, runs in Lambda/containers; the maintained core — import pypdf, never the dead PyPDF2, which was merged back into it.
Fill an interactive PDF formpypdfupdate_page_form_field_values writes AcroForm fields; can flatten.
Generate a Word/DOCX from a templatedocxtpl (0.20.x)A real .docx becomes a Jinja2 template; author in Word, tag, render.
Generate a PDF from scratchReportLabCanvas / Platypus flowables for laid-out PDFs.
OCR a scan, local / no API budgetDocling (or Marker)Layout + reading order + table structure, fully local, wraps Tesseract/RapidOCR.
OCR messy scans / handwriting / hard tables, API okMistral OCRmistral-ocr-2512 (OCR 3), ~$2 / 1,000 pages, tuned for forms + handwriting.
Fastest extract / easiest page→PNG rasterPyMuPDF ⚠️ AGPLFast, but AGPL: shipping it imposes an open-source obligation or needs a paid license. Flag this before recommending.

Extraction recipes

Text + tables with pdfplumber, straight to CSV:

python
import csv
import pdfplumber

rows = []
with pdfplumber.open("invoice.pdf") as pdf:
    for page in pdf.pages:
        for table in page.extract_tables():
            rows.extend(table)

with open("out.csv", "w", newline="") as f:
    csv.writer(f).writerows(rows)

Raw text, merge, split, rotate with pypdf:

python
from pypdf import PdfReader, PdfWriter

# raw text
text = "\n".join(p.extract_text() or "" for p in PdfReader("doc.pdf").pages)

# merge two files
w = PdfWriter()
for src in ("a.pdf", "b.pdf"):
    w.append(src)
with open("merged.pdf", "wb") as f:
    w.write(f)

# split first 3 pages + rotate one
w2 = PdfWriter()
reader = PdfReader("doc.pdf")
for page in reader.pages[:3]:
    w2.add_page(page)
w2.pages[0].rotate(90)
with open("first3.pdf", "wb") as f:
    w2.write(f)

Form filling (AcroForm)

Dump the field names first — guessing them is the #1 reason a fill silently does nothing:

python
from pypdf import PdfReader

fields = PdfReader("form.pdf").get_fields() or {}
for name, f in fields.items():
    print(name, "->", f.get("/FT"))  # /Tx text, /Btn checkbox/radio, /Ch choice

Then write the values. Set auto_regenerate=False and bake with flatten=True if it must not be editable:

python
from pypdf import PdfReader, PdfWriter

reader = PdfReader("form.pdf")
writer = PdfWriter()
writer.append(reader)

for page in writer.pages:
    writer.update_page_form_field_values(
        page,
        {"applicant_name": "Eric Risco", "agree": "/Yes"},  # checkbox = its on-state
        auto_regenerate=False,  # else a spurious "save changes?" prompt fires on open
    )

# flatten=True bakes the values and drops the editable widgets
with open("filled.pdf", "wb") as f:
    writer.write(f)

auto_regenerate defaults to True for legacy reasons, and you almost never want it. Checkbox/radio values are the field's /V on-state (often /Yes), not True — read the field to find it.

Generation

DOCX from a Word template you authored and tagged with Jinja2 ({{ client }}, {% tr for row in items %} on a table row, InlineImage for pictures):

python
from docxtpl import DocxTemplate

doc = DocxTemplate("contract_template.docx")
doc.render({
    "client": "Acme SL",
    "date": "2026-06-02",
    "items": [{"desc": "Audit", "amount": "1.200,00 €"}],
})
doc.save("contract_2026-06-02.docx")

PDF from scratch with ReportLab Platypus:

python
from reportlab.lib.pagesizes import A4
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer
from reportlab.lib.styles import getSampleStyleSheet

styles = getSampleStyleSheet()
doc = SimpleDocTemplate("report.pdf", pagesize=A4)
doc.build([
    Paragraph("Quarterly Report", styles["Title"]),
    Spacer(1, 12),
    Paragraph("Generated automatically from the data dict.", styles["BodyText"]),
])
Show full SKILL.md (352 more words)Show less

OCR

Branch on cost and privacy. Local, no API budget, or data must not leave the machine → Docling/Marker. Messy scans, handwriting, brutal tables, and an API budget is fine → Mistral OCR.

Local with Docling (wraps Tesseract / RapidOCR, exports Markdown preserving tables):

python
from docling.document_converter import DocumentConverter

result = DocumentConverter().convert("scan.pdf")
markdown = result.document.export_to_markdown()
open("scan.md", "w").write(markdown)

Hosted with Mistral OCR (~$2 / 1,000 pages, 50% off via Batch API; outputs interleaved text+images as Markdown):

python
from mistralai import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
resp = client.ocr.process(
    model="mistral-ocr-2512",
    document={"type": "document_url", "document_url": signed_url},
)
markdown = "\n\n".join(p.markdown for p in resp.pages)

Never trust OCR output blind. OCR confuses 0/O, 1/l/I, and drops or shifts decimal points — a 1.234,50 can come back as 1234,50 or 1,234.50. Always spot-check totals, dates, and ID numbers against the rendered page before you hand the text downstream. For clean scans with no budget, plain pytesseract is the zero-cost baseline, but it is weak on layout/tables versus the pipelines above.

Scale

Batch jobs: parallelize per-file, cap concurrency on the hosted API (rate limits + cost), and use Mistral's Batch API for the 50% discount on large runs. Engine install matrix, exact version pins, the full licensing table, the Docling-vs-Marker-vs-Mistral feature/cost comparison, and troubleshooting (encrypted PDFs, mangled AcroForm field names, multi-column reading order, CJK/handwriting) live in references/engines.md — read it before a non-trivial install.

Anti-patterns

Anti-patternWhy it is wrongDo instead
Pipe every PDF straight to OCROCR'ing a digital PDF is slow, costs money, and adds errors to text you could extract losslesslyStep 0: check the text layer first; OCR only image-only PDFs
import PyPDF2Unmaintained; merged into pypdf years ago — a stale-code smellfrom pypdf import PdfReader, PdfWriter
Recommend PyMuPDF without a word about its licensePyMuPDF is AGPL; shipping it silently creates an open-source obligationFlag AGPL; prefer pdfplumber/pypdf, or get a commercial license knowingly
Leave auto_regenerate=True on a form fillMarks the AcroForm dirty → a spurious "save changes?" prompt for every userPass auto_regenerate=False
Trust OCR'd totals/numbers as-is0/O, 1/l, shifted decimals silently corrupt amountsSpot-check totals/dates/IDs against the page image
Hand-roll a regex to pull typed fields from the MarkdownBrittle, re-implements a sibling, breaks on layout driftOutput clean Markdown, hand it to structured-extraction
Use Mistral OCR when the user said "no cloud / local only"Sends documents off-machine, violating the privacy constraintUse Docling/Marker + Tesseract/RapidOCR locally

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/document-processing of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/engines.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Document Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Document Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Document Processing this skillericrisco/rsc-harness156—~2.4kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Nutrient Document Processingaffaan-m/ECC274k4 repos~1.5kAutomated safety check: PassMIT
Extracting Lab Tablesmaziyarpanahi/openmed5.5k—~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.

    274k GitHub starsUsed in 4 repos~1.5k tokens
    Documents & OfficeAuto-check passed
  • Extracting Lab Tables

    maziyarpanahi/openmed

    Detects and extracts tabular laboratory panels from PDFs, scans, and images into structured rows ready for OpenMed and FHIR.

    5.5k GitHub stars~2.1k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Document Processing

What does Document Processing do?

A skill your agent uses when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of…. Document Processing is an agent skill from ericrisco/rsc-harness. Use when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of image-only scans.

When should I use Document Processing?

Document Processing fits situations like: the deliverable is a documents bytes; its literal content — text/tables out of PDFs; acroForm fill and flatten; page merge/split.

How do I install Document Processing in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill document-processing -a claude-code`. Or copy the skill folder (skills/document-processing in ericrisco/rsc-harness) into .claude/skills/document-processing in your project. Claude Code loads it when a task matches its description.

How do I install Document Processing in Codex?

Run `npx skills add ericrisco/rsc-harness --skill document-processing -a codex`. Or copy the skill folder (skills/document-processing in ericrisco/rsc-harness) into .agents/skills/document-processing in your project. Codex loads it when a task matches its description.

Can I use Document Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill document-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-processing, .gemini/skills/document-processing, .github/skills/document-processing and .opencode/skills/document-processing in your project.

What does Document Processing need to run?

Going by SKILL.md and its folder, Document Processing needs a shell for the scripts in its folder and credentials named MISTRAL_API_KEY. Our summary lists: Python 3; A Bash shell; A credential in MISTRAL_API_KEY.

Does Document Processing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Document Processing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Document Processing use?

Document Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Document Processing use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Document Processing?

Skills that share tags, products or a category with Document Processing: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and Nutrient Document Processing (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Document Processing?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.