Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
A skill your agent uses when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of…
$ npx skills add ericrisco/rsc-harness --skill document-processing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness document-processing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/document-processing .claude/skills/document-processing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "document-processing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/document-processing into .claude/skills/document-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-processing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/document-processingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill document-processing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness document-processing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/document-processing .agents/skills/document-processing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "document-processing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/document-processing into .agents/skills/document-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-processing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill document-processing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness document-processing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/document-processing .cursor/skills/document-processing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "document-processing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/document-processing into .cursor/skills/document-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-processing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/document-processing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill document-processing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness document-processing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/document-processing .gemini/skills/document-processing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "document-processing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/document-processing into .gemini/skills/document-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-processing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness document-processingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill document-processing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/document-processing .github/skills/document-processing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "document-processing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/document-processing into .github/skills/document-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-processing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill document-processing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness document-processing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/document-processing .opencode/skills/document-processing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "document-processing" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/document-processing into .opencode/skills/document-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-processing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
document-processingA skill your agent uses when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of…
Document Processing is an agent skill from ericrisco/rsc-harness. Use when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of image-only scans. NOT schema-typed fields pulled from text (that is structured-extraction), NOT signature routing (e-signature) or spreadsheet cells/formulas (spreadsheet-ops).
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/engines.md`).
It sits in Documents & Office, covering PDF, Excel spreadsheets and Document parsing. It works with Microsoft Word and pypdf. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
MISTRAL_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Document Processing loads about 2.4k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 893 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 893 words, ~2,442 tokens.
.claude/skills/document-processing/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.File in, content out — or data in, file out. You open a byte stream (PDF, DOCX, scan) and either pull the content out, or you build a new document from a template and a data dict. That is the whole job: the deliverable is bytes of a document or the literal content of one.
The boundary test, apply it first:
{parties: [...], total: 1234.50}) → that is structured-extraction. This skill stops at "clean Markdown out of the file"; the schema-constrained extraction runs on that Markdown.Everything else routes too: signing with an audit trail → e-signature, spreadsheet grids/formulas/XLSX-as-data → spreadsheet-ops, indexing for cross-document Q&A → rag (this skill produces the text rag ingests, it does not index it), downloading the files off a site → data-scraper.
The most expensive mistake in this skill is OCR'ing a PDF that already has a text layer. A digital PDF (exported from Word, a browser, a report tool) carries selectable text — extracting it is free, instant, and lossless. OCR is slow, costs money or GPU, and introduces errors. Never OCR a PDF you can extract.
Check before you pick an engine:
import pdfplumber
with pdfplumber.open("doc.pdf") as pdf:
txt = pdf.pages[0].extract_text() or ""
if len(txt.strip()) > 20:
print("text layer present -> extract directly (pdfplumber / pypdf)")
else:
print("image-only or empty -> this is an OCR job")If extract_text() returns empty (or near-empty) across the first few pages, it is a scan or image-only PDF and you go to the OCR branch. Symptom from the user's side: "the text copies out as garbage / random symbols" usually means a broken/embedded font, not a missing text layer — try pypdf extraction too before assuming OCR.
| Goal | Use | Why |
|---|---|---|
| Extract text + tables with layout | pdfplumber | Layout-aware; extract_tables() returns rows/cols as Python lists → pandas/CSV. |
| Raw text, merge, split, rotate, page ops | pypdf (6.12.2) | Pure-Python, no C deps, runs in Lambda/containers; the maintained core — import pypdf, never the dead PyPDF2, which was merged back into it. |
| Fill an interactive PDF form | pypdf | update_page_form_field_values writes AcroForm fields; can flatten. |
| Generate a Word/DOCX from a template | docxtpl (0.20.x) | A real .docx becomes a Jinja2 template; author in Word, tag, render. |
| Generate a PDF from scratch | ReportLab | Canvas / Platypus flowables for laid-out PDFs. |
| OCR a scan, local / no API budget | Docling (or Marker) | Layout + reading order + table structure, fully local, wraps Tesseract/RapidOCR. |
| OCR messy scans / handwriting / hard tables, API ok | Mistral OCR | mistral-ocr-2512 (OCR 3), ~$2 / 1,000 pages, tuned for forms + handwriting. |
| Fastest extract / easiest page→PNG raster | PyMuPDF ⚠️ AGPL | Fast, but AGPL: shipping it imposes an open-source obligation or needs a paid license. Flag this before recommending. |
Text + tables with pdfplumber, straight to CSV:
import csv
import pdfplumber
rows = []
with pdfplumber.open("invoice.pdf") as pdf:
for page in pdf.pages:
for table in page.extract_tables():
rows.extend(table)
with open("out.csv", "w", newline="") as f:
csv.writer(f).writerows(rows)Raw text, merge, split, rotate with pypdf:
from pypdf import PdfReader, PdfWriter
# raw text
text = "\n".join(p.extract_text() or "" for p in PdfReader("doc.pdf").pages)
# merge two files
w = PdfWriter()
for src in ("a.pdf", "b.pdf"):
w.append(src)
with open("merged.pdf", "wb") as f:
w.write(f)
# split first 3 pages + rotate one
w2 = PdfWriter()
reader = PdfReader("doc.pdf")
for page in reader.pages[:3]:
w2.add_page(page)
w2.pages[0].rotate(90)
with open("first3.pdf", "wb") as f:
w2.write(f)Dump the field names first — guessing them is the #1 reason a fill silently does nothing:
from pypdf import PdfReader
fields = PdfReader("form.pdf").get_fields() or {}
for name, f in fields.items():
print(name, "->", f.get("/FT")) # /Tx text, /Btn checkbox/radio, /Ch choiceThen write the values. Set auto_regenerate=False and bake with flatten=True if it must not be editable:
from pypdf import PdfReader, PdfWriter
reader = PdfReader("form.pdf")
writer = PdfWriter()
writer.append(reader)
for page in writer.pages:
writer.update_page_form_field_values(
page,
{"applicant_name": "Eric Risco", "agree": "/Yes"}, # checkbox = its on-state
auto_regenerate=False, # else a spurious "save changes?" prompt fires on open
)
# flatten=True bakes the values and drops the editable widgets
with open("filled.pdf", "wb") as f:
writer.write(f)auto_regenerate defaults to True for legacy reasons, and you almost never want it. Checkbox/radio values are the field's /V on-state (often /Yes), not True — read the field to find it.
DOCX from a Word template you authored and tagged with Jinja2 ({{ client }}, {% tr for row in items %} on a table row, InlineImage for pictures):
from docxtpl import DocxTemplate
doc = DocxTemplate("contract_template.docx")
doc.render({
"client": "Acme SL",
"date": "2026-06-02",
"items": [{"desc": "Audit", "amount": "1.200,00 €"}],
})
doc.save("contract_2026-06-02.docx")PDF from scratch with ReportLab Platypus:
from reportlab.lib.pagesizes import A4
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer
from reportlab.lib.styles import getSampleStyleSheet
styles = getSampleStyleSheet()
doc = SimpleDocTemplate("report.pdf", pagesize=A4)
doc.build([
Paragraph("Quarterly Report", styles["Title"]),
Spacer(1, 12),
Paragraph("Generated automatically from the data dict.", styles["BodyText"]),
])Branch on cost and privacy. Local, no API budget, or data must not leave the machine → Docling/Marker. Messy scans, handwriting, brutal tables, and an API budget is fine → Mistral OCR.
Local with Docling (wraps Tesseract / RapidOCR, exports Markdown preserving tables):
from docling.document_converter import DocumentConverter
result = DocumentConverter().convert("scan.pdf")
markdown = result.document.export_to_markdown()
open("scan.md", "w").write(markdown)Hosted with Mistral OCR (~$2 / 1,000 pages, 50% off via Batch API; outputs interleaved text+images as Markdown):
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
resp = client.ocr.process(
model="mistral-ocr-2512",
document={"type": "document_url", "document_url": signed_url},
)
markdown = "\n\n".join(p.markdown for p in resp.pages)Never trust OCR output blind. OCR confuses 0/O, 1/l/I, and drops or shifts decimal points — a 1.234,50 can come back as 1234,50 or 1,234.50. Always spot-check totals, dates, and ID numbers against the rendered page before you hand the text downstream. For clean scans with no budget, plain pytesseract is the zero-cost baseline, but it is weak on layout/tables versus the pipelines above.
Batch jobs: parallelize per-file, cap concurrency on the hosted API (rate limits + cost), and use Mistral's Batch API for the 50% discount on large runs. Engine install matrix, exact version pins, the full licensing table, the Docling-vs-Marker-vs-Mistral feature/cost comparison, and troubleshooting (encrypted PDFs, mangled AcroForm field names, multi-column reading order, CJK/handwriting) live in references/engines.md — read it before a non-trivial install.
| Anti-pattern | Why it is wrong | Do instead |
|---|---|---|
| Pipe every PDF straight to OCR | OCR'ing a digital PDF is slow, costs money, and adds errors to text you could extract losslessly | Step 0: check the text layer first; OCR only image-only PDFs |
import PyPDF2 | Unmaintained; merged into pypdf years ago — a stale-code smell | from pypdf import PdfReader, PdfWriter |
| Recommend PyMuPDF without a word about its license | PyMuPDF is AGPL; shipping it silently creates an open-source obligation | Flag AGPL; prefer pdfplumber/pypdf, or get a commercial license knowingly |
Leave auto_regenerate=True on a form fill | Marks the AcroForm dirty → a spurious "save changes?" prompt for every user | Pass auto_regenerate=False |
| Trust OCR'd totals/numbers as-is | 0/O, 1/l, shifted decimals silently corrupt amounts | Spot-check totals/dates/IDs against the page image |
| Hand-roll a regex to pull typed fields from the Markdown | Brittle, re-implements a sibling, breaks on layout drift | Output clean Markdown, hand it to structured-extraction |
| Use Mistral OCR when the user said "no cloud / local only" | Sends documents off-machine, violating the privacy constraint | Use Docling/Marker + Tesseract/RapidOCR locally |
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references) in skills/document-processing of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Document Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Document Processing this skillericrisco/rsc-harness | 156 | — | ~2.4k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markitdownjimmc414/Kosmos | 594 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| MineruNebutra/MinerU-Skill | 122 | — | ~504 | Automated safety check: Pass | MIT | |
| Nutrient Document Processingaffaan-m/ECC | 274k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Extracting Lab Tablesmaziyarpanahi/openmed | 5.5k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
affaan-m/ECC
Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.
maziyarpanahi/openmed
Detects and extracts tabular laboratory panels from PDFs, scans, and images into structured rows ready for OpenMed and FHIR.
wentorai/Research-Claw
Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of…. Document Processing is an agent skill from ericrisco/rsc-harness. Use when the deliverable is a document's bytes or its literal content — text/tables out of PDFs, AcroForm fill and flatten, page merge/split, PDF/DOCX from templates, OCR of image-only scans.
Document Processing fits situations like: the deliverable is a documents bytes; its literal content — text/tables out of PDFs; acroForm fill and flatten; page merge/split.
Run `npx skills add ericrisco/rsc-harness --skill document-processing -a claude-code`. Or copy the skill folder (skills/document-processing in ericrisco/rsc-harness) into .claude/skills/document-processing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill document-processing -a codex`. Or copy the skill folder (skills/document-processing in ericrisco/rsc-harness) into .agents/skills/document-processing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill document-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-processing, .gemini/skills/document-processing, .github/skills/document-processing and .opencode/skills/document-processing in your project.
Going by SKILL.md and its folder, Document Processing needs a shell for the scripts in its folder and credentials named MISTRAL_API_KEY. Our summary lists: Python 3; A Bash shell; A credential in MISTRAL_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Document Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Document Processing: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and Nutrient Document Processing (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.