Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.
$ npx skills add mrmps/classifier-dev --skill document-intake-routing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mrmps/classifier-dev document-intake-routing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/document-intake-routing .claude/skills/document-intake-routing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "document-intake-routing" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/document-intake-routing into .claude/skills/document-intake-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-intake-routing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mrmps/classifier-dev/tree/main/skills/document-intake-routingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mrmps/classifier-dev --skill document-intake-routing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mrmps/classifier-dev document-intake-routing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/document-intake-routing .agents/skills/document-intake-routing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "document-intake-routing" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/document-intake-routing into .agents/skills/document-intake-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-intake-routing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mrmps/classifier-dev --skill document-intake-routing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mrmps/classifier-dev document-intake-routing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/document-intake-routing .cursor/skills/document-intake-routing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "document-intake-routing" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/document-intake-routing into .cursor/skills/document-intake-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-intake-routing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mrmps/classifier-dev.git --path skills/document-intake-routing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mrmps/classifier-dev --skill document-intake-routing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mrmps/classifier-dev document-intake-routing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/document-intake-routing .gemini/skills/document-intake-routing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "document-intake-routing" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/document-intake-routing into .gemini/skills/document-intake-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-intake-routing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mrmps/classifier-dev document-intake-routingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mrmps/classifier-dev --skill document-intake-routing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/document-intake-routing .github/skills/document-intake-routing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "document-intake-routing" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/document-intake-routing into .github/skills/document-intake-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-intake-routing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mrmps/classifier-dev --skill document-intake-routing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mrmps/classifier-dev document-intake-routing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/document-intake-routing .opencode/skills/document-intake-routing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "document-intake-routing" agent skill from https://github.com/mrmps/classifier-dev/tree/main/skills/document-intake-routing into .opencode/skills/document-intake-routing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "document-intake-routing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
document-intake-routingLabel each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.
Document Intake Routing is an agent skill from mrmps/classifier-dev. Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person. Use when a mailroom, loan file, claim packet or vendor upload arrives as one long PDF. Triggers on "route these scanned forms", "which form type is each page", "split this intake packet", "sort these scans by form type".
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering PDF. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b9211dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pdftotextFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
classifier.devFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Document Intake Routing loads about 1.5k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 573 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mrmps/classifier-dev at commit b9211dd, republished under its MIT licence (© mrmps). 573 words, ~1,497 tokens.
.claude/skills/document-intake-routing/SKILL.md (or your agent's skills folder).A packet is one PDF holding several documents: a cover sheet, two forms, a pile
of attachments, some blank backs. Extraction fails on it because it is pointed
at the wrong page. Classify pages first, then extract what you are sure of.
The classifier below is classifier.dev: keyless HTTP, a label per page, a
calibrated confidence, no generated text.
Each request carries page text and your label names: no file, no image, no file name. The service states it stores no input text and forwards it to the model that answers (https://classifier.dev/privacy) — still a third party, and page one of a tax or loan packet is where the name, SSN and account number sit. Redact first, every time:
import re
PATTERNS = [(r"(?i)\b(?:bearer|basic)\s+[\w.\-+/=]{8,}", "<CRED>"),
(r"(?i)\b[\w.-]*(?:key|token|secret|password|pwd)[\w.-]*\s*[=:]\s*[^\s\"',&]{6,}", "<CRED>"),
(r"\b[A-Za-z0-9+/]{32,}={0,2}\b", "<BLOB>"), (r"\b[0-9a-f]{16,}\b", "<BLOB>"),
(r"[\w.+-]+@[\w-]+\.[\w.]{2,}", "<EMAIL>"), (r"\b(?:\d[ -]?){13,16}\b", "<CARD>"),
(r"\b\d{3}-\d{2}-\d{4}\b", "<SSN>"), (r"\b\d{7,}\b", "<NUM>")]
def redact(t):
for p, tag in PATTERNS:
t = re.sub(p, tag, t)
return t
def sendable(t): # never send a mostly-redacted page
kept = len(re.sub(r"<[A-Z]+>", "", redact(t)))
return kept >= 25 and kept >= 0.5 * len(t)On a filled page:
Name as shown on your income tax return: Jane Q. Public. Social security
number <SSN>. Account <CARD>. Email <EMAIL> W-9 Request for Taxpayer Form ...That page classified the same either way: type 1.00 raw and 1.00 redacted, role
0.97 against 0.91. Redaction costs almost nothing. It catches shapes, not names:
Jane Q. Public survives it.
If these packets are confidential under the policy you work under, do not call
out at all. The workflow is a classifier plus two thresholds, and anything
giving a calibrated confidence fits: your own model over the same labels, or a
local zero-shot model. classifier.dev is the keyless example because it needs
no setup. Same steps, other call, no network.
Skip it too for a one-document PDF you already know, for pulling fields out (extraction, not classification), and for image-only scans: OCR first, since a page with no text layer classifies as nothing.
pip install pdfplumber==0.11.4import pdfplumber
with pdfplumber.open("packet.pdf") as pdf:
raw = [" ".join((p.extract_text() or "").split()) for p in pdf.pages]
pages = [(i, redact(t[:1200])) for i, t in enumerate(raw, 1) if sendable(t)]1,200 characters is enough. pdftotext -f N -l N packet.pdf - does the same in
the shell. Empty pages drop out here: a whitespace-only input fails the whole
batch with empty_input.
One question per pass: a combined set ("W-9 instructions page") multiplies the labels and splits the probability mass.
import json, urllib.request
FORMS = ["IRS Form W-9 request for taxpayer identification number",
"IRS Form W-4 employee withholding certificate", "commercial invoice",
"bank statement", "none of these"]
ROLES = ["cover or transmittal page", "filled form page with input fields",
"instructions or attachment page", "blank or nearly empty page"]
def classify(labels, instructions=""):
body = {"labels": labels, "inputs": [t for _, t in pages],
"instructions": instructions}
req = urllib.request.Request("https://classifier.dev/v1/classify",
json.dumps(body).encode(),
{"content-type": "application/json", "user-agent": "intake/1.0"})
return json.load(urllib.request.urlopen(req))["results"]
forms = classify(FORMS, "Identify this page's document type.")
roles = classify(ROLES)One request per pass, up to 1,000 pages each; tens of types is normal, 100 is
the ceiling. Set a user-agent: Python's default is blocked at the edge.
Name the document, not your code for it: the same 11 pages against
DOC_TYPE_01 ... DOC_TYPE_06, OTHER all came back DOC_TYPE_01, at 0.47, 0.34,
0.30, 0.26 — one bucket, nothing usable. Carry none of these too, since every
call returns one of your labels; unrelated prose against five types picked it
at 0.94.
"tier": "smart", which re-runs answers under 0.7 on a reasoning model.confidence: null (no comparable score is available) — do
not extract; send the page to a person.fw9.pdf (6 pp) and fw4.pdf (5 pp) from irs.gov/pub/irs-pdf/, redacted,
both passes over 11 pages, 240 ms and 184 ms (types cut short):
fw9.pdf p1 W-9 1.00 | cover or transmittal page 0.37
fw9.pdf p2 W-9 1.00 | instructions or attachment page 0.98
fw4.pdf p1 W-4 1.00 | filled form page with input fields 0.49Every type came back at 0.98 or above. The roles are the interesting column: both first pages sit at 0.37 and 0.49, because a form page that is also the front of a document is genuinely both. Those two go to a person, found without reading the packet.
Every page was redacted first and carries a type, a role and two confidences; pages group into documents by runs of one type; the extractor sees only what is above 0.9, a person the rest.
© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/document-intake-routing of mrmps/classifier-dev.
Open the folder on GitHubat commit b9211dd
Document Intake Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Document Intake Routing this skillmrmps/classifier-dev | 424 | — | ~1.5k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Gzh Designisjiamu/gzh-design-skill | 3.9k | 1 repos | ~2.2k | Automated safety check: Pass | AGPL-3.0 | |
| GenOffice Document CLIgenspark-ai/genoffice | 8.8k | — | ~19k | Automated safety check: Pass | Apache-2.0 | |
| Harness Book Best Practicewquguru/harness-books | 3.2k | — | ~4.1k | Automated safety check: Pass | None | |
| Bookforge Korean Ebook PDF Makergongnyang/bookforge | 314 | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
isjiamu/gzh-design-skill
微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…
genspark-ai/genoffice
Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.
wquguru/harness-books
Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.
gongnyang/bookforge
Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
mrmps/classifier-dev
Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.
mrmps/classifier-dev
Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.
mrmps/classifier-dev
Check user-generated text against a written policy before it is published.
mrmps/classifier-dev
Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…
mrmps/classifier-dev
Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…
mrmps/classifier-dev
Type candidate (subject, sentence, object) triples against a fixed relation schema and flag triples that contradict each other, batched, with a calibrated confidence per edge so only confident edges…
Categories
Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person. Document Intake Routing is an agent skill from mrmps/classifier-dev. Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.
Document Intake Routing fits situations like: vendor upload arrives as one long PDF; route these scanned forms; which form type is each page; split this intake packet.
Run `npx skills add mrmps/classifier-dev --skill document-intake-routing -a claude-code`. Or copy the skill folder (skills/document-intake-routing in mrmps/classifier-dev) into .claude/skills/document-intake-routing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mrmps/classifier-dev --skill document-intake-routing -a codex`. Or copy the skill folder (skills/document-intake-routing in mrmps/classifier-dev) into .agents/skills/document-intake-routing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill document-intake-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-intake-routing, .gemini/skills/document-intake-routing, .github/skills/document-intake-routing and .opencode/skills/document-intake-routing in your project.
Going by SKILL.md and its folder, Document Intake Routing needs the command-line tools its instructions call (pdftotext). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Document Intake Routing is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Document Intake Routing: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.8k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 6, 2026.
Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.