Agent skill

Document Intake Routing

by mrmps in mrmps/classifier-dev

Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

MITAuto-check passedDocuments & Office

Install Document Intake Routing

skills CLI
$ npx skills add mrmps/classifier-dev --skill document-intake-routing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mrmps/classifier-dev document-intake-routing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/document-intake-routing .claude/skills/document-intake-routing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
document-intake-routing
GitHub stars
424
Token cost
~1.5k tokens
SKILL.md length
573 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

  • Works in 3 steps: one text per page → two passes over the same pages → the gate
  • Vendor upload arrives as one long PDF
  • SKILL.md covers What leaves the machine, When not to use it, Step 1: one text per page and Step 2: two passes over the…, plus 3 more sections
  • Calls pdftotext; reaches classifier.dev

What it does

Document Intake Routing is an agent skill from mrmps/classifier-dev. Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person. Use when a mailroom, loan file, claim packet or vendor upload arrives as one long PDF. Triggers on "route these scanned forms", "which form type is each page", "split this intake packet", "sort these scans by form type".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.

When your agent uses it

  • Vendor upload arrives as one long PDF
  • Route these scanned forms
  • Which form type is each page
  • Split this intake packet

Example prompts

  • “route these scanned forms”
  • “which form type is each page”
  • “split this intake packet”
  • “/document-intake-routing”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. one text per page
  2. two passes over the same pages
  3. the gate

What it can do on your machine

Read from SKILL.md and the folder at commit b9211dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pdftotext

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • classifier.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Document Intake Routing loads about 1.5k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 573 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mrmps/classifier-dev at commit b9211dd, republished under its MIT licence (© mrmps). 573 words, ~1,497 tokens.

Download SKILL.mdSave it as .claude/skills/document-intake-routing/SKILL.md (or your agent's skills folder).
name
document-intake-routing
description
Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person. Use when a mailroom, loan file, claim packet or vendor upload arrives as one long PDF. Triggers on "route these scanned forms", "which form type is each page", "split this intake packet", "sort these scans by form type".
license
MIT

Route an intake packet page by page

A packet is one PDF holding several documents: a cover sheet, two forms, a pile of attachments, some blank backs. Extraction fails on it because it is pointed at the wrong page. Classify pages first, then extract what you are sure of. The classifier below is classifier.dev: keyless HTTP, a label per page, a calibrated confidence, no generated text.

What leaves the machine

Each request carries page text and your label names: no file, no image, no file name. The service states it stores no input text and forwards it to the model that answers (https://classifier.dev/privacy) — still a third party, and page one of a tax or loan packet is where the name, SSN and account number sit. Redact first, every time:

python
import re
PATTERNS = [(r"(?i)\b(?:bearer|basic)\s+[\w.\-+/=]{8,}", "<CRED>"),
    (r"(?i)\b[\w.-]*(?:key|token|secret|password|pwd)[\w.-]*\s*[=:]\s*[^\s\"',&]{6,}", "<CRED>"),
    (r"\b[A-Za-z0-9+/]{32,}={0,2}\b", "<BLOB>"), (r"\b[0-9a-f]{16,}\b", "<BLOB>"),
    (r"[\w.+-]+@[\w-]+\.[\w.]{2,}", "<EMAIL>"), (r"\b(?:\d[ -]?){13,16}\b", "<CARD>"),
    (r"\b\d{3}-\d{2}-\d{4}\b", "<SSN>"), (r"\b\d{7,}\b", "<NUM>")]

def redact(t):
    for p, tag in PATTERNS:
        t = re.sub(p, tag, t)
    return t

def sendable(t):                     # never send a mostly-redacted page
    kept = len(re.sub(r"<[A-Z]+>", "", redact(t)))
    return kept >= 25 and kept >= 0.5 * len(t)

On a filled page:

Name as shown on your income tax return: Jane Q. Public. Social security
number <SSN>. Account <CARD>. Email <EMAIL> W-9 Request for Taxpayer Form ...

That page classified the same either way: type 1.00 raw and 1.00 redacted, role 0.97 against 0.91. Redaction costs almost nothing. It catches shapes, not names: Jane Q. Public survives it.

When not to use it

If these packets are confidential under the policy you work under, do not call out at all. The workflow is a classifier plus two thresholds, and anything giving a calibrated confidence fits: your own model over the same labels, or a local zero-shot model. classifier.dev is the keyless example because it needs no setup. Same steps, other call, no network.

Skip it too for a one-document PDF you already know, for pulling fields out (extraction, not classification), and for image-only scans: OCR first, since a page with no text layer classifies as nothing.

Step 1: one text per page

pip install pdfplumber==0.11.4
python
import pdfplumber
with pdfplumber.open("packet.pdf") as pdf:
    raw = [" ".join((p.extract_text() or "").split()) for p in pdf.pages]
pages = [(i, redact(t[:1200])) for i, t in enumerate(raw, 1) if sendable(t)]

1,200 characters is enough. pdftotext -f N -l N packet.pdf - does the same in the shell. Empty pages drop out here: a whitespace-only input fails the whole batch with empty_input.

Show full SKILL.md (272 more words)Show less

Step 2: two passes over the same pages

One question per pass: a combined set ("W-9 instructions page") multiplies the labels and splits the probability mass.

python
import json, urllib.request
FORMS = ["IRS Form W-9 request for taxpayer identification number",
         "IRS Form W-4 employee withholding certificate", "commercial invoice",
         "bank statement", "none of these"]
ROLES = ["cover or transmittal page", "filled form page with input fields",
         "instructions or attachment page", "blank or nearly empty page"]

def classify(labels, instructions=""):
    body = {"labels": labels, "inputs": [t for _, t in pages],
            "instructions": instructions}
    req = urllib.request.Request("https://classifier.dev/v1/classify",
        json.dumps(body).encode(),
        {"content-type": "application/json", "user-agent": "intake/1.0"})
    return json.load(urllib.request.urlopen(req))["results"]

forms = classify(FORMS, "Identify this page's document type.")
roles = classify(ROLES)

One request per pass, up to 1,000 pages each; tens of types is normal, 100 is the ceiling. Set a user-agent: Python's default is blocked at the edge.

Name the document, not your code for it: the same 11 pages against DOC_TYPE_01 ... DOC_TYPE_06, OTHER all came back DOC_TYPE_01, at 0.47, 0.34, 0.30, 0.26 — one bucket, nothing usable. Carry none of these too, since every call returns one of your labels; unrelated prose against five types picked it at 0.94.

Step 3: the gate

  • 0.9 and above — route to the extractor for that type.
  • 0.5 to 0.9 — extract, but queue for review, or re-ask with "tier": "smart", which re-runs answers under 0.7 on a reasoning model.
  • below 0.5, or confidence: null (no comparable score is available) — do not extract; send the page to a person.

A worked run

fw9.pdf (6 pp) and fw4.pdf (5 pp) from irs.gov/pub/irs-pdf/, redacted, both passes over 11 pages, 240 ms and 184 ms (types cut short):

fw9.pdf p1  W-9 1.00 | cover or transmittal page          0.37
fw9.pdf p2  W-9 1.00 | instructions or attachment page    0.98
fw4.pdf p1  W-4 1.00 | filled form page with input fields 0.49

Every type came back at 0.98 or above. The roles are the interesting column: both first pages sit at 0.37 and 0.49, because a form page that is also the front of a document is genuinely both. Those two go to a person, found without reading the packet.

Done looks like

Every page was redacted first and carries a type, a role and two confidences; pages group into documents by runs of one type; the extractor sees only what is above 0.9, a person the rest.

© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/document-intake-routing of mrmps/classifier-dev.

Open the folder on GitHubat commit b9211dd

Compare with similar skills

Document Intake Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Document Intake Routing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Document Intake Routing this skillmrmps/classifier-dev424—~1.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice8.8k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3141 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.8k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    314 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from mrmps/classifier-dev

All 21 skills in this repo
  • Bulk Classify

    mrmps/classifier-dev

    Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

    424 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Computer Use Action Picker

    mrmps/classifier-dev

    Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Content Moderation Gate

    mrmps/classifier-dev

    Check user-generated text against a written policy before it is published.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Headline Filter Map Reduce

    mrmps/classifier-dev

    Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Type candidate (subject, sentence, object) triples against a fixed relation schema and flag triples that contradict each other, batched, with a calibrated confidence per edge so only confident edges…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Questions about Document Intake Routing

What does Document Intake Routing do?

Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person. Document Intake Routing is an agent skill from mrmps/classifier-dev. Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

When should I use Document Intake Routing?

Document Intake Routing fits situations like: vendor upload arrives as one long PDF; route these scanned forms; which form type is each page; split this intake packet.

How do I install Document Intake Routing in Claude Code?

Run `npx skills add mrmps/classifier-dev --skill document-intake-routing -a claude-code`. Or copy the skill folder (skills/document-intake-routing in mrmps/classifier-dev) into .claude/skills/document-intake-routing in your project. Claude Code loads it when a task matches its description.

How do I install Document Intake Routing in Codex?

Run `npx skills add mrmps/classifier-dev --skill document-intake-routing -a codex`. Or copy the skill folder (skills/document-intake-routing in mrmps/classifier-dev) into .agents/skills/document-intake-routing in your project. Codex loads it when a task matches its description.

Can I use Document Intake Routing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill document-intake-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-intake-routing, .gemini/skills/document-intake-routing, .github/skills/document-intake-routing and .opencode/skills/document-intake-routing in your project.

What does Document Intake Routing need to run?

Going by SKILL.md and its folder, Document Intake Routing needs the command-line tools its instructions call (pdftotext). Our summary lists: Python 3.

Does Document Intake Routing access the network?

SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Document Intake Routing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Document Intake Routing use?

Document Intake Routing is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Document Intake Routing use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Document Intake Routing?

Skills that share tags, products or a category with Document Intake Routing: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.8k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Document Intake Routing?

mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 6, 2026.

Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.