Agent skill

Bulk Classify

by mrmps in mrmps/classifier-dev

Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

MITAuto-check passedBackend & APIs

Install Bulk Classify

skills CLI
$ npx skills add mrmps/classifier-dev --skill bulk-classify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mrmps/classifier-dev bulk-classify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src .claude/skills/bulk-classify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bulk-classify
GitHub stars
424
Token cost
~3.1k tokens
SKILL.md length
1,489 words
Files
230
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

  • Bucketing more items than are worth putting in context — search results before you read them
  • SKILL.md covers When this is worth a network…, Quickstart, From a shell and Parameters, plus 7 more sections
  • Runs TypeScript scripts from its folder; calls npm and jq; reaches classifier.dev
  • Past conversations

What it does

Bulk Classify is an agent skill from mrmps/classifier-dev. Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer. Use when triaging, filtering, routing or bucketing more items than are worth putting in context — search results before you read them, log lines, tickets, files, diffs, past conversations. Triggers on "filter these", "which of these are relevant", "triage", "bucket", "route", "categorise", or any loop that would otherwise read N items to keep a few.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 232 other files (for example `account-openapi.ts`, `admin.ts` and `admission.ts`).

It sits in Backend & APIs, covering REST APIs. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.

When your agent uses it

  • Bucketing more items than are worth putting in context — search results before you read them
  • Past conversations
  • Which of these are relevant
  • Any loop that would otherwise read N items to keep a few

Example prompts

  • “filter these”
  • “which of these are relevant”
  • “triage”
  • “/bulk-classify”

Requirements

  • Python 3
  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit b9211dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • npm
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • classifier.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bulk Classify loads about 3.1k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,489 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mrmps/classifier-dev at commit b9211dd, republished under its MIT licence (© mrmps). 1,489 words, ~3,069 tokens.

Download SKILL.mdSave it as .claude/skills/bulk-classify/SKILL.md (or your agent's skills folder). This skill also uses 229 other files; get the full folder from GitHub.
name
bulk-classify
description
Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer. Use when triaging, filtering, routing or bucketing more items than are worth putting in context — search results before you read them, log lines, tickets, files, diffs, past conversations. Triggers on "filter these", "which of these are relevant", "triage", "bucket", "route", "categorise", or any loop that would otherwise read N items to keep a few.
license
MIT

Classify at scale without reading

classifier.dev assigns text to your categories. No key, no signup, no SDK. One HTTP call takes up to a thousand texts at a time and comes back in about a second, each with a confidence you can act on.

When this is worth a network call

You are a language model. You can already classify any text you can see, for free. The question is whether you want this text in your context at all.

Reach for this when reading the input is the expensive part:

  • Filtering before reading. You have 40 search snippets and want the 6 worth opening. Classifying them yourself means pulling all 40 into context first, which is the cost you were trying to avoid. One call returns 40 labels and you read only the survivors.
  • Cascade pre-filter. Cheaply drop the obvious no's, then spend real reasoning on what is left.
  • Streams you would never read line by line. Log lines, error buckets, inbound tickets, changed files in a large diff, ten thousand URLs' titles.
  • Label-based routing. Use the fast tier to route texts into your categories. Model updates, fallback, and smart reasoning can change answers; neither tier guarantees identical results across calls.

Do not bother when you have a handful of items already in context, or the judgement needs reasoning about things the text does not state. Under about five items you have already paid the context cost, so just decide yourself.

Quickstart

One text, bare label back:

curl "https://classifier.dev/relevant,not+relevant/Redis+beats+Postgres+for+queues"
relevant

The same call as query parameters, when code is building the URL:

curl "https://classifier.dev/?labels=relevant,not+relevant&text=Redis+beats+Postgres+for+queues"
relevant

Many texts in one call. This is the path that matters:

curl https://classifier.dev -d '{
  "labels": ["relevant", "not relevant"],
  "inputs": ["first snippet", "second snippet", "third snippet"]
}'

Returns results in input order. Single-label results have {label, confidence, scores}. Confidence and scores can be null; check before comparing a threshold. Multi-label results use {labels: [...], scores: {...}} instead. Up to 1,000 texts per call; 400 news headlines measured at 650ms end to end. For more, fan out calls in parallel; the limit is 3,000 classifications a minute. Each result also names the model that answered it. At the batch level, modelsUsed lists every serving model and model is mixed when more than one model answered the batch.

From a shell

When the text is already in files, or the answer feeds another command, the CLI saves you writing the batching and the JSON:

npm i -g classifier-dev

classify bug,feature,praise < feedback.txt          # label<TAB>confidence<TAB>text, input order
classify relevant,"not relevant" --review 0.7 < snippets.txt   # only the unsure ones
classify db,web,ml --count < titles.txt             # a histogram instead of rows
classify a,b --json < items.txt | jq -c 'select(.confidence == null or .confidence < 0.8)'

It batches a thousand inputs per request, four requests at a time, and streams rows as they land, so | head on a large file returns at once. Retries 429 and 5xx with backoff. --help has the rest.

Reach for the HTTP API instead when the text is already in memory, when you need the full score map per item, or when you are inside a language runtime where one fetch is simpler than a subprocess.

Parameters

FieldNotes
labels2–100 categories. Required.
inputOne text; above 32,000 characters, default/explicit Jev uses paid Fast-only long context.
inputsUp to 1,000 texts in one call.
tierfast (default) or smart: re-asks low-confidence answers of a reasoning model.
instructionsExtra criteria — "judge only the service, ignore the food".
multiReturn every label that applies, with a score per label.
max_labelsCap on how many multi-label answers come back.
verbose=1On GET, returns JSON instead of a bare label.
textOn GET, the text as a query parameter: /?labels=a,b&text=.... input and q work too; classes and categories for labels.

Long context requires a workspace key backed by paid balance or an active paid subscription; signup credit and anonymous access do not qualify. Use POST with at most 250,000 original cl100k_base context tokens across inputs, 20 documents, 32 decisions and a 1 MB body. Price: $0.084/M original context tokens counted once across inputs, independent of dimensions and actual screening/final usage. Final Jev reads selected whole chunks in source order; eligible evidence may be omitted when the budget fills. usage.long_context discloses selection. No evidence returns 422 long_context_no_evidence without charge. Explicit model: "chunklaya" remains a separate legacy opt-in.

On GET every option goes in the query string, whichever form carries the labels and text; the two forms mix (/a,b?text=...). If a GET is malformed the error comes with usage: and try: — try is a URL built from what you sent that would have worked. Follow it rather than re-reading the docs.

Labels are read semantically, so name them in words: urgent bug classifies better than p0.

Confidence you can act on

The model is a decision model, not an LLM prompted to classify: it returns a calibrated probability for every label. Measured on 400 six-way emotion items, answers at confidence ≥ 0.9 were right 82% of the time; answers below 0.5 were right 29% of the time. So:

python
for text, r in zip(texts, results):
    if r["confidence"] is not None and r["confidence"] >= 0.8:
        act(r["label"])
    else:
        look_yourself(text)      # or send it through tier "smart"

tier: "smart" does that routing server-side: every single-label answer under 0.7 confidence is re-asked of a fast reasoning model and replaced, marked escalated: true, with usage.escalated telling you how many. Measured: four-way news 87.5% → 90.0% by re-asking 12% of items. It costs a few seconds per escalated item, so a batch on smart is slower in proportion to how uncertain it is. Escalated answers have confidence: null, scores: null, and unscored: the reasoning model does not return comparable probabilities. These answers belong in review when your workflow requires a confidence gate.

Show full SKILL.md (570 more words)Show less

Many labels at once

To tag instead of sorting (an article against fifty topics, a ticket against every subsystem it touches), ask for every label that applies:

curl https://classifier.dev -d '{
  "input": "...",
  "labels": ["machine learning", "databases", "... up to 100 ..."],
  "multi": true,
  "max_labels": 10
}'

Results carry labels (an array, most likely first) plus scores, one probability per label. POST multi-label results omit the singular label and confidence keys. Labels at or above 0.7 are returned; use scores to pick your own threshold. On GET, add ?multi=1 and they come back one per line. Measured F1 0.887 on a seven-task set with recall 0.99, in ~200ms. The tier makes no difference here, so leave it on fast.

Two things that will bite you

1. Every call returns one of your labels, always. There is no "none of the above" unless you supply one. Text that fits nothing still gets confidently sorted into your best-matching category: "the weather is nice today" against bug / feature / praise must land in one of those categories. If "none of these" is a real outcome, add it as a label. Hoping for a low score does not create a missing category.

2. Confidence predicts accuracy, not fit. It tells you how likely the chosen label is right among your labels, which is exactly what you want for routing. It does not tell you whether the text belongs to any of them; see point 1. Scores express the model's choice among the labels you supplied. They do not validate the input or prove the choice is correct, so supply labels suitable for every kind of input your caller may send. Confidence and scores can be null when the provider returns none or the smart tier replaces the scored answer.

Recipe: filter search results before reading them

python
import json, urllib.request

def keep_relevant(question, snippets):
    body = json.dumps({
        "labels": ["relevant", "not relevant"],
        "inputs": snippets,                      # up to 1,000
        "instructions": (
            f"Relevant means it helps answer: {question}. "
            "Include background and contrasting alternatives."
        ),
    }).encode()
    req = urllib.request.Request(
        "https://classifier.dev",
        data=body,
        headers={
            "content-type": "application/json",
            "user-agent": "my-agent/1.0",
        },
    )
    results = json.load(urllib.request.urlopen(req))["results"]
    # A dropped item is invisible, so keep anything the model was unsure about.
    return [s for s, r in zip(snippets, results)
            if r["label"] == "relevant" or r["confidence"] is None
            or r["confidence"] < 0.8]

Then read only what comes back. The snippets you dropped never enter context.

Bias a filter toward keeping. You never learn what you lost, so recall matters more than precision here. The confidence gate above does that directly; "When in doubt, keep it" in the instructions also measurably helps.

Python's standard urllib, curl and Node fetch work without a custom User-Agent. A descriptive agent name is optional. For a JSON error, read code, action and retryable: a 403 can mean the free service detected an anonymous proxy network, which requires a funded workspace key. A 429 carries Retry-After. An HTML error is an edge/network failure; report its status and request ID rather than assuming classification ran.

Report a problem with classifier.dev

If classifier.dev itself returns a wrong result shape, contradicts its docs, or creates repeated integration friction, report that to the service instead of classifying the report as input. Read the live policy first:

GET https://classifier.dev/.well-known/agent-feedback.json

For a short report, send one category and one useful sentence. No key is needed:

curl https://classifier.dev/api/v1/observations -d '{
  "category": "docs_mismatch",
  "summary": "The documented response field was absent from POST /v1/classify.",
  "surface": "/v1/classify"
}'

Use POST /api/v1/feedback when you have reproduction steps or other evidence. It accepts the envelope and limits described by the discovery document. Both routes return a receipt; poll GET /api/v1/receipts/{id} to confirm it landed. Never include credentials, private input text, or unrelated user data in a report or its evidence.

Limits

Per IP per minute: 3,000 classifications on fast, 200 on smart; per day 20,000 and 2,000. A batch of 400 counts as 400. 429 when exceeded, with RateLimit-Limit on every response. Errors are JSON on POST, {"error": "...", "code": "..."}, and plain text on GET unless you add ?verbose=1 or send Accept: application/json.

Reference

  • GET / — full docs, plain text
  • GET /openapi.json — OpenAPI 3.1
  • GET /benchmark — measured accuracy, calibration, cost and latency

© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 229 other files in src of mrmps/classifier-dev.

  • SKILL.md
  • account-openapi.ts
  • admin.ts
  • admission.ts
  • agents.ts
  • alerts.ts
  • billing-identity.ts
  • brand.ts
  • chat-analytics.ts
  • chat.ts
  • chatui.ts
  • classification-usage.ts
  • components/account-menu.tsx
  • components/app-shell.tsx
  • components/credit-balance.tsx
  • components/evilcharts/LICENSE
  • components/evilcharts/README.md
  • components/evilcharts/charts/recharts-area-chart.tsx
  • … and 212 more

Open the folder on GitHubat commit b9211dd

Compare with similar skills

Bulk Classify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bulk Classify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bulk Classify this skillmrmps/classifier-dev424—~3.1kAutomated safety check: PassMIT
API DesignerJeffallan/claude-skills12k2 repos~2kAutomated safety check: PassMIT
Paperclippaperclipai/paperclip98k—~9.6kAutomated safety check: PassMIT
Nodejs Backend Patternsever-works/ever-works15817 repos~4kAutomated safety check: PassAGPL-3.0
OpenAPI to MCP Servermcp-use/mcp-use11k—~5.2kAutomated safety check: PassApache-2.0
Use Yaakmountain-loop/yaak19k—~1.9kAutomated safety check: PassMIT

Similar skills

  • API Designer

    Jeffallan/claude-skills

    Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.

    12k GitHub starsUsed in 2 repos~2k tokens
    Backend & APIsAuto-check passed
  • Paperclip

    paperclipai/paperclip

    Interact with the Paperclip control plane API for task coordination and governance.

    98k GitHub stars~9.6k tokensUpdated today
    Backend & APIsAuto-check passed
  • Nodejs Backend Patterns

    ever-works/ever-works

    Build production-ready Node.js backend services with Express/Fastify, implementing middleware patterns, error handling, authentication, database integration, and API design best practices.

    158 GitHub starsUsed in 17 repos~4k tokens
    Backend & APIsAuto-check passed
  • OpenAPI to MCP Server

    mcp-use/mcp-use

    Turns an OpenAPI or Swagger spec into an MCP server with the mcp-use TypeScript SDK, mapping each operation to a tool, wiring auth, testing and deploying.

    11k GitHub stars~5.2k tokensUpdated today
    Backend & APIsAuto-check passed
  • Use Yaak

    mountain-loop/yaak

    A skill your agent uses when the user mentions Yaak, a Yaak workspace, or the yaak command, or asks to call, hit, or smoke test HTTP/REST endpoints, save or organize API requests for reuse or manual…

    19k GitHub stars~1.9k tokensUpdated today
    Backend & APIsAuto-check passed
  • Covers the RuView `wifi-densepose` command line binary, its Axum REST API and the WebAssembly builds for browsers and ESP32, for embedding or scripting RuView.

    97k GitHub stars~1.2k tokensUpdated today
    Backend & APIsAuto-check: notes

More from mrmps/classifier-dev

All 21 skills in this repo
  • Computer Use Action Picker

    mrmps/classifier-dev

    Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Content Moderation Gate

    mrmps/classifier-dev

    Check user-generated text against a written policy before it is published.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Document Intake Routing

    mrmps/classifier-dev

    Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Headline Filter Map Reduce

    mrmps/classifier-dev

    Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Type candidate (subject, sentence, object) triples against a fixed relation schema and flag triples that contradict each other, batched, with a calibrated confidence per edge so only confident edges…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Bulk Classify

What does Bulk Classify do?

Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer. Bulk Classify is an agent skill from mrmps/classifier-dev. Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

When should I use Bulk Classify?

Bulk Classify fits situations like: bucketing more items than are worth putting in context — search results before you read them; past conversations; which of these are relevant; any loop that would otherwise read N items to keep a few.

How do I install Bulk Classify in Claude Code?

Run `npx skills add mrmps/classifier-dev --skill bulk-classify -a claude-code`. Or copy the skill folder (src in mrmps/classifier-dev) into .claude/skills/bulk-classify in your project. Claude Code loads it when a task matches its description.

How do I install Bulk Classify in Codex?

Run `npx skills add mrmps/classifier-dev --skill bulk-classify -a codex`. Or copy the skill folder (src in mrmps/classifier-dev) into .agents/skills/bulk-classify in your project. Codex loads it when a task matches its description.

Can I use Bulk Classify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill bulk-classify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bulk-classify, .gemini/skills/bulk-classify, .github/skills/bulk-classify and .opencode/skills/bulk-classify in your project.

What does Bulk Classify need to run?

Going by SKILL.md and its folder, Bulk Classify needs TypeScript for the scripts in its folder and the command-line tools its instructions call (npm and jq). Our summary lists: Python 3; Node.js.

Does Bulk Classify access the network?

SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bulk Classify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bulk Classify use?

Bulk Classify is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bulk Classify use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bulk Classify?

Skills that share tags, products or a category with Bulk Classify: API Designer (Jeffallan/claude-skills, 12k stars), Paperclip (paperclipai/paperclip, 98k stars), Nodejs Backend Patterns (ever-works/ever-works, 158 stars) and OpenAPI to MCP Server (mcp-use/mcp-use, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bulk Classify?

mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 6, 2026.

Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.