Agent skill

RAG Reranker

by mrmps in mrmps/classifier-dev

Rerank retrieved passages by whether each one answers the question, drop the near-misses before they reach the prompt, and check every citation in the answer against the passage it points at.

MITAuto-check passedAI & LLM Engineering

Install RAG Reranker

skills CLI
$ npx skills add mrmps/classifier-dev --skill rag-reranker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mrmps/classifier-dev rag-reranker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-reranker .claude/skills/rag-reranker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-reranker
GitHub stars
424
Token cost
~1.5k tokens
SKILL.md length
616 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Rerank retrieved passages by whether each one answers the question, drop the near-misses before they reach the prompt, and check every citation in the answer against the passage it points at.

  • Works in 3 steps: one call over every candidate → rank, then cut → check the citations you generated
  • Which chunks answer this
  • SKILL.md covers When not to use it, Step 1: one call over every…, Step 2: rank, then cut and Measured, plus 3 more sections
  • Reaches classifier.dev

What it does

RAG Reranker is an agent skill from mrmps/classifier-dev. Rerank retrieved passages by whether each one answers the question, drop the near-misses before they reach the prompt, and check every citation in the answer against the passage it points at. Use after a vector search returns twenty plausible chunks, or when an answer needs its citations verified. Triggers on "rerank these", "which chunks answer this", "check these citations", "my RAG cites the wrong chunk".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Citation management, Retrieval-augmented generation and Vector databases. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.

When your agent uses it

  • Which chunks answer this
  • Check these citations
  • My RAG cites the wrong chunk

Example prompts

  • “rerank these”
  • “which chunks answer this”
  • “check these citations”
  • “/rag-reranker”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. one call over every candidate
  2. rank, then cut
  3. check the citations you generated

What it can do on your machine

Read from SKILL.md and the folder at commit 629df75. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • classifier.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Reranker loads about 1.5k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 616 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mrmps/classifier-dev at commit 629df75, republished under its MIT licence (© mrmps). 616 words, ~1,500 tokens.

Download SKILL.mdSave it as .claude/skills/rag-reranker/SKILL.md (or your agent's skills folder).
name
rag-reranker
description
Rerank retrieved passages by whether each one answers the question, drop the near-misses before they reach the prompt, and check every citation in the answer against the passage it points at. Use after a vector search returns twenty plausible chunks, or when an answer needs its citations verified. Triggers on "rerank these", "which chunks answer this", "check these citations", "my RAG cites the wrong chunk".
license
MIT

Rerank and verify with a yes/no call

Vector search returns what is near the question. Much of it is about the right subject and answers nothing, which costs context and invites citations pointing at passages that never made the claim.

classifier.dev answers one question per passage — does this answer it — with a calibrated confidence, keyless, in one batched call. It never writes text; the answer stays yours to write.

When not to use it

Not for first-stage retrieval (it does not embed or search), not for five passages already in context, and not for ranking a set that is entirely relevant — every passage gets a label, so a uniformly good set all comes back answers the question.

Step 1: one call over every candidate

python
import json, urllib.request

def rerank(question, passages):
    body = {"labels": ["answers the question",
                       "related but does not answer",
                       "unrelated"],
            "inputs": passages,                    # up to 1,000 per request
            "instructions": f"The question is: {question}"}
    req = urllib.request.Request(
        "https://classifier.dev/v1/classify", data=json.dumps(body).encode(),
        headers={"content-type": "application/json", "user-agent": "rag/1.0"})
    return json.load(urllib.request.urlopen(req))["results"]

The question goes in instructions, never in the labels: labels stay reusable, the question changes per call. Set a user-agent; Python's stdlib default is refused at the edge with 403.

Step 2: rank, then cut

Two signals come back per passage. Use both.

  • scores["answers the question"] is the ranking. Sort by it and find the cliff; retrieval sets usually have one.
  • label is the cut. Keep every passage labelled answers the question, drop the rest, but keep any reject whose confidence is under 0.5 as cheap insurance — a dropped passage is invisible to you afterwards.
  • confidence says how much to trust that cut: 0.9 and above act on it, 0.5 to 0.9 keep it but have your model check it, below 0.5 never let it be the only source of a claim.

Measured

20 passages retrieved for "Why was the Mars Climate Orbiter lost in 1999?" from three Wikipedia articles, one call: 253 ms wall, 177 ms server, sorted by the answers score:

0.99  answers the question         conf 0.99  Mars Climate Orbiter lead: names the mismatch
0.56  answers the question         conf 0.35  ... began the orbital insertion manoeuvre ...
0.08  related but does not answer  conf 0.86  The loss took place two and a half months ...
0.05  related but does not answer  conf 0.45  The cause of the communication loss is ...
0.01  unrelated                    conf 0.63  In 1832, Gauss used the astronomical second ...

Two passages cleared the cliff; 18 sat at 0.08 or below. Eight came back unrelated, six from the Metric_system article retrieval pulled in because the answer happens to be about units. The prompt went from 20 passages to 2, the one holding the answer at 0.99.

The second row is the lesson: a real answer passage at score 0.56 but confidence 0.35, because it says "complications arising from human error" without naming the cause. Keep it, rank it second, never let it be the only citation. A "tier": "smart" pass over the same 20 took 8.5 s, escalated 14 and moved no reject above 0.9 — the middle band is where honest uncertainty lives.

Show full SKILL.md (216 more words)Show less

Step 3: check the citations you generated

After your model writes the answer, pair each claim with the passage it cites and ask a different question — same API, one batch:

python
body = {"labels": ["the passage states this claim",
                   "the passage is about this but does not state it",
                   "the passage contradicts this claim"],
        "inputs": [f"CLAIM: {c}\n\nPASSAGE: {p}" for c, p in cited],
        "instructions": "Judge only whether the passage supports the claim. "
                        "Do not use outside knowledge."}

Four claims against their cited passages, 214 ms (claims abbreviated):

0.79  is about this but does not state it  "lost because ground software sent
                                            pound-force where newtons were expected"
0.39  contradicts this claim               "lost because a solar array failed to deploy"
1.00  states this claim                    "radio contact was lost 49 seconds early"
0.31  contradicts this claim               "the Polar Lander was lost the same way"

The first row is the useful one. The passage says "a measurement mismatch between SI units and US customary units"; the claim names pound-force and newton-seconds, detail the passage never gives. At 0.79 it is in the review band — right, for a claim that is true but not sourced by that passage. Cut the citation or cite something that states it.

Two things that will bite you

confidence is not scores[label]. One item scored 0.52 for its chosen label and came back at confidence: 0.04: the score is raw preference, the confidence is the calibrated probability it is right. Rank on the score, gate on the confidence.

It cannot tell you what a passage says. It picks among your labels. Anything that needs words — the answer, a summary, a citation rewrite — is your job.

Done looks like

Every passage carries a label, a score and a confidence; the prompt holds only the ones labelled as answering; every claim has a citation checked at 0.9 or above, and the rest are removed or flagged for a person.

© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/rag-reranker of mrmps/classifier-dev.

Open the folder on GitHubat commit 629df75

Compare with similar skills

RAG Reranker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Reranker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Reranker this skillmrmps/classifier-dev424—~1.5kAutomated safety check: PassMIT
Scholar RAGjoshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesCustom licence
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0
Convex Agentswaynesutton/builder-skills404—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 19 days ago
    Research & ScienceAuto-check: notes
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Convex Agents

    waynesutton/builder-skills

    Builds AI agents on the Convex agent component: threads, messages, tools that call queries and mutations, streaming, RAG with vector search, and workflows for multi step jobs.

    404 GitHub stars~2.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Sciverse

    opendatalab/Sciverse-Agent-Tools

    A skill your agent uses when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original…

    119 GitHub stars~3k tokensUpdated 18 days ago
    AI & LLM EngineeringAuto-check passed

More from mrmps/classifier-dev

All 21 skills in this repo
  • Bulk Classify

    mrmps/classifier-dev

    Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

    424 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Computer Use Action Picker

    mrmps/classifier-dev

    Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Content Moderation Gate

    mrmps/classifier-dev

    Check user-generated text against a written policy before it is published.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Document Intake Routing

    mrmps/classifier-dev

    Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Headline Filter Map Reduce

    mrmps/classifier-dev

    Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Questions about RAG Reranker

What does RAG Reranker do?

Rerank retrieved passages by whether each one answers the question, drop the near-misses before they reach the prompt, and check every citation in the answer against the passage it points at. RAG Reranker is an agent skill from mrmps/classifier-dev. Rerank retrieved passages by whether each one answers the question, drop the near-misses before they reach the prompt, and check every citation in the answer against the passage it points at.

When should I use RAG Reranker?

RAG Reranker fits situations like: which chunks answer this; check these citations; my RAG cites the wrong chunk.

How do I install RAG Reranker in Claude Code?

Run `npx skills add mrmps/classifier-dev --skill rag-reranker -a claude-code`. Or copy the skill folder (skills/rag-reranker in mrmps/classifier-dev) into .claude/skills/rag-reranker in your project. Claude Code loads it when a task matches its description.

How do I install RAG Reranker in Codex?

Run `npx skills add mrmps/classifier-dev --skill rag-reranker -a codex`. Or copy the skill folder (skills/rag-reranker in mrmps/classifier-dev) into .agents/skills/rag-reranker in your project. Codex loads it when a task matches its description.

Can I use RAG Reranker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill rag-reranker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-reranker, .gemini/skills/rag-reranker, .github/skills/rag-reranker and .opencode/skills/rag-reranker in your project.

What does RAG Reranker need to run?

SKILL.md names no scripts, command-line tools or credentials: RAG Reranker is instructions for the agent only. Our summary lists: Python 3.

Does RAG Reranker access the network?

SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is RAG Reranker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Reranker use?

RAG Reranker is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Reranker use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Reranker?

Skills that share tags, products or a category with RAG Reranker: Scholar RAG (joshzyj/open-scholar-skill, 168 stars), Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ms Agent Framework RAG (shuyu-labs/WebCode, 278 stars) and Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Reranker?

mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 7, 2026.

Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.