Agent skill

Agentsop Reranker Stage

by agentsope in agentsope/SkillAlchemy

Adds and tunes a reranker stage for RAG using the retrieve-wide, rerank-narrow pattern.

MITAuto-check passedAI & LLM Engineering

Install Agentsop Reranker Stage

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-reranker-stage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-reranker-stage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-reranker-stage .claude/skills/agentsop-reranker-stage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-reranker-stage
GitHub stars
436
Token cost
~4.9k tokens
SKILL.md length
2,321 words
Files
4 (incl. references)
Skills in repo
46
Repo updated
First seen
Licence
MIT

At a glance

Adds and tunes a reranker stage for RAG using the retrieve-wide, rerank-narrow pattern.

  • Works in 7 steps: 何时激活 (Activation Rules) → 核心心智模型 (Core Mental Model) → SOP 工作流 (Agentic Protocol) → …
  • Relevant documents appear in the initial top-N but are buried by noise
  • SKILL.md covers 1 · 何时激活 (Activation Rules), 2 · 核心心智模型 (Core Mental Model), 3 · SOP 工作流 (Agentic Protocol) and 4 · 操作模型 (Operation Models), plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentsop Reranker Stage is an agent skill from agentsope/SkillAlchemy. Adds and tunes a reranker stage for RAG using the retrieve-wide, rerank-narrow pattern. Use when relevant documents appear in the initial top-N but are buried by noise, top-1 precision or MRR is low despite adequate recall, or too many marginal chunks consume context. Covers cross-encoder, API, and local rerankers; N-to-k selection; and latency/cost tradeoffs. Do not use when retrieval recall itself is failing.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation. It works with LlamaIndex. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Relevant documents appear in the initial top-N but are buried by noise
  • Top-1 precision
  • MRR is low despite adequate recall
  • Too many marginal chunks consume context

Example prompts

  • “/agentsop-reranker-stage”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 何时激活 (Activation Rules)
  2. 核心心智模型 (Core Mental Model)
  3. SOP 工作流 (Agentic Protocol)
  4. 操作模型 (Operation Models)
  5. 困境决策案例 (Dilemma Cases)
  6. 反模式与边界 (Anti-patterns & Boundaries)
  7. 跨框架对照 (Cross-framework Mapping)

What it can do on your machine

Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop Reranker Stage loads about 4.9k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 2,321 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 2,321 words, ~4,856 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-reranker-stage/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
agentsop-reranker-stage
description
Adds and tunes a reranker stage for RAG using the retrieve-wide, rerank-narrow pattern. Use when relevant documents appear in the initial top-N but are buried by noise, top-1 precision or MRR is low despite adequate recall, or too many marginal chunks consume context. Covers cross-encoder, API, and local rerankers; N-to-k selection; and latency/cost tradeoffs. Do not use when retrieval recall itself is failing.
version
0.1.0
when_to_use
a RAG pipeline's answer quality has plateaued and the relevant doc is present in top-k but not ranked first, the LLM context window is pressured by too many…
when_not_to_use
retrieval RECALL is the problem (the right doc is NOT in top-N at all) — fix retrieval/chunking/hybrid first, top-k already small (<=5) and answers are…
overlay
true
cross_links
llamaindex, hybrid-retrieval

Reranker Stage · SOP

Third-person analytical view of how a mature RAG pipeline thinks about the reranker. The skill is for an LLM agent that writes / reviews / debugs retrieval code — it teaches the cross-framework reranking discipline, not one vendor's API. For the per-framework API, descend to [[llamaindex]] (node postprocessors) or [[agentsop-hybrid-retrieval]] (the recall stage that feeds the reranker).

This is the C4 gap skill in the Phase-D enhance pass. The reranker SOP existed only buried inside [[llamaindex]] (OP-03 AddReranker, Stage 3 step 7, anti-pattern A6). It is the highest-ROI single addition to a naive RAG pipeline, so it earns a standalone overlay.


1 · 何时激活 (Activation Rules)

Activate when any holds:

  1. A RAG pipeline's answer quality has plateaued after the cheap knobs (prompt, embedding model, chunk size) are exhausted — [[llamaindex]] Stage 3 lists reranking as the last optimization step, deliberately.
  2. Diagnostics show the relevant document is in top-k but buried — high hit-rate, low MRR, wrong top-1. This is LlamaIndex failure modes #1 / #10 ([[llamaindex]] OP-03).
  3. The LLM context window is under pressure — too many marginal chunks inflate cost, latency, and "lost-in-the-middle" degradation. A reranker lets you retrieve 50 and feed 5.
  4. A user asks where to add a reranker, how to tune N vs k, or API vs local.

Do not activate (boundary — see §6):

  • Recall is the bottleneck: the right doc is not in top-N at all. A reranker can only reorder what retrieval already found — fix retrieval, hybrid ([[agentsop-hybrid-retrieval]]), or chunking first.
  • top-k is already small (≤5) and answers are correct — no plateau.
  • A hard sub-100ms path where the extra round-trip is unaffordable and quality is already acceptable.

2 · 核心心智模型 (Core Mental Model)

The one sentence

Retrieve wide for recall with a cheap bi-encoder; rerank narrow for precision with an expensive cross-encoder that sees query + document together — something the bi-encoder structurally could not do.

Why two stages exist at all

The retriever (bi-encoder / vector search) embeds the query and every document separately, offline. Similarity is a dot product of two vectors that never met. This is fast (vectors are precomputed; ANN search is sub-linear) but lossy: the document's vector is a single "topic average" computed without knowledge of the query.

A cross-encoder takes [query, document] as a single joint input and runs full attention across both, emitting one relevance score. It sees exactly which query token matches which document token. This is far more accurate — and far more expensive: it cannot be precomputed, so it runs once per (query, candidate) pair at query time. Scoring 1M docs this way is infeasible; scoring 20-50 is cheap.

  query ─┐                               query ─┐
         ├─ dot product (precomputed)            ├─► [CROSS-ENCODER] ─► score
  doc  ─┘   ← bi-encoder, FAST, lossy     doc  ─┘   joint attention, SLOW, sharp
       RECALL stage (retrieve top-50)         PRECISION stage (rerank → top-5)

The reranker is the bridge: it spends cross-encoder accuracy on a small candidate set the bi-encoder produced cheaply. Wide net, sharp knife.

The order law (inherited from [[llamaindex]] Stage 3)

Prompts first, reranking last. Reranking is high-impact but expensive — exhaust the cheap knobs (prompt, embed model, chunk size, hybrid) before spending per-query cross-encoder latency. But once those are spent, the reranker is usually the single biggest remaining lever (5-15pp faithfulness lift on noisy corpora — [[llamaindex]] OP-03).

What a reranker is NOT
  • Not a recall fix — it reorders, never retrieves (§6).
  • Not a chunking fix — it scores whole candidates, it does not resize them.
  • Not free — every reranked candidate is an inference (local) or a billed unit (API).

3 · SOP 工作流 (Agentic Protocol)

Each stage gates the next. Never skip the baseline measurement.

Stage 0 — Confirm the lever is real

Before adding anything, prove the symptom is precision, not recall:

  1. Run the existing pipeline against ~30-50 labeled QA pairs.
  2. Record hit-rate@N (is the gold doc in top-N?) and MRR (how high?).
  3. If hit-rate is low → recall problem → STOP, fix retrieval / hybrid ([[agentsop-hybrid-retrieval]]) / chunking. A reranker will not help.
  4. If hit-rate is high but MRR is low → precision problem → the gold doc is buried → a reranker is the right lever. Proceed.
Stage 1 — Retrieve wide

Raise the retriever's top_k (or top_n) to 20-50. This is the "recall" stage: cast a wide net so the reranker has the gold doc to find. Hybrid retrieval ([[agentsop-hybrid-retrieval]]) feeds the reranker an even better candidate pool because it adds lexical recall the dense retriever misses.

Stage 2 — Insert the reranker

Add a reranker as a post-retrieval step (LlamaIndex node postprocessor; LangChain ContextualCompressionRetriever — §7). Pick the model per §4 OP-03. It consumes the wide candidate list and re-scores every candidate against the query with a cross-encoder.

Stage 3 — Keep narrow

Truncate to top-k = 3-5 after rerank. This is what reaches the synthesizer. The whole point: the LLM now sees a small, high-precision context instead of a large noisy one.

Stage 4 — Measure the lift, gate the change

Re-run the same eval set. Compare before vs after on {MRR, faithfulness, relevancy, p95 latency, per-query cost}. Keep the reranker only if the precision lift justifies the added latency/cost (§5). A reranker that adds 300ms for +1pp is not always worth shipping. Pin N and k as tuned constants.


4 · 操作模型 (Operation Models)

Each operation: Trigger / Action / Output / Evidence. Full machine-readable list in intermediate/operation_candidates.json.

OP-01 ConfirmPrecisionNotRecall
  • Trigger: Considering a reranker; haven't proven the symptom is precision.
  • Action: Measure hit-rate@N and MRR on labeled QA. High hit-rate + low MRR ⇒ precision problem ⇒ reranker is right. Low hit-rate ⇒ recall problem ⇒ STOP.
  • Output: Go/no-go decision backed by a number, not a hunch.
  • Evidence: [[llamaindex]] OP-03 (#1/#10 = right doc in top-k, wrong top-1); [[agentsop-hybrid-retrieval]] for the recall path.
OP-02 WidenThenNarrow
  • Trigger: Reranker confirmed; pipeline still on naive top_k=4.
  • Action: Retrieve top-N = 20-50, rerank, keep top-k = 3-5. Embed the two numbers as named, evaluated constants.
  • Output: A two-stage retrieve→rerank pipeline with a high-precision tail.
  • Evidence: [[llamaindex]] Stage 3 step 7 ("widen top_k to 20-50, rerank to 3-5"); OP-03.
OP-03 ChooseRerankerModel
  • Trigger: Reranker stage exists; model not yet chosen.
  • Action: Pick by constraint:
    • Cohere Rerank / Voyage rerank (API) — fastest to ship, no GPU, strong multilingual; cost per 1k searches, data leaves your boundary.
    • bge-reranker (bge-reranker-v2-m3 / large) — local — open-weights, self-hosted, no per-call fee, strong on multilingual; needs a GPU for low latency, you own ops.
    • SentenceTransformer cross-encoder (e.g. ms-marco-MiniLM) — light, CPU-runnable for small N, the lowest-dependency local option; weaker than bge-large but cheap.
    • ColBERT (late-interaction) — middle ground: token-level interaction, precomputable, scales to larger N than a full cross-encoder.
  • Output: A model justified by latency budget, cost ceiling, data-residency, and language mix.
  • Evidence: [[llamaindex]] OP-03 (CohereRerank / SentenceTransformerRerank / ColBERT named); external: "cohere rerank", "bge-reranker", "cross-encoder rerank RAG".
OP-04 BudgetLatencyAndCost
  • Trigger: Before shipping; reranker adds a per-query inference/billing unit.
  • Action: Measure p95 added by the rerank call at the chosen N. API rerankers add a network round-trip (~tens-hundreds ms) + per-search cost; local models add GPU/CPU inference time. Latency scales with N, not k — so over-large N is the latency killer (§6).
  • Output: p95 and $/query deltas attached to the change; ship only if the precision lift clears the bar.
  • Evidence: [[llamaindex]] Stage 3 ("high-impact but expensive — exhaust cheap knobs first"); §5 Dilemma 1.
OP-05 TuneNvsK
  • Trigger: Reranker live; N/k still at defaults; want to optimize the precision/latency frontier.
  • Action: Sweep N ∈ {20, 30, 50} holding k fixed (recall ceiling of the candidate pool), then sweep k ∈ {3, 5, 8} holding N fixed (how much context the LLM sees). Pick the smallest N that saturates hit-rate and the smallest k that saturates faithfulness.
  • Output: Tuned (N, k) on the cost/quality frontier, not guessed.
  • Evidence: [[llamaindex]] OP-02 TuneChunkSize (same sweep-and-pin discipline applied to N/k); Stage 3 step 7.
OP-06 RerankAfterHybrid
  • Trigger: Traffic has lexical-identity queries (codes, SKUs, symbols) AND a precision plateau.
  • Action: Use hybrid retrieval ([[agentsop-hybrid-retrieval]], BM25 + dense) for the wide stage, then rerank its fused candidate list. Hybrid maximizes recall into the pool; rerank maximizes precision out of it. They compose.
  • Output: Best-of-both — lexical recall + cross-encoder precision.
  • Evidence: [[agentsop-hybrid-retrieval]] (recall stage); [[llamaindex]] OP-04 AddHybridBM25 + OP-03 AddReranker (sequential in Stage 3).
OP-07 GateOnEval
  • Trigger: Any reranker add/change.
  • Action: Compare before/after on {MRR, faithfulness, relevancy, p95, $/query}. Keep only on net-positive. Treat as a regression test for future retriever changes.
  • Output: Quantitative justification; the reranker is now eval-gated.
  • Evidence: [[llamaindex]] OP-10 EvalLoop, Stage 2 ("eval loop before optimizing anything"), Stage 4.

5 · 困境决策案例 (Dilemma Cases)

Show full SKILL.md (976 more words)Show less
Dilemma 1 — Rerank latency vs answer quality

困境: A reranker reliably lifts precision but adds a per-query stage: network round-trip (API) or GPU inference (local). On a latency-sensitive surface (chat, autocomplete) the added p95 may violate the SLA even when quality improves.

约束: Cross-encoder cost is per (query, candidate) pair and scales with N ([[llamaindex]] Stage 3: reranking is "high-impact but expensive"). Latency is dominated by N, not k. The bi-encoder stage was chosen precisely because it is fast; the reranker reintroduces query-time compute.

决策步骤:

  1. Measure baseline p95 and the SLA headroom.
  2. Measure rerank-stage p95 at the smallest viable N (start N=20).
  3. If it fits headroom and quality lifts ≥ a meaningful threshold → ship.
  4. If it does not fit → shrink N (OP-05), switch to a lighter model (MiniLM cross-encoder, ColBERT), or rerank async/cache for repeat queries.
  5. If still over budget and quality is already acceptable → do not rerank (§6 boundary).

结果: Reranking is the highest-ROI lever only when latency headroom exists. The decision is SLA-driven, not quality-driven in isolation. Smaller N often recovers most of the lift at a fraction of the latency.

可提取的操作: OP-04, OP-05. Anti-pattern A3 (over-large N).

Dilemma 2 — API reranker (Cohere/Voyage) vs local model (bge / cross-encoder)

困境: The hosted API ships in an afternoon, needs no GPU, and tracks SOTA — but bills per search and sends query + candidates to a third party. A local bge-reranker has zero per-call fee and keeps data in-boundary — but needs a GPU, ops ownership, and model-update discipline.

约束: Per-query cost (API) vs fixed infra cost + ops (local); data-residency / compliance; latency (API adds network hop, local adds inference); team's GPU/MLOps capacity.

决策步骤:

  1. Data residency hard constraint (PII, regulated)? → local (bge / ColBERT), decision over.
  2. Estimate query volume × API price vs GPU rental. Low/spiky volume → API usually cheaper; high steady volume → local amortizes.
  3. No GPU and no MLOps appetite? → API (or CPU MiniLM for tiny N).
  4. Multilingual corpus? Both Cohere and bge-reranker-v2-m3 are strong — let cost/residency decide.
  5. Whichever: wrap the call behind a single rerank(query, nodes) -> nodes seam so swapping API↔local is a one-line change.

结果: Default to the API to validate the lift cheaply (prove the reranker helps before investing in infra), then migrate to local once volume, cost, or residency justify it. The abstraction seam makes the migration safe.

可提取的操作: OP-03, OP-04. Anti-pattern A5 (vendor lock-in, no seam).


6 · 反模式与边界 (Anti-patterns & Boundaries)

Anti-patterns
#Anti-patternCorrect move
A1Reranking to fix recall — gold doc isn't in top-NFix retrieval / hybrid ([[agentsop-hybrid-retrieval]]) / chunking; a reranker only reorders what's already retrieved
A2Naive similarity_top_k=N then feed all N to the LLM, no rerankWiden N and rerank to top-3-5 ([[llamaindex]] A6)
A3Over-large N (rerank 200+ candidates)Latency scales with N; pick the smallest N that saturates hit-rate (OP-05)
A4Add reranker first, before prompt/embed/chunk/hybridOrder law: reranking is last ([[llamaindex]] Stage 3); cheapest knobs first
A5Hard-wire one vendor SDK throughout the pipelineHide behind a rerank(query, nodes) seam so API↔local swaps in one line (Dilemma 2)
A6Ship reranker without before/after evalGate on {MRR, faithfulness, p95, $/query} (OP-07); a reranker that costs latency for no lift is removed
A7Keep N=k (rerank n candidates, return n)Reranking only helps when k < N — you must discard the low-scored tail
A8Re-embed / re-chunk hoping to fix "wrong top-1"If the right doc is present but buried, that's a rerank job, not a re-ingest
Boundaries — when not to add a reranker
  • B1 — Recall is the bottleneck: hit-rate@N low ⇒ the answer isn't in the pool. Reranking is a no-op. Fix retrieval first (OP-01, [[agentsop-hybrid-retrieval]]).
  • B2 — Already narrow & correct: top-k ≤ 5 and answers right ⇒ no plateau, no lever.
  • B3 — Hard real-time / sub-100ms: the extra round-trip blows the budget and quality is acceptable ⇒ skip (Dilemma 1).
  • B4 — Tiny static corpus (prompt-stuffable, <100k tokens): no retrieval stage to rerank ([[llamaindex]] B1).
PR-review smells (instant red flags)
  • index.as_query_engine(similarity_top_k=20) with no node postprocessor → A2 ([[llamaindex]] PR-smell).
  • Reranker added but top_k still 4 → A7 (N=k, reranker is a no-op).
  • A vendor rerank SDK imported in >1 module → A5 (no seam).
  • A reranker PR with no eval delta in the description → A6.
  • "Added reranker to improve recall" in a commit message → A1 (category error).

7 · 跨框架对照 (Cross-framework Mapping)

The reranker is one stage with the same shape everywhere: consume a wide candidate list, re-score with a cross-encoder, truncate to top-k.

Framework / vendorReranker primitiveNotes
LlamaIndex ([[llamaindex]])Node postprocessor: CohereRerank, SentenceTransformerRerank, ColbertRerank, LLMRerank passed as node_postprocessors=[...] to the query engine; widen similarity_top_k, set top_n on the rerankerThe canonical reference; OP-03 AddReranker, Stage 3 step 7, A6
LangChainContextualCompressionRetriever wrapping a base retriever with a CohereRerank / CrossEncoderReranker / LLMChainExtractor compressorBase retriever returns N, compressor reranks/filters to k
Cohere Rerank APIcohere.rerank(query, documents, top_n, model="rerank-v3.5")Hosted cross-encoder; multilingual; per-search billing
Voyage rerank APIvoyageai.rerank(query, documents, model="rerank-2", top_k)Hosted; pairs well with Voyage embeddings
bge-reranker (local)FlagReranker("BAAI/bge-reranker-v2-m3") / via sentence-transformers CrossEncoderOpen-weights, self-hosted, no per-call fee, GPU recommended
SentenceTransformers cross-encoderCrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2").predict([(q, d), ...])Lightest local option; CPU-viable for small N
ColBERT / RAGatouilleLate-interaction reranker; token-level scoring, precomputableScales to larger N than a full cross-encoder
HaystackTransformersSimilarityRanker / CohereRanker component in the pipelineSame wide→narrow shape, pipeline-component form

Activate this skill for the reranking decision (whether, where, how wide, which model, what it costs). Descend to [[llamaindex]] for node-postprocessor wiring, and to [[agentsop-hybrid-retrieval]] for the recall stage that feeds it.


References

  • references/R1-source-evidence.md — every cited claim resolved to a source line.
  • intermediate/operation_candidates.json — machine-readable operation list.
Primary sources (cited inline above)
  • [[llamaindex]] SKILL — OP-03 AddReranker, OP-02 TuneChunkSize, OP-04 AddHybridBM25, OP-10 EvalLoop; Stage 2/3/4; anti-patterns A6/A3; failure modes #1/#10; the "prompts first, reranking last" order law.
  • [[agentsop-hybrid-retrieval]] — the wide/recall stage (BM25 + dense) that feeds the reranker; lexical-identity recall.
  • External: "cohere rerank" (rerank-v3.5, hosted cross-encoder, per-search billing, multilingual); "bge-reranker" (BAAI bge-reranker-v2-m3 / large, open-weights local cross-encoder); "cross-encoder rerank RAG" (bi-encoder retrieve → cross-encoder rerank, joint query+doc attention, the two-stage recall→precision pattern).

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/agentsop-reranker-stage of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-source-evidence.md

Open the folder on GitHubat commit d0f0355

Compare with similar skills

Agentsop Reranker Stage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop Reranker Stage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop Reranker Stage this skillagentsope/SkillAlchemy436—~4.9kAutomated safety check: PassMIT
LlamaindexOrchestra-Research/AI-Research-SKILLs13k2 repos~3.7kAutomated safety check: PassMIT
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k2 repos~1.6kAutomated safety check: PassMIT
Llamaindexmagnus919/agent-skills115—~3kAutomated safety check: PassMIT
FAISS Similarity SearchOrchestra-Research/AI-Research-SKILLs13k6 repos~1.3kAutomated safety check: PassMIT
Pinecone Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k5 repos~2kAutomated safety check: PassMIT

Similar skills

  • Llamaindex

    Orchestra-Research/AI-Research-SKILLs

    Data framework for building LLM applications with RAG. An agent skill from Orchestra-Research/AI-Research-SKILLs.

    13k GitHub starsUsed in 2 repos~3.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Llamaindex

    magnus919/agent-skills

    Build LLM applications with the LlamaIndex framework. An agent skill from magnus919/agent-skills.

    115 GitHub stars~3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 6 repos~1.3k tokens
    DatabasesAuto-check passed
  • Pinecone Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to use Pinecone, a managed vector database, for production RAG, semantic search and recommendations: indexes, upserts, queries, filters and namespaces.

    13k GitHub starsUsed in 5 repos~2k tokens
    DatabasesAuto-check passed
  • RAG Skills

    llama-farm/llamafarm

    RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers.

    836 GitHub stars~1.3k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from agentsope/SkillAlchemy

All 46 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    436 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    436 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    436 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    436 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    436 GitHub stars~5.4k tokensUpdated yesterday
    Auto-check: notes
  • Agentsop Multiscale Chunking

    agentsope/SkillAlchemy

    Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

    436 GitHub stars~4.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Agentsop Reranker Stage

What does Agentsop Reranker Stage do?

Adds and tunes a reranker stage for RAG using the retrieve-wide, rerank-narrow pattern. Agentsop Reranker Stage is an agent skill from agentsope/SkillAlchemy. Adds and tunes a reranker stage for RAG using the retrieve-wide, rerank-narrow pattern.

When should I use Agentsop Reranker Stage?

Agentsop Reranker Stage fits situations like: relevant documents appear in the initial top-N but are buried by noise; top-1 precision; MRR is low despite adequate recall; too many marginal chunks consume context.

How do I install Agentsop Reranker Stage in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-reranker-stage -a claude-code`. Or copy the skill folder (skills/agentsop-reranker-stage in agentsope/SkillAlchemy) into .claude/skills/agentsop-reranker-stage in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop Reranker Stage in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-reranker-stage -a codex`. Or copy the skill folder (skills/agentsop-reranker-stage in agentsope/SkillAlchemy) into .agents/skills/agentsop-reranker-stage in your project. Codex loads it when a task matches its description.

Can I use Agentsop Reranker Stage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-reranker-stage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-reranker-stage, .gemini/skills/agentsop-reranker-stage, .github/skills/agentsop-reranker-stage and .opencode/skills/agentsop-reranker-stage in your project.

What does Agentsop Reranker Stage need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentsop Reranker Stage is instructions for the agent only.

Does Agentsop Reranker Stage access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentsop Reranker Stage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop Reranker Stage use?

Agentsop Reranker Stage is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop Reranker Stage use?

About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop Reranker Stage?

Skills that share tags, products or a category with Agentsop Reranker Stage: Llamaindex (Orchestra-Research/AI-Research-SKILLs, 13k stars), Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), Llamaindex (magnus919/agent-skills, 115 stars) and FAISS Similarity Search (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop Reranker Stage?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.