Install the "agentsop-hybrid-retrieval" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-hybrid-retrieval into .claude/skills/agentsop-hybrid-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-hybrid-retrieval", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "agentsop-hybrid-retrieval" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-hybrid-retrieval into .agents/skills/agentsop-hybrid-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-hybrid-retrieval", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "agentsop-hybrid-retrieval" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-hybrid-retrieval into .cursor/skills/agentsop-hybrid-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-hybrid-retrieval", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "agentsop-hybrid-retrieval" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-hybrid-retrieval into .gemini/skills/agentsop-hybrid-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-hybrid-retrieval", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "agentsop-hybrid-retrieval" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-hybrid-retrieval into .github/skills/agentsop-hybrid-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-hybrid-retrieval", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "agentsop-hybrid-retrieval" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-hybrid-retrieval into .opencode/skills/agentsop-hybrid-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-hybrid-retrieval", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Works in 7 steps: 何时激活 (Activation Rules) → 核心心智模型 (Core Mental Model) → SOP 工作流 (Agentic Protocol) → …
Tasks that involve Operations and SOPs
SKILL.md covers 1. 何时激活 (Activation Rules), 2. 核心心智模型 (Core Mental Model), 3. SOP 工作流 (Agentic Protocol) and 4. 操作模型 (Operation Models), plus 3 more sections
Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
What it does
Agentsop Hybrid Retrieval is an agent skill from agentsope/SkillAlchemy. Enhancement-overlay SOP for adding sparse (BM25 / keyword) retrieval alongside dense (embedding) retrieval. Activate when a calling agent is building, reviewing, or debugging a retrieval pipeline whose corpus contains exact-match tokens — identifiers, error codes, SKUs, API/function names, proper nouns, citations, rare jargon — that pure dense embedding silently misses. Encodes the single decision rule (hybrid is traffic-driven, not theoretical: add sparse only when the query share that depends on exact tokens is…
Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in AI & LLM Engineering, covering Operations and SOPs, Embeddings and Retrieval-augmented generation. It works with LlamaIndex. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
When your agent uses it
Tasks that involve Operations and SOPs
Tasks that involve Embeddings
Tasks that involve Retrieval-augmented generation
Example prompts
“add keyword search for completeness”
“/agentsop-hybrid-retrieval”
Requirements
Python 3
Workflow steps
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Network
Links to these hosts (documentation or services it may open):
developers.llamaindex.ai
llamaindex.ai
tianpan.co
python.langchain.com
qdrant.tech
docs.weaviate.io
docs.pinecone.io
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Agentsop Hybrid Retrieval loads about 6.5k tokens when it runs, and up to ~7.9k if it reads all its reference files. Until then it costs about 204 tokens; SKILL.md has 2,845 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~204
When it runs· the whole SKILL.md, loaded when a task matches
~6.5k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~7.9k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/agentsop-hybrid-retrieval/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
agentsop-hybrid-retrieval
description
Enhancement-overlay SOP for adding sparse (BM25 / keyword) retrieval alongside
dense (embedding) retrieval. Activate when a calling agent is building, reviewing,
or debugging a retrieval pipeline whose corpus contains exact-match tokens —
identifiers, error codes, SKUs, API/function names, proper nouns, citations, rare
jargon — that pure dense embedding silently misses. Encodes the single decision
rule (**hybrid is traffic-driven, not theoretical: add sparse only when the query
share that depends on exact tokens is non-trivial**), the wiring of
QueryFusionRetriever-style fusion (RRF vs alpha-weighted), and per-query-type alpha
tuning. Frame the work as recovering lexical identity that dense pooling destroys,
not as "add keyword search for completeness". Cross-links [[llamaindex]].
the corpus contains exact-match tokens: identifiers, error codes, SKUs, part numbers, API/function names, proper nouns, legal/medical citations, rare jargon…
when_not_to_use
traffic is purely semantic / conceptual with <5% lexical-identity queries — hybrid is over-engineering, the corpus has no stable identifiers and queries never…
Hybrid Retrieval · Dense + Sparse SOP
Third-person operating model for a coder agent that owns retrieval recall on a
corpus where both meaning and exact tokens matter. The audience is the LLM
agent writing or reviewing retrieval code — not an end user.
One sentence: Dense captures meaning, sparse captures exact tokens; hybrid
wins when both matter — but only fuse them when traffic actually carries
exact-match queries, and tune the blend per query type or hybrid loses to dense.
1. 何时激活 (Activation Rules)
Activate this skill when any of the following holds:
The corpus contains exact-match tokens that a query may reference verbatim:
error codes (ERR_SSL_PROTOCOL), SKUs / part numbers (A1-2293-X), API or
function names (as_query_engine), proper nouns, legal/medical citations
(42 U.S.C. § 1983), version strings, rare jargon, ticket IDs.
A bug report says "I searched the exact code/name/string and got nothing", or
"the right document exists but dense retrieval ranks it below fuzzy near-misses".
PR review surfaces a retriever serving lexical-identity traffic but wired
dense-only (index.as_retriever(...) / similarity_search(...) with no
sparse leg).
You are tuning recall and have already exhausted the cheap dense knobs (prompt,
embedding model, chunk size) per the [[llamaindex]] optimization ladder — hybrid
is the next rung.
The user mentions hybrid search, BM25, sparse retrieval, RRF, QueryFusionRetriever,
EnsembleRetriever, or vector-store-native hybrid (Qdrant/Weaviate/Pinecone).
Do not activate when:
Traffic is purely semantic ("what does X mean?", "summarize the policy") with a
lexical-identity share <5% — adding BM25 doubles index footprint for no gain.
The corpus has no stable identifiers and no query ever quotes an exact string.
No dense baseline + eval loop exists yet. Hybrid is a Stage-3 optimization
([[llamaindex]] Stage 3 step 4): baseline and measure before fusing.
2. 核心心智模型 (Core Mental Model)
Three principles. Violating any of them is why teams "try hybrid and conclude it
didn't help".
Principle 1 — Dense and sparse fail in opposite directions
Dense embedding models destroy lexical identity by pooling token representations:
querying a specific error string yields a vector that captures "document about SSL
errors" rather than "document containing this exact string". BM25 does the
inverse — it scores against an inverted index of exact tokens and is blind to
synonyms and paraphrase. (TianPan, Hybrid search in production, 2026; cited in
[[llamaindex]] Dilemma 2.)
Operational corollary: the symptom "I pasted the exact code and got nothing"
is not a bug in the embedding model — it is the embedding model working as
designed. The fix is a second retriever that indexes tokens, not a better
embedding.
Principle 2 — Hybrid is traffic-driven, not theoretical
Whether to add sparse is decided by the query-type distribution of real traffic,
not by a belief that "more retrievers = better". The decision threshold is the
lexical share: the fraction of queries whose correct answer hinges on an exact
token. Below ~5% → dense-only. 5–50% → hybrid. Above ~50% (code, logs, legal) →
invert to sparse-first with dense as a rerank signal. (Cited in [[llamaindex]]
OP-04 / Dilemma 2.)
Principle 3 — One global alpha underperforms; tune per query type
alpha is the dense↔sparse blend (alpha=1 → pure dense, alpha=0 → pure sparse).
A semantic query wants high alpha; a lexical-identity query wants low alpha. A single
global alpha picked to help lexical queries hurts the semantic slice — which is
exactly why a flat alpha "loses to pure dense" and teams wrongly conclude hybrid
failed. Tune alpha per query type, or route per type and pick alpha per route.
(LlamaIndex alpha-tuning blog; [[llamaindex]] Dilemma 2 结果.)
The fusion method (RRF vs weighted) and the fusion parameter (alpha) are two
separate decisions. RRF is score-scale-robust and parameter-light; weighted fusion
is tunable but requires score normalization. Choosing neither deliberately is the
third silent failure.
3. SOP 工作流 (Agentic Protocol)
Four stages. Each gates the next. This overlay assumes a working dense baseline +
eval loop already exists (see [[llamaindex]] Stages 1–2); do not start here.
Stage 0 — Decide if the corpus needs sparse at all
Sample ≥50 real queries (or representative synthetic ones if pre-launch).
Label each: semantic ("what does X mean?") / lexical ("find error code
ABC-123", "the function named foo_bar") / mixed.
Compute the lexical share:
< 5% → dense-only. Stop; hybrid is over-engineering. Record the decision.
5–50% → add hybrid (Stages 1–3).
> 50% (code search, log search, legal citation lookup) → invert: BM25-first,
dense as a fallback / rerank signal.
Confirm the corpus actually carries the tokens queries reference (an identifier
that never appears verbatim in any chunk can't be recovered by BM25 either).
Artifact: a 3-line decision note co-located with the retriever recording the lexical
share and the chosen shape.
Stage 1 — Wire BM25 + dense as two retrievers
python
from llama_index.core.retrievers import QueryFusionRetriever
from llama_index.retrievers.bm25 import BM25Retriever
dense = index.as_retriever(similarity_top_k=10) # embedding leg
sparse = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=10)
Hard rules:
Both legs index the same node set. A BM25 leg built over a different/stale node
set silently drops recall (it can only return tokens it has).
Retrieve wider per leg (top_k 10–20 each) than the final cut — fusion needs
candidates to combine. Narrow after fusion (and after any reranker).
BM25 needs the raw nodes/corpus, not just the vector store — persist or rebuild it
alongside the index, or use a vector store with native hybrid (Stage 1-alt).
Stage 1-alt — vector-store-native hybrid. If the vector store does sparse
internally (Qdrant sparse_vectors, Weaviate hybrid(alpha=...), Pinecone sparse-dense,
pgvector + ts_rank), prefer it: one round-trip, one consistency domain, no separate
BM25 index to keep in sync. Use the framework fusion path only when the store has no
native hybrid.
Stage 2 — Fuse (RRF or alpha-weighted)
Pick the fusion method deliberately (OP-02):
python
# Reciprocal Rank Fusion — score-scale-robust, parameter-light. Good default.
fused = QueryFusionRetriever(
[dense, sparse],
mode="reciprocal_rerank", # RRF: combine by rank, ignore raw score scales
similarity_top_k=8,
num_queries=1, # set >1 to also fan out query rewrites
use_async=True,
)
# Relative-score / weighted fusion — tunable blend; requires score normalization.
fused = QueryFusionRetriever(
[dense, sparse],
mode="relative_score",
retriever_weights=[0.6, 0.4], # ~ alpha=0.6 toward dense
similarity_top_k=8,
)
Default to RRF unless you have a labeled set to tune weights on — RRF sidesteps
the dense-cosine vs BM25-score scale mismatch that breaks naive weighted sums.
Stage 3 — Tune alpha by query type (the gate)
Build a small labeled set per query type (semantic / lexical / mixed) — reuse the
Stage-0 sample.
Evaluate alpha (or retriever_weights) at {0.0, 0.25, 0.5, 0.75, 1.0} on each
type separately, scoring recall / MRR / hit-rate per type.
Expect: lexical queries peak near low alpha (sparse-leaning), semantic near high
alpha (dense-leaning), mixed in between.
If a single global alpha cannot satisfy both types without regressing one →
route by query type and apply per-route alpha (hand off classification to
[[agentsop-query-routing]] if present), or split into two retrievers selected per query.
Gate: hybrid must lift the lexical slice with no regression on the semantic
slice vs the dense baseline. If it regresses semantic, the alpha is wrong, not
hybrid.
The most common false negative: tuning one global alpha, watching semantic recall
drop, and reverting to dense. The correct read is "alpha was global; tune per type".
4. 操作模型 (Operation Models)
Format: Trigger / Action / Output / Evidence.
OP-01 WhenHybridChecklist
Trigger: Deciding whether to add sparse to a dense pipeline.
Action: Run the Stage-0 lexical-share checklist. Confirm: (a) corpus has
exact-match tokens, (b) traffic references them, (c) lexical share ≥5%, (d) those
tokens appear verbatim in chunks. All four must hold.
Output: A go/no-go with the lexical share recorded; "no" is a valid, common
outcome.
Evidence: [[llamaindex]] Dilemma 2 决策步骤; OP-04 ("trigger is traffic-driven,
not theoretical").
OP-02 ChooseFusionMethod
Trigger: Two legs wired; need to combine their result lists.
Action: Default RRF (mode="reciprocal_rerank") — rank-based, robust to the
dense-cosine vs BM25-score scale gap, no weight to tune. Use weighted /
relative-score only when you have a labeled set to fit weights and have normalized
scores. Never naive-sum un-normalized scores.
Output: One deliberate fusion method + the reason it was chosen.
Trigger: Hybrid approved; sparse leg not yet built.
Action: Build BM25Retriever.from_defaults(nodes=...) over the same node
set as the dense index; persist/rebuild it as a deployment artifact alongside the
vector index. Set per-leg similarity_top_k wide (10–20).
Output: A sparse retriever consistent with the dense index, returning enough
candidates to fuse.
Trigger: The vector store supports sparse internally.
Action: Use the store's native hybrid (Qdrant sparse vectors / Prefetch +
fusion, Weaviate hybrid(query, alpha=...), Pinecone sparse-dense vectors, pgvector
full-text + vector) instead of a separate BM25 index. One round-trip, one
consistency domain.
Output: Hybrid with no second index to keep in sync; lower operational surface.
Trigger: Fusion wired; recall not yet optimized; or hybrid "loses to dense".
Action: Evaluate alpha / weights at {0, 0.25, 0.5, 0.75, 1.0} on labeled
subsets per query type, not globally. Lexical → low alpha, semantic → high.
Output: Per-type alpha curve; the alpha that lifts lexical without regressing
semantic.
Evidence: LlamaIndex alpha-tuning in hybrid search blog; [[llamaindex]]
Dilemma 2 ("tune per type, not globally — otherwise hybrid underperforms dense").
OP-06 RouteThenAlpha
Trigger: No single global alpha satisfies both query types without regression.
Action: Classify the query type first, then dispatch to a retriever configured
with the per-type alpha (or pure dense / pure sparse). Hand classification to a
router ([[agentsop-query-routing]] if available).
Output: Each query gets its optimal blend; semantic and lexical slices both peak.
Evidence: [[llamaindex]] Dilemma 2 结果 (per-type tuning); routing as the
mechanism to apply it.
Action: Make BM25 the primary retriever; use dense as a fallback / rerank
signal to catch paraphrase, not as the lead leg. Equivalent to alpha pinned low.
Output: Recall dominated by exact-token matching where that is what users want;
dense recovers the semantic minority.
Evidence: [[llamaindex]] Dilemma 2 step 4 ("BM25-first, dense as fallback rerank
signal"); TianPan production write-up.
OP-08 HybridEvalGate
Trigger: Before merging any hybrid change.
Action: Run the eval loop on both slices: assert lexical-slice recall/MRR
rises vs dense baseline AND semantic-slice metrics do not regress. Gate merge on
both. A lexical lift bought with a semantic regression is not a win.
Output: Quantitative proof hybrid helped where intended and harmed nothing else.
Dilemma 1 — Hybrid (BM25 + dense) vs pure dense: worth the complexity?
困境: Dense is the modern default; BM25 looks like "the old keyword thing". Adding
hybrid doubles the index footprint, adds a BM25 index to keep in sync, and introduces
alpha tuning. Is the complexity justified, or is a better embedding model enough?
约束:
Pure dense silently fails on exact identifiers, code, error strings, SKUs, API
names, rare jargon (TianPan 2026, quoted in Principle 1) — and a better embedding
model does not fix it; lexical loss is a property of pooled representations.
Hybrid adds a second index, a fusion step, and a tuning parameter — real ops cost.
决策步骤:
Build the query-type taxonomy from real traffic: semantic / lexical / mixed
(Stage 0).
Lexical share <5% → dense-only; the complexity is not justified.
Evaluate alpha at {0, 0.25, 0.5, 0.75, 1.0} on labeled per-type subsets.
结果: Hybrid lifts the lexical slice with no degradation on the semantic
slice — if alpha is tuned per type. A single global alpha often loses to pure dense
on semantic queries, which is why teams sometimes wrongly conclude "hybrid didn't
help". The decision is traffic-driven, not theoretical. (Verbatim from
[[llamaindex]] Dilemma 2.)
可提取的操作: OP-01, OP-05, OP-07, OP-08.
Dilemma 2 — Alpha for keyword-heavy vs semantic queries: one knob, two demands
困境: After wiring hybrid, a single alpha must serve both a user pasting
ERR_TLS_CERT_INVALID (wants exact-token match, low alpha) and a user asking "why is
my connection failing?" (wants meaning, high alpha). Picking alpha=0.5 helps neither
fully; picking alpha to win lexical regresses semantic, and vice versa.
约束:
alpha=1 = pure dense, alpha=0 = pure sparse; the optimum differs by query
type, not by corpus.
Choosing alpha to maximize average recall across mixed traffic can land in a valley
that underperforms pure dense on the semantic majority — the classic "hybrid hurt
us" report.
RRF reduces but does not eliminate the tension: it still implicitly weights the two
rankings.
决策步骤:
Split the labeled set by query type (semantic / lexical / mixed).
Sweep alpha per type; observe lexical peaks low, semantic peaks high.
If one global alpha exists that lifts lexical with zero semantic regression →
pin it (cheapest).
Else route by query type and apply per-route alpha / pure-leg selection
(OP-06) — classify first, blend second.
For >50% lexical corpora, skip the balancing act: invert to BM25-first (OP-07).
结果: Per-type alpha (or per-type routing) makes both slices peak simultaneously.
The error to avoid is treating alpha as a single global hyperparameter — that is the
documented cause of "hybrid underperforms dense". (LlamaIndex alpha-tuning blog;
[[llamaindex]] Dilemma 2.)
可提取的操作: OP-05, OP-06, OP-08.
6. 反模式与边界 (Anti-patterns & Boundaries)
Top anti-patterns (instant red flags in code review)
#
Anti-pattern
Why it's wrong
Correct move
A1
Adding hybrid by default on a corpus that is purely semantic
Doubles index footprint and ops surface for no recall gain; <5% lexical traffic
OP-01: gate on lexical share; dense-only is the right answer for semantic corpora
B2 No dense baseline + eval loop yet — baseline first; hybrid is a Stage-3
optimization, not a starting point ([[llamaindex]] Stage 3).
B3 The recall problem is actually chunking, embedding-model mismatch, or missing
metadata filters — fix the cheaper knob first ([[llamaindex]] optimization ladder).
B4 The exact tokens users query don't appear verbatim in any chunk — BM25 can't
recover what isn't indexed; fix ingestion, not retrieval.
B5 Top-1 is wrong but the right doc is in top-k — that's a reranking problem,
not a recall problem ([[llamaindex]] OP-03 AddReranker).
PR review smells
QueryFusionRetriever([...]) with no mode= set and no comment on fusion choice.
A single hard-coded alpha / retriever_weights with no per-type eval behind it.
BM25Retriever.from_defaults(nodes=other_nodes) where other_nodes differs from
the dense index's node set.
Per-leg similarity_top_k equal to the final desired count (no headroom for fusion).
Hybrid added but the eval suite only reports a single aggregate recall number.
Hybrid hand-rolled with dense_scores + bm25_scores summed without normalization.
7. 跨框架对照 (Cross-Framework Reference Table)
How "dense + sparse, fused" looks across the common stacks. Cross-links [[llamaindex]].
7.1 LlamaIndex — QueryFusionRetriever
python
from llama_index.core.retrievers import QueryFusionRetriever
from llama_index.retrievers.bm25 import BM25Retriever
dense = index.as_retriever(similarity_top_k=10)
sparse = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=10)
fused = QueryFusionRetriever(
[dense, sparse],
mode="reciprocal_rerank", # or "relative_score" / "dist_based_score"
retriever_weights=[0.6, 0.4], # used by weighted modes; ~alpha toward dense
similarity_top_k=8,
num_queries=1, # >1 also fans out query rewrites
)
Fusion modes: reciprocal_rerank (RRF, default-recommended), relative_score,
dist_based_score (the latter two are weighted). ([[llamaindex]] OP-04.)
EnsembleRetriever fuses with Reciprocal Rank Fusion under the hood; weights
biases the RRF contribution per retriever (the LangChain analogue of alpha).
7.3 Raw RRF (framework-agnostic)
python
def rrf(result_lists, k=60, top_k=8):
scores = {}
for results in result_lists: # each = ranked list of doc ids
for rank, doc_id in enumerate(results):
scores[doc_id] = scores.get(doc_id, 0) + 1.0 / (k + rank + 1)
return sorted(scores, key=scores.get, reverse=True)[:top_k]
fused = rrf([dense_ids, bm25_ids]) # rank-based, no score normalization
RRF combines by rank, so dense-cosine and BM25 raw scores never need to share a
scale — the reason it is the safe default (OP-02). k≈60 is the canonical constant
(Cormack et al., 2009).
tsvector full-text + vector distance, combined in SQL
hand-weighted in the ORDER BY expression
Milvus
hybrid search with WeightedRanker / RRFRanker
ranker choice + weights
Prefer native hybrid when available (OP-04): one round-trip, one consistency domain,
no separate BM25 index to keep in sync. Weaviate's alpha is the literal parameter the
[[llamaindex]] alpha-tuning blog generalizes.
Agentsop Hybrid Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Agentsop Hybrid Retrieval compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Agentsop Hybrid Retrieval this skillagentsope/SkillAlchemy
Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.
Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
Enhancement-overlay SOP for adding sparse (BM25 / keyword) retrieval alongside dense (embedding) retrieval. Agentsop Hybrid Retrieval is an agent skill from agentsope/SkillAlchemy. Enhancement-overlay SOP for adding sparse (BM25 / keyword) retrieval alongside dense (embedding) retrieval.
When should I use Agentsop Hybrid Retrieval?
Agentsop Hybrid Retrieval fits situations like: tasks that involve Operations and SOPs; tasks that involve Embeddings; tasks that involve Retrieval-augmented generation.
How do I install Agentsop Hybrid Retrieval in Claude Code?
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a claude-code`. Or copy the skill folder (skills/agentsop-hybrid-retrieval in agentsope/SkillAlchemy) into .claude/skills/agentsop-hybrid-retrieval in your project. Claude Code loads it when a task matches its description.
How do I install Agentsop Hybrid Retrieval in Codex?
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a codex`. Or copy the skill folder (skills/agentsop-hybrid-retrieval in agentsope/SkillAlchemy) into .agents/skills/agentsop-hybrid-retrieval in your project. Codex loads it when a task matches its description.
Can I use Agentsop Hybrid Retrieval in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-hybrid-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-hybrid-retrieval, .gemini/skills/agentsop-hybrid-retrieval, .github/skills/agentsop-hybrid-retrieval and .opencode/skills/agentsop-hybrid-retrieval in your project.
What does Agentsop Hybrid Retrieval need to run?
SKILL.md names no scripts, command-line tools or credentials: Agentsop Hybrid Retrieval is instructions for the agent only. Our summary lists: Python 3.
Does Agentsop Hybrid Retrieval access the network?
SKILL.md names 7 domains. As links in the text: developers.llamaindex.ai, llamaindex.ai, tianpan.co, python.langchain.com, qdrant.tech, docs.weaviate.io and docs.pinecone.io. This is read from the text; nothing was executed.
Is Agentsop Hybrid Retrieval safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Agentsop Hybrid Retrieval use?
Agentsop Hybrid Retrieval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Agentsop Hybrid Retrieval use?
About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.
What are the alternatives to Agentsop Hybrid Retrieval?
Skills that share tags, products or a category with Agentsop Hybrid Retrieval: Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), Cohere Core Workflow A (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), FAISS Similarity Search (Orchestra-Research/AI-Research-SKILLs, 13k stars) and RAG Skills (llama-farm/llamafarm, 836 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Agentsop Hybrid Retrieval?
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.