Official agent skill

Qdrant Hybrid Search Prefetches

by qdrant in qdrant/skills

Constructing prefetch queries for hybrid retrieval, including sparse/dense and multi-field setups, and choosing a sparse embedding model.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Qdrant Hybrid Search Prefetches

skills CLI
$ npx skills add qdrant/skills --skill qdrant-hybrid-search-prefetches -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install qdrant/skills qdrant-hybrid-search-prefetches --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/qdrant/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qdrant-search-quality/search-strategies/hybrid-search/search-types .claude/skills/qdrant-hybrid-search-prefetches && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qdrant-hybrid-search-prefetches
GitHub stars
254
Token cost
~2.4k tokens
SKILL.md length
1,126 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Constructing prefetch queries for hybrid retrieval, including sparse/dense and multi-field setups, and choosing a sparse embedding model.

  • Works in 2 steps: The same vector representations but… → Different vector representations but the…
  • Someone asks dense and sparse in one search?
  • SKILL.md covers Missed Keyword Matches, Need to Combine Multiple… and What NOT to Do
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Qdrant Hybrid Search Prefetches is an agent skill from qdrant/skills, published by the product's own GitHub organization. Constructing prefetch queries for hybrid retrieval, including sparse/dense and multi-field setups, and choosing a sparse embedding model. Use when someone asks 'dense and sparse in one search?', 'how to combine multiple fields for retrieval?', 'payloads or sparse vectors for lexical?', 'which sparse embedding model to use?', or 'BM25 vs SPLADE?'

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Vector databases, Retrieval-augmented generation and Embeddings. It works with Qdrant. The repository describes itself as: Agent skills for Qdrant vector search: scaling, performance optimization, search quality, monitoring, deployment, model migration, version upgrades, and SDK usage across Python…. The licence is Apache-2.0.

When your agent uses it

  • Someone asks dense and sparse in one search?
  • How to combine multiple fields for retrieval?
  • Sparse vectors for lexical?
  • Which sparse embedding model to use?

Example prompts

  • “dense and sparse in one search?”
  • “how to combine multiple fields for retrieval?”
  • “payloads or sparse vectors for lexical?”
  • “/qdrant-hybrid-search-prefetches”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The same vector representations but different queries or filters.
  2. Different vector representations but the same raw query.

What it can do on your machine

Read from SKILL.md and the folder at commit 1780b6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • skills.qdrant.tech

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Qdrant Hybrid Search Prefetches loads about 2.4k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 1,126 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from qdrant/skills at commit 1780b6d, republished under its Apache-2.0 licence (© qdrant). 1,126 words, ~2,414 tokens.

Download SKILL.mdSave it as .claude/skills/qdrant-hybrid-search-prefetches/SKILL.md (or your agent's skills folder).
name
qdrant-hybrid-search-prefetches
description
Constructing prefetch queries for hybrid retrieval, including sparse/dense and multi-field setups, and choosing a sparse embedding model. Use when someone asks 'dense and sparse in one search?', 'how to combine multiple fields for retrieval?', 'payloads or sparse vectors for lexical?', 'which sparse embedding model to use?', or 'BM25 vs SPLADE?'

Different Searches in One Query API Request

Each prefetch runs exactly one search per one query.

Understand if user wants to run several parallel searches on:

  1. The same vector representations but different queries or filters.
  2. Different vector representations but the same raw query.

If first, help user to design logic of constructing query or/and filters on application side and then check Combining Searches. Don't forget to create indices on filterable payload fields, immediately after collection creation, prior to building HNSW, so filterable HNSW could be constructed.

If second, use named vectors, which allow to store multiple vector types per point in one collection. On Qdrant 1.18 or newer, named vectors can be added to or removed from an existing collection without recreating it Update vector schema; on 1.17 or older, they can be configured only at collection creation. To choose vectors, check following recommendations.

Missed Keyword Matches

Use when: pure vector search misses exact term or keyword matches and you need lexical retrieval alongside semantic search.

Most likely you need a sparse vector for exact text search alongside the dense one. Qdrant uses sparse vectors for lexical searches, as payload filtering doesn't provide any ranking score.

Choose a Sparse Vector for Text
  • BM25 statistical representations, built into Qdrant core (computed server-side). Good baseline, works out-of-domain, usually for long texts. Can be used for non-English content, but needs to be configured per language (tokenization, stemming, stopwords, etc) at indexing and retrieval time. More in Text Search Guide
  • BM42 learned sparse, based on BM25, but better for small chunks of text & with meaning understanding. Works only on English. Requires fine-tuning for domain-specific retrieval. Requires FastEmbed (Python/REST only, not available in all SDKs). Not maintained.
  • miniCOIL learned sparse, BM25 with additional understanding of words meaning in context. Works only on English. Requires fine-tuning for domain-specific retrieval. Requires FastEmbed. Usage shown in FastEmbed miniCOIL documentation.
  • SPLADE++ learned sparse with term expansion. Heavier inference and resources usage but better performance due to term expansion. Requires fine-tuning for domain-specific retrieval. Provided in Qdrant Cloud Inference and FastEmbed versions work only on English. To use with FastEmbed, check FastEmbed SPLADE documentation.
  • External learned sparse embeddings, for example BAAI/bge-m3.

What to remember when using sparse vectors for lexical search:

  • tokenization and stemming affect exact matches, especially on custom codes, terms, etc.
  • IDF statistics are computed over a corpus. On Qdrant 1.18 and older that corpus is the data in the whole shard being queried, which distorts scoring in multi-tenant collections: tenants with different vocabularies get merged into one statistic, so a term's IDF no longer reflects how rare it is inside either tenant's data. On 1.19+ you can narrow the corpus the statistics are computed over down to a single tenant, so IDF reflects term rarity within that tenant. See Per-tenant IDF statistics.

What to remember when using Qdrant BM25 and miniCOIL (based on BM25):

  • avg_len in formula is not computed server-side, it is a user responsibility and passed as a parameter. Calibrate per field — defaults assume document-length text; short fields (titles, tags) need a much smaller value or BM25 scoring is skewed (avg_len=256 against a 10-word title overweights term frequency).
  • BM25 might be not good for small chunks of text, as BM25 algorithm was initially created for search on long documents; consider adjusting document statistics in sparse vectors (TF & IDF, k, b).
  • Qdrant BM25 vectors are configured per language, so consider customizing stop words, stemming & tokenization when users documents mix several languages or carefully configure vectors per point when they are monolingual. To disable text processing entirely for language-neutral content: on Qdrant 1.19 or newer, use stemmer: {"type": "none"} plus an empty stopwords set explicitly (by default, BM25 applies English stemming and stopword removal); on 1.18 or older, use language: none (deprecated as of 1.19).

More on Sparse Vectors for Text Search

Show full SKILL.md (489 more words)Show less

Need to Combine Multiple Representations of the Same Item

Use when: the same item is embedded in multiple ways (e.g. different models, languages, modalities, or different fields like title/abstract/chunk) and you want to search across different representations in one request (don't have to be all of them, can be even one).

Use multiple named vector prefetches, each prefetch covers one representation.

A representation only earns its own prefetch if it carries signal independent of the others — e.g. title vocabulary the body never repeats, or an abstract treated as a single semantic unit vs. individual chunks. Don't add a prefetch per field reflexively; verify each candidate contributes content the other vectors don't.

When a representation's signal is mostly lexical — keyword-driven titles, codes, tags, or other short fields — prefer a sparse named vector (e.g. BM25) over an additional dense embedding. Server-side BM25 in Qdrant avoids the inference cost of another dense model and stores far less per point. Skip this when the field carries paraphrase or conceptual signal that exact-term matching would miss.

  • End-to-end worked example fusing title, abstract, chunk, and sparse-title named vectors with RRF and document-level grouping in one Query API call: Multi-Representation Search tutorial
  • If you have groups and subgroups of representations (document -> chunk, image -> patch), you could use searching in groups. To not store identical payloads several times, check Lookup in Groups. Index the grouping payload field (e.g. document_id) as a keyword payload index before grouping.
  • When grouping chunk-level points back to documents, each prefetch only contributes the candidates it returned — so size per-prefetch limit well above the final document limit (rule of thumb: prefetch_limit ≥ final_limit × expected_chunks_per_document), otherwise a few documents with many chunks saturate the candidate pool and relevant documents drop silently. Validate grouped recall on a labeled sample.
  • When per-document vectors (title, abstract) would be duplicated across every chunk-level point, the duplication can dominate storage at scale. Keeping them denormalized in one collection makes queries simpler (single Query API call, every representation reachable from any point); a sidecar collection joined via Lookup in Groups is the alternative when storage matters.

You can also search directly on multivectors, a matrix of dense vectors, in a prefetch.

However, it comes with several considerations, as multivectors were designed to support late interaction models using max similarity metric, so it's impossible to retrieve the list of individual max similarity scores for each query vector.

Moreover, multivectors are rarely a good pick for prefetch:

There are ways to make multivector retrieval cheaper (MUVERA, pooling), you can see more in "Evaluating Tradeoffs of Multi-stage Multi-vector Search"

What NOT to Do

  • Choose any search method (for example, BM25) without evaluation of its quality & resources used.
  • Use any search method (for example, BM25) without paying attention to the specifics of their configuration and applicability to the use case.

© qdrant, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/qdrant-search-quality/search-strategies/hybrid-search/search-types of qdrant/skills.

Open the folder on GitHubat commit 1780b6d

Compare with similar skills

Qdrant Hybrid Search Prefetches next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Qdrant Hybrid Search Prefetches compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Qdrant Hybrid Search Prefetches this skillqdrant/skills254—~2.4kAutomated safety check: PassApache-2.0
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Hunt RAG Vectorelementalsouls/Claude-BugHunter4.8k—~2.6kAutomated safety check: PassMIT
Qdrant Search Qualitygithub/awesome-copilot40k1 repos~336Automated safety check: PassMIT
RAG ArchitectJeffallan/claude-skills12k—~2kAutomated safety check: PassMIT
RAG Patternssoftspark/ai-toolkit179—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Qdrant Search Quality

    github/awesome-copilot

    Official

    Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~336 tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub stars~2k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • RAG Patterns

    softspark/ai-toolkit

    RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.

    179 GitHub stars~1.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from qdrant/skills

All 33 skills in this repo
  • Qdrant Clients SDK

    qdrant/skills

    Official

    Qdrant provides client SDKs for various programming languages, allowing easy integration with Qdrant deployments.

    254 GitHub starsUsed in 2 repos~752 tokens
    Auto-check: notes
  • Qdrant Advisor

    qdrant/skills

    Official

    Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.

    254 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed
  • Official

    Guides Qdrant deployment selection. An agent skill from qdrant/skills.

    254 GitHub starsUsed in 2 repos~976 tokens
    Auto-check passed
  • Official

    Guides Qdrant search strategy selection. An agent skill from qdrant/skills.

    254 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed
  • Official

    Diagnoses and guides Qdrant horizontal scaling decisions. An agent skill from qdrant/skills.

    254 GitHub starsUsed in 2 repos~833 tokens
    Auto-check passed

Works with

Questions about Qdrant Hybrid Search Prefetches

What does Qdrant Hybrid Search Prefetches do?

Constructing prefetch queries for hybrid retrieval, including sparse/dense and multi-field setups, and choosing a sparse embedding model. Qdrant Hybrid Search Prefetches is an agent skill from qdrant/skills, published by the product's own GitHub organization. Constructing prefetch queries for hybrid retrieval, including sparse/dense and multi-field setups, and choosing a sparse embedding model.

When should I use Qdrant Hybrid Search Prefetches?

Qdrant Hybrid Search Prefetches fits situations like: someone asks dense and sparse in one search?; how to combine multiple fields for retrieval?; sparse vectors for lexical?; which sparse embedding model to use?.

How do I install Qdrant Hybrid Search Prefetches in Claude Code?

Run `npx skills add qdrant/skills --skill qdrant-hybrid-search-prefetches -a claude-code`. Or copy the skill folder (skills/qdrant-search-quality/search-strategies/hybrid-search/search-types in qdrant/skills) into .claude/skills/qdrant-hybrid-search-prefetches in your project. Claude Code loads it when a task matches its description.

How do I install Qdrant Hybrid Search Prefetches in Codex?

Run `npx skills add qdrant/skills --skill qdrant-hybrid-search-prefetches -a codex`. Or copy the skill folder (skills/qdrant-search-quality/search-strategies/hybrid-search/search-types in qdrant/skills) into .agents/skills/qdrant-hybrid-search-prefetches in your project. Codex loads it when a task matches its description.

Can I use Qdrant Hybrid Search Prefetches in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add qdrant/skills --skill qdrant-hybrid-search-prefetches -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qdrant-hybrid-search-prefetches, .gemini/skills/qdrant-hybrid-search-prefetches, .github/skills/qdrant-hybrid-search-prefetches and .opencode/skills/qdrant-hybrid-search-prefetches in your project.

What does Qdrant Hybrid Search Prefetches need to run?

SKILL.md names no scripts, command-line tools or credentials: Qdrant Hybrid Search Prefetches is instructions for the agent only.

Does Qdrant Hybrid Search Prefetches access the network?

SKILL.md names 1 domain. As links in the text: skills.qdrant.tech. This is read from the text; nothing was executed.

Is Qdrant Hybrid Search Prefetches safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Qdrant Hybrid Search Prefetches use?

Qdrant Hybrid Search Prefetches is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Qdrant Hybrid Search Prefetches use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Qdrant Hybrid Search Prefetches?

Skills that share tags, products or a category with Qdrant Hybrid Search Prefetches: RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars), Qdrant Search Quality (github/awesome-copilot, 40k stars) and RAG Architect (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Qdrant Hybrid Search Prefetches?

qdrant (a GitHub organization, an official publisher) maintains it in qdrant/skills, which has 254 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 9, 2026.

Source: qdrant/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.