Official agent skill

Qdrant Search Quality Diagnosis

by qdrant in qdrant/skills

Diagnoses Qdrant search quality issues. An agent skill from qdrant/skills.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Qdrant Search Quality Diagnosis

skills CLI
$ npx skills add qdrant/skills --skill qdrant-search-quality-diagnosis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install qdrant/skills qdrant-search-quality-diagnosis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/qdrant/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qdrant-search-quality/diagnosis .claude/skills/qdrant-search-quality-diagnosis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qdrant-search-quality-diagnosis
GitHub stars
254
Token cost
~2.3k tokens
SKILL.md length
1,000 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses Qdrant search quality issues. An agent skill from qdrant/skills.

  • Someone reports results are bad
  • SKILL.md covers Need a Labeled Baseline to…, Don't Know What's Wrong Yet, Approximate Search Worse Than… and Wrong Embedding Model, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Not relevant results

What it does

Qdrant Search Quality Diagnosis is an agent skill from qdrant/skills, published by the product's own GitHub organization. Diagnoses Qdrant search quality issues. Use when someone reports 'results are bad', 'wrong results', 'not relevant results', 'missing matches', 'recall is low', 'approximate search worse than exact', 'which embedding model', 'should I fine-tune my embedding model', 'quality dropped after quantization', 'how to measure retrieval quality', 'build a golden set', 'ground truth dataset', 'how to score recall@k', or 'is my improvement real / statistically significant'. Also use when search quality degrades without…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Vector databases, Embeddings and LLM inference and serving. It works with Qdrant. The repository describes itself as: Agent skills for Qdrant vector search: scaling, performance optimization, search quality, monitoring, deployment, model migration, version upgrades, and SDK usage across Python…. The licence is Apache-2.0.

When your agent uses it

  • Someone reports results are bad
  • Not relevant results
  • Missing matches
  • Approximate search worse than exact

Example prompts

  • “results are bad”
  • “wrong results”
  • “not relevant results”
  • “/qdrant-search-quality-diagnosis”

What it can do on your machine

Read from SKILL.md and the folder at commit 1780b6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • skills.qdrant.tech

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Qdrant Search Quality Diagnosis loads about 2.3k tokens when it runs. Until then it costs about 141 tokens; SKILL.md has 1,000 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~141
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from qdrant/skills at commit 1780b6d, republished under its Apache-2.0 licence (© qdrant). 1,000 words, ~2,333 tokens.

Download SKILL.mdSave it as .claude/skills/qdrant-search-quality-diagnosis/SKILL.md (or your agent's skills folder).
name
qdrant-search-quality-diagnosis
description
Diagnoses Qdrant search quality issues. Use when someone reports 'results are bad', 'wrong results', 'not relevant results', 'missing matches', 'recall is low', 'approximate search worse than exact', 'which embedding model', 'should I fine-tune my embedding model', 'quality dropped after quantization', 'how to measure retrieval quality', 'build a golden set', 'ground truth dataset', 'how to score recall@k', or 'is my improvement real / statistically significant'. Also use when search quality degrades without obvious changes.

How to Diagnose Bad Search Quality

Before diagnosing or evaluating search quality, establish two different ground truths: a labeled query set measures whether the results are the right ones; exact KNN measures whether ANN finds what brute force would (no labels needed). The first diagnoses relevance, the second diagnoses the index.

Need a Labeled Baseline to Score Quality or Validate a Gain

Use when: user has no golden set, asks "how do I know if my search is good?", or needs to gate releases on a retrieval metric. Every fix in the sections below should be validated this way.

  • Pick the metric by usage: Recall@k for RAG, MRR/Hits@1 for single-answer, NDCG@k for re-ranking Choosing the metric
  • Build a labeled query set (human, log-based, or LLM-synthetic) and score retrieval with ranx Measuring Retrieval Relevance
  • When tuning and evaluating search quality, set a minimum target gain, then make sure your labeled query set is large enough to measure it reliably. Quantify the uncertainty around a measured gain with a confidence interval (e.g. bootstrap the per-query gains, 95%): if it includes zero, the improvement is inconclusive and may be due to query-to-query variation. Adding labeled queries narrows the interval; how many you need depends on the size of the gain you are targeting and on query-to-query variation, not on collection size Before tuning a collection
  • If you use the labeled queries to tune or select a configuration, evaluate the final choice on held-out queries. A gain that doesn't survive the held-out evaluation isn't reliable evidence of an improvement.
  • For full RAG pipelines, also score generation with Ragas and use the retrieval-vs-generation 2x2 to isolate regressions Pipeline Output Quality
  • Gate CI on a per-metric threshold to catch regressions from embedding-model swaps, prompt changes, or index config changes

Don't Know What's Wrong Yet

Use when: results are irrelevant or missing expected matches and you need to isolate the cause.

  • For a no-code quick check, use the Web UI's ANN Recall tab to compare approximate vs exact recall@k Web UI ANN Recall
  • For the same comparison in code (CI gating, regression tests), run each query twice: once approximate, once with exact=true. Compute recall@k from the overlap ANN recall in CI
  • Target >95% recall@k in production. Exact search bad = model, data, or search pipeline problem. Exact good, approximate bad = tune HNSW.
  • Match the embedding model to your data: the context window should fit your chunk length with room for special tokens, the model must cover every language in your corpus, and its training domain should resemble yours (e.g. a code-trained model for code) How to choose an embedding model
  • Make sure documents are properly chunked; splitting chunks mid-sentence alone can drop quality by 30-40%
  • Check if quantization degrades quality (compare with and without)
  • Check if filters are too restrictive (then you might need to use ACORN)
  • If duplicate results from chunked documents, use Grouping API to deduplicate Grouping

Payload filtering and sparse vector search are different things. Metadata (dates, categories, tags) goes in payload for filtering. Text content goes in sparse vectors for search.

Approximate Search Worse Than Exact

Use when: exact search returns good results but HNSW approximation misses them.

Binary quantization requires rescore. Without it, quality loss is severe. Use oversampling to recover recall: the docs report 0.98 recall with 2x oversampling on 4096-dimensional and 4x on 1536-dimensional embeddings, so start around 2-4x and tune on your data. Always test quantization impact on your data before production. Quantization

Show full SKILL.md (359 more words)Show less

Wrong Embedding Model

Use when: exact search also returns bad results.

  • Check Qdrant team recommendations on how to choose an embedding model.

  • Test top 3 MTEB models on 100-1000 sample queries Hosted Qdrant inference. Score them against a labeled set to compare apples to apples Measuring Retrieval Relevance.

  • If your data is strongly hierarchical (taxonomies, product catalogs, part-whole relationships), consider hyperbolic (Poincaré) embeddings. They capture tree structure in far fewer dimensions than flat ones. In Qdrant, use Euclidean HNSW to pull a candidate set from the original Poincaré coordinates, then a Formula Query to rescore with the real hyperbolic distance. How to serve hyperbolic embeddings with Qdrant.

  • Consider fine-tuning an embedding model for your specific use case only after trying better-suited models and retrieval/pipeline tuning and confirming that the embedding model remains the bottleneck. Fine-tuning is most useful when general-purpose embeddings fail to capture important domain- or task-specific distinctions and you have good labeled query-document pairs. Fine-tuning requires re-embedding and re-indexing the collection Model migration.

Unoptimized Search Pipeline

Use when: exact search also returns bad results and model choice is confirmed by user.

  • Optimize search according to the advanced search-strategies skill
  • Check candidate depth: if your retrieval pipeline includes a first-stage retriever that feeds a reranker or fusion stage, test whether increasing the prefetch limit improves your quality metric. A downstream ranker cannot recover relevant documents that never enter its candidate set. For hybrid search, start around limit=100-200 and test larger values against your labeled queries Candidate depth

What NOT to Do

  • Tune Qdrant before verifying the model is right for the task (most quality issues are model issues)
  • Use binary quantization without rescore (severe quality loss)
  • Set hnsw_ef lower than results requested (guaranteed bad recall)
  • Skip payload indexes on filtered fields then blame quality (HNSW can't traverse filtered-out nodes, and filterable HNSW is built only if payload indexes were set up prior)
  • Deploy without baseline recall or other search relevance metrics (no way to measure regressions)
  • Compare two configs on a query set too small to resolve the difference between them (the result is noise, not evidence)
  • Confuse payload filtering with sparse vector search (different things, different config)

© qdrant, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/qdrant-search-quality/diagnosis of qdrant/skills.

Open the folder on GitHubat commit 1780b6d

Compare with similar skills

Qdrant Search Quality Diagnosis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Qdrant Search Quality Diagnosis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Qdrant Search Quality Diagnosis this skillqdrant/skills254—~2.3kAutomated safety check: PassApache-2.0
Qdrant Search Quality Diagnosisgithub/awesome-copilot40k1 repos~928Automated safety check: PassMIT
Qdrant Search Qualitygithub/awesome-copilot40k1 repos~336Automated safety check: PassMIT
Vector DBericrisco/rsc-harness180—~2.8kAutomated safety check: PassMIT
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0

Similar skills

  • Qdrant Search Quality Diagnosis

    github/awesome-copilot

    Official

    Diagnoses Qdrant search quality issues. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~928 tokens
    AI & LLM EngineeringAuto-check passed
  • Qdrant Search Quality

    github/awesome-copilot

    Official

    Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~336 tokens
    AI & LLM EngineeringAuto-check passed
  • Vector DB

    ericrisco/rsc-harness

    A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

    180 GitHub stars~2.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed

More from qdrant/skills

All 33 skills in this repo
  • Qdrant Clients SDK

    qdrant/skills

    Official

    Qdrant provides client SDKs for various programming languages, allowing easy integration with Qdrant deployments.

    254 GitHub starsUsed in 2 repos~752 tokens
    Auto-check: notes
  • Qdrant Advisor

    qdrant/skills

    Official

    Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.

    254 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Guides Qdrant deployment selection. An agent skill from qdrant/skills.

    254 GitHub starsUsed in 2 repos~976 tokens
    Auto-check passed
  • Official

    Guides Qdrant search strategy selection. An agent skill from qdrant/skills.

    254 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Official

    Diagnoses and guides Qdrant horizontal scaling decisions. An agent skill from qdrant/skills.

    254 GitHub starsUsed in 2 repos~833 tokens
    Auto-check passed

Works with

Questions about Qdrant Search Quality Diagnosis

What does Qdrant Search Quality Diagnosis do?

Diagnoses Qdrant search quality issues. An agent skill from qdrant/skills. Qdrant Search Quality Diagnosis is an agent skill from qdrant/skills, published by the product's own GitHub organization. Diagnoses Qdrant search quality issues.

When should I use Qdrant Search Quality Diagnosis?

Qdrant Search Quality Diagnosis fits situations like: someone reports results are bad; not relevant results; missing matches; approximate search worse than exact.

How do I install Qdrant Search Quality Diagnosis in Claude Code?

Run `npx skills add qdrant/skills --skill qdrant-search-quality-diagnosis -a claude-code`. Or copy the skill folder (skills/qdrant-search-quality/diagnosis in qdrant/skills) into .claude/skills/qdrant-search-quality-diagnosis in your project. Claude Code loads it when a task matches its description.

How do I install Qdrant Search Quality Diagnosis in Codex?

Run `npx skills add qdrant/skills --skill qdrant-search-quality-diagnosis -a codex`. Or copy the skill folder (skills/qdrant-search-quality/diagnosis in qdrant/skills) into .agents/skills/qdrant-search-quality-diagnosis in your project. Codex loads it when a task matches its description.

Can I use Qdrant Search Quality Diagnosis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add qdrant/skills --skill qdrant-search-quality-diagnosis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qdrant-search-quality-diagnosis, .gemini/skills/qdrant-search-quality-diagnosis, .github/skills/qdrant-search-quality-diagnosis and .opencode/skills/qdrant-search-quality-diagnosis in your project.

What does Qdrant Search Quality Diagnosis need to run?

SKILL.md names no scripts, command-line tools or credentials: Qdrant Search Quality Diagnosis is instructions for the agent only.

Does Qdrant Search Quality Diagnosis access the network?

SKILL.md names 1 domain. As links in the text: skills.qdrant.tech. This is read from the text; nothing was executed.

Is Qdrant Search Quality Diagnosis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Qdrant Search Quality Diagnosis use?

Qdrant Search Quality Diagnosis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Qdrant Search Quality Diagnosis use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Qdrant Search Quality Diagnosis?

Skills that share tags, products or a category with Qdrant Search Quality Diagnosis: Qdrant Search Quality Diagnosis (github/awesome-copilot, 40k stars), Qdrant Search Quality (github/awesome-copilot, 40k stars), Vector DB (ericrisco/rsc-harness, 180 stars) and Codebase Management (giancarloerra/SocratiCode, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Qdrant Search Quality Diagnosis?

qdrant (a GitHub organization, an official publisher) maintains it in qdrant/skills, which has 254 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 9, 2026.

Source: qdrant/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.