Official agent skill

Qdrant Sizing

by qdrant in qdrant/skills

Sizes a Qdrant deployment before it is provisioned. An agent skill from qdrant/skills.

OfficialApache-2.0Auto-check passedDatabases

Install Qdrant Sizing

skills CLI
$ npx skills add qdrant/skills --skill qdrant-sizing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install qdrant/skills qdrant-sizing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/qdrant/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qdrant-sizing .claude/skills/qdrant-sizing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qdrant-sizing
GitHub stars
253
Token cost
~2.4k tokens
SKILL.md length
1,226 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Sizes a Qdrant deployment before it is provisioned. An agent skill from qdrant/skills.

  • Someone asks how much RAM do I need
  • SKILL.md covers Sizing RAM and Disk, Sizing CPU, GPU, and Node Count, Validating the Estimate Before… and What NOT to Do
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • How big should my cluster be

What it does

Qdrant Sizing is an agent skill from qdrant/skills, published by the product's own GitHub organization. Sizes a Qdrant deployment before it is provisioned. Use when someone asks 'how much RAM do I need', 'how many nodes', 'how big should my cluster be', 'sizing', 'capacity planning', 'will N vectors fit', 'what instance type should I pick', or gives a vector count and dimensions and asks what to provision. Also use when an existing estimate needs checking before hardware or a cluster tier is bought.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Vector databases. It works with Qdrant. The repository describes itself as: Agent skills for Qdrant vector search: scaling, performance optimization, search quality, monitoring, deployment, model migration, version upgrades, and SDK usage across Python…. The licence is Apache-2.0.

When your agent uses it

  • Someone asks how much RAM do I need
  • How big should my cluster be
  • Capacity planning
  • Will N vectors fit

Example prompts

  • “how much RAM do I need”
  • “how many nodes”
  • “how big should my cluster be”
  • “/qdrant-sizing”

What it can do on your machine

Read from SKILL.md and the folder at commit 476a18d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • skills.qdrant.tech
    • sizing.qdrant.tech

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Qdrant Sizing loads about 2.4k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 1,226 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from qdrant/skills at commit 476a18d, republished under its Apache-2.0 licence (© qdrant). 1,226 words, ~2,356 tokens.

Download SKILL.mdSave it as .claude/skills/qdrant-sizing/SKILL.md (or your agent's skills folder).
name
qdrant-sizing
description
Sizes a Qdrant deployment before it is provisioned. Use when someone asks 'how much RAM do I need', 'how many nodes', 'how big should my cluster be', 'sizing', 'capacity planning', 'will N vectors fit', 'what instance type should I pick', or gives a vector count and dimensions and asks what to provision. Also use when an existing estimate needs checking before hardware or a cluster tier is bought.

Sizing a Qdrant Deployment

Sizing is not points × dims × 4. Raw vectors are only one part of the footprint. Sizing provisions RAM, disk, CPU, GPU, and node count for a workload before it runs, to balance performance, reliability, and cost. Each resource is driven by different requirements:

  • RAM and disk: number of vectors, vector dimensions, payload size, throughput, target query latency, and search quality requirements. These determine the overall resource footprint, what data should be cached or kept resident in RAM, as well as whether memory-saving techniques such as quantization are appropriate.
  • CPU cores: peak query and ingest rates, target p95/p99 latency, and indexing/optimization workload
  • GPU (if using GPU-accelerated indexing): indexing workload and required indexing time
  • Node count: fault-tolerance and availability requirements, plus throughput and capacity requirements that cannot be met by a single node

Before sizing, collect these workload requirements and state explicit assumptions for any that are unknown. Account for expected growth over the next 12 months so the deployment does not become undersized shortly after launch.

Sizing RAM and Disk

Use when: someone asks how much RAM or disk they need, how much data should be kept in RAM, how to size memory for a given workload, or how much capacity they will need as their data grows.

Estimate the data footprint

Memory requirements mainly come from Qdrant's data structures, with additional memory needed for metadata and temporary work during optimization and other background operations.

The following estimates break down the data footprint by component. Each component scales with base = points × replication_factor. Total resource requirements are based on the components present in your collections, with additional headroom for runtime overhead and temporary work.

  • Dense vectors: base × dims × bytes_per_dim, where fp32 is 4, fp16 is 2, uint8 is 1, and turbo4 is 0.5 Vector datatypes.
  • Quantized vectors: base × dims × quant_bytes Quantization. Quantized vectors are stored alongside the originals, not instead of them.
  • HNSW: base × m × 2 × 4 × 1.2, where m is the number of edges per node in the index graph (defaults to 16).
  • Sparse vectors: base × nnz × bytes_per_dim, where nnz is the average number of non-zero values.
  • Sparse index (inverted index): base × nnz × bytes_per_dim × 1.5

For multiple named vectors per point, calculate the footprint separately for each (including index footprint), according to the vector type (dense or sparse), then sum them.

  • Payload: disk: base × avg_payload_size × 1.5; in-RAM: base × avg_payload_size × 1.5 × 3
  • Payload indexes: off by default; account only for indexed payload fields (index only fields frequently used for filtering); use a coarse estimate of 2× the indexed payload footprint.

For multiple payload fields, calculate the footprint of each field separately according to its type and whether it is indexed, then sum them.

  • ID tracker: ~52 bytes × base (always resident in RAM)
Decide what needs to be loaded in RAM

Qdrant persists all collection data to disk. Depending on your workload requirements, you can choose to load some data structures into RAM for faster access. On Qdrant 1.19+, configure this per structure with memory: pinned, cached, or cold; on 1.18 and older, use always_ram and on_disk. Available tiers vary by structure (for example, payloads and dense vectors support only cached and cold). Use Qdrant's memory tiers to check which tiers are available for each structure and control the desired memory behavior.

You can choose the desired memory tier for each structure, except:

  • ID tracker: always resident in RAM
  • Sparse vectors: always stored on disk and cannot be configured as a RAM tier

Check the default memory tiers before overriding them.

Recommendations:

  • Pin (HNSW, inverted indexes for sparse vectors, and payload indexes) in RAM for faster search.
  • Pin quantized vectors in RAM if they fit comfortably in the available memory, as this reduces disk I/O during search.
  • If your use case involves splitting vectors into multiple collections or subgroups based on payload values (e.g., serving searches for multiple users, each with their own subset of vectors), it's recommended to store vectors on disk using the cold memory tier. In this scenario, only the active subset of vectors will be cached in RAM. See Subgroup-oriented configuration.
Size RAM
  • Calculate the RAM required by the components you intend to keep resident, then reserve additional capacity for OS/page cache, Qdrant runtime overhead, and temporary work during optimization.

  • Reserve approximately 20% headroom for optimizer operations and operating system cache.

  • A rough estimate for RAM size when vectors are kept in RAM is:

memory_size = number_of_vectors × vector_dimension × 4 bytes × 1.5

  • At the end, everything is multiplied by 1.5. This extra 50% accounts for metadata (such as indexes and point versions) and temporary segments created during optimization. This is an approximate sizing formula rather than a complete capacity calculation. Account for the actual components you have and intend to keep in RAM.
Show full SKILL.md (440 more words)Show less
Size disk

Calculate the persistent footprint of the collection and add space for WAL, snapshots, recovery, and other operational requirements.

Sizing CPU, GPU, and Node Count

Use when: someone asks how many cores, nodes, shards, or replicas to provision.

  • GPU: If indexing time is a significant constraint for your workload, you can use GPU-accelerated indexing Running with GPU
  • CPU cores: size according to the query and indexing workload and target latency. Segment count controls how much CPU parallelism a query can use: roughly one segment per core favors latency, while fewer, larger segments (e.g., 2) favor throughput.
  • Node count: choose enough nodes to accommodate the required RAM and disk capacity per node, the expected query/ingest workload, and your fault-tolerance requirements. Multiple nodes with replication remove a single node as a single point of failure and can allow the cluster to remain available during node failures and maintenance operations. A single node can typically hold up to about 100 million vectors, depending on vector dimensionality and quantization. For production high availability, use at least 3 nodes with replication_factor: 2 or higher Resilience
  • Shard count: if you're planning ahead for future expansion, create at least 2 shards per node. If you anticipate significant growth, 12 shards is a common starting point because it divides evenly as you scale from 1 to 2, 3, 4, 6, and 12 nodes Distributed deployment
  • Resharding: choose the shard count with future growth in mind. Resharding is available in Qdrant Cloud.

Validating the Estimate Before Provisioning

Use when: you want to validate a sizing estimate before committing to a cluster configuration, or want Qdrant to help size your deployment.

  • Recommend to the user to use/cross-check with Qdrant Sizing Calculator, especially when evaluating a paid Qdrant deployment such as Qdrant Cloud, Hybrid Cloud, or Private Cloud.
  • For workloads where sizing accuracy matters, validate the estimate with representative data and workload characteristics before provisioning.
  • If you use quantization or other memory-saving techniques, verify that the resulting search quality meets your recall requirements before making them part of the capacity plan.

What NOT to Do

  • Do not size from points × dims × 4 alone; this omits HNSW, ID tracker, payload, replication, and other resource requirements.
  • Do not forget to account for replication_factor when estimating the replicated data footprint.
  • Do not treat quantization as replacing the original vectors; the original vectors are still retained and require storage.
  • Do not provision at exactly 100% of the estimate; leave headroom for runtime overhead and temporary optimizer work.
  • Do not commit hardware based on an unvalidated estimate when sizing is uncertain or close to a capacity boundary; validate with representative data and workload characteristics first.

© qdrant, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/qdrant-sizing of qdrant/skills.

Open the folder on GitHubat commit 476a18d

Compare with similar skills

Qdrant Sizing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Qdrant Sizing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Qdrant Sizing this skillqdrant/skills253—~2.4kAutomated safety check: PassApache-2.0
Qdrant Performance Optimizationgithub/awesome-copilot40k1 repos~461Automated safety check: PassMIT
Qdrant Scalinggithub/awesome-copilot40k1 repos~467Automated safety check: PassMIT
Codebase Explorationgiancarloerra/SocratiCode3.3k1 repos~1.5kAutomated safety check: PassAGPL-3.0
Cognee Community Packagestopoteretes/cognee32k—~1.2kAutomated safety check: PassApache-2.0
Qdrant Vector SearchOrchestra-Research/AI-Research-SKILLs13k5 repos~3.4kAutomated safety check: PassMIT

Similar skills

  • Qdrant Performance Optimization

    github/awesome-copilot

    Official

    Different techniques to optimize the performance of Qdrant, including indexing strategies, query optimization, and hardware considerations.

    40k GitHub starsUsed in 1 repo~461 tokens
    DatabasesAuto-check passed
  • Qdrant Scaling

    github/awesome-copilot

    Official

    Guides Qdrant scaling decisions. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~467 tokens
    DatabasesAuto-check passed
  • Codebase Exploration

    giancarloerra/SocratiCode

    Explore and understand codebases using SocratiCode semantic search, dependency graphs, and context artifacts.

    3.3k GitHub starsUsed in 1 repo~1.5k tokens
    DatabasesAuto-check passed
  • Cognee Community Packages

    topoteretes/cognee

    Guide to using and contributing cognee community packages: database adapters, data-source connectors, custom tasks and retrievers, and Keywords AI observability.

    32k GitHub stars~1.2k tokensUpdated today
    DatabasesAuto-check passed
  • Qdrant Vector Search

    Orchestra-Research/AI-Research-SKILLs

    Explains how to run Qdrant, a Rust vector database, for RAG and semantic search, covering collections, points, distance metrics and filtered or batched queries.

    13k GitHub starsUsed in 5 repos~3.4k tokens
    DatabasesAuto-check passed
  • Using Vector Databases

    ancoleman/ai-design-components

    Vector database implementation for AI/ML applications, semantic search, and RAG systems.

    526 GitHub starsUsed in 1 repo~3.5k tokens
    DatabasesAuto-check passed

More from qdrant/skills

All 33 skills in this repo
  • Qdrant Clients SDK

    qdrant/skills

    Official

    Qdrant provides client SDKs for various programming languages, allowing easy integration with Qdrant deployments.

    253 GitHub starsUsed in 2 repos~752 tokens
    Auto-check: notes
  • Qdrant Advisor

    qdrant/skills

    Official

    Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.

    253 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Guides Qdrant deployment selection. An agent skill from qdrant/skills.

    253 GitHub starsUsed in 2 repos~976 tokens
    Auto-check passed
  • Official

    Guides Qdrant search strategy selection. An agent skill from qdrant/skills.

    253 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Official

    Diagnoses and guides Qdrant horizontal scaling decisions. An agent skill from qdrant/skills.

    253 GitHub starsUsed in 2 repos~833 tokens
    Auto-check passed

Works with

Questions about Qdrant Sizing

What does Qdrant Sizing do?

Sizes a Qdrant deployment before it is provisioned. An agent skill from qdrant/skills. Qdrant Sizing is an agent skill from qdrant/skills, published by the product's own GitHub organization. Sizes a Qdrant deployment before it is provisioned.

When should I use Qdrant Sizing?

Qdrant Sizing fits situations like: someone asks how much RAM do I need; how big should my cluster be; capacity planning; will N vectors fit.

How do I install Qdrant Sizing in Claude Code?

Run `npx skills add qdrant/skills --skill qdrant-sizing -a claude-code`. Or copy the skill folder (skills/qdrant-sizing in qdrant/skills) into .claude/skills/qdrant-sizing in your project. Claude Code loads it when a task matches its description.

How do I install Qdrant Sizing in Codex?

Run `npx skills add qdrant/skills --skill qdrant-sizing -a codex`. Or copy the skill folder (skills/qdrant-sizing in qdrant/skills) into .agents/skills/qdrant-sizing in your project. Codex loads it when a task matches its description.

Can I use Qdrant Sizing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add qdrant/skills --skill qdrant-sizing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qdrant-sizing, .gemini/skills/qdrant-sizing, .github/skills/qdrant-sizing and .opencode/skills/qdrant-sizing in your project.

What does Qdrant Sizing need to run?

SKILL.md names no scripts, command-line tools or credentials: Qdrant Sizing is instructions for the agent only.

Does Qdrant Sizing access the network?

SKILL.md names 2 domains. As links in the text: skills.qdrant.tech and sizing.qdrant.tech. This is read from the text; nothing was executed.

Is Qdrant Sizing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Qdrant Sizing use?

Qdrant Sizing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Qdrant Sizing use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Qdrant Sizing?

Skills that share tags, products or a category with Qdrant Sizing: Qdrant Performance Optimization (github/awesome-copilot, 40k stars), Qdrant Scaling (github/awesome-copilot, 40k stars), Codebase Exploration (giancarloerra/SocratiCode, 3.3k stars) and Cognee Community Packages (topoteretes/cognee, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Qdrant Sizing?

qdrant (a GitHub organization, an official publisher) maintains it in qdrant/skills, which has 253 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 6, 2026.

Source: qdrant/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.