Official agent skill

Qdrant Memory Usage Optimization

by qdrant in qdrant/skills

Diagnoses and reduces Qdrant memory usage. An agent skill from qdrant/skills.

OfficialApache-2.0Auto-check passedDatabases

Install Qdrant Memory Usage Optimization

skills CLI
$ npx skills add qdrant/skills --skill qdrant-memory-usage-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install qdrant/skills qdrant-memory-usage-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/qdrant/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qdrant-performance-optimization/memory-usage-optimization .claude/skills/qdrant-memory-usage-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qdrant-memory-usage-optimization
GitHub stars
253
Used in
2 other repos
Token cost
~1.6k tokens
SKILL.md length
703 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses and reduces Qdrant memory usage. An agent skill from qdrant/skills.

  • Someone reports memory too high
  • SKILL.md covers Memory usage monitoring, How much memory is needed for… and How to minimize memory footprint
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • RAM keeps growing

What it does

Qdrant Memory Usage Optimization is an agent skill from qdrant/skills, published by the product's own GitHub organization. Diagnoses and reduces Qdrant memory usage. Use when someone reports 'memory too high', 'RAM keeps growing', 'node crashed', 'out of memory', 'memory leak', or asks 'why is memory usage so high?', 'how to reduce RAM?'. Also use when memory doesn't match calculations, quantization didn't help, or nodes crash during recovery.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Vector databases and Performance optimization. It works with Qdrant. The repository describes itself as: Agent skills for Qdrant vector search: scaling, performance optimization, search quality, monitoring, deployment, model migration, version upgrades, and SDK usage across Python…. The licence is Apache-2.0.

When your agent uses it

  • Someone reports memory too high
  • RAM keeps growing
  • Asks why is memory usage so high?
  • How to reduce RAM?

Example prompts

  • “memory too high”
  • “RAM keeps growing”
  • “node crashed”
  • “/qdrant-memory-usage-optimization”

What it can do on your machine

Read from SKILL.md and the folder at commit 476a18d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • skills.qdrant.tech

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Qdrant Memory Usage Optimization loads about 1.6k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 703 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from qdrant/skills at commit 476a18d, republished under its Apache-2.0 licence (© qdrant). 703 words, ~1,587 tokens.

Download SKILL.mdSave it as .claude/skills/qdrant-memory-usage-optimization/SKILL.md (or your agent's skills folder).
name
qdrant-memory-usage-optimization
description
Diagnoses and reduces Qdrant memory usage. Use when someone reports 'memory too high', 'RAM keeps growing', 'node crashed', 'out of memory', 'memory leak', or asks 'why is memory usage so high?', 'how to reduce RAM?'. Also use when memory doesn't match calculations, quantization didn't help, or nodes crash during recovery.

Understanding memory usage

Qdrant operates with two types of memory:

  • Resident memory (aka RSSAnon) - memory used for internal data structures like the ID tracker, plus components that stay fully in RAM. On Qdrant 1.19 or newer this is controlled per-component with memory: pinned (e.g. quantized vectors, payload indexes); see memory tier legacy settings for deployments on version 1.18 or older.

  • OS page cache - memory used for caching disk reads, which can be released when needed. Original vectors are normally stored in page cache, so the service won't crash if RAM is full, but performance may degrade. On Qdrant 1.19 or newer this corresponds to memory: cached (pre-warmed into page cache at startup) or memory: cold (lazy disk reads, not pre-warmed); on 1.18 or older it's controlled via the on_disk boolean on vectors, HNSW config, sparse vector index, and payload index. See Memory Tiers docs (available on 1.19+).

It is normal for the OS page cache to occupy all available RAM, but if resident memory is above 80% of total RAM, it is a sign of a problem.

Memory usage monitoring

  • Qdrant exposes memory usage through the /metrics endpoint. See Monitoring docs.
<!-- ToDo: Talk about memory usage of each components once API is available -->

How much memory is needed for Qdrant?

Optimal memory usage depends on the use case.

For a detailed breakdown of memory usage at large scale, see Large scale memory usage example.

Payload indexes and HNSW graph also require memory, along with vectors themselves, so it's important to consider them in calculations.

Additionally, Qdrant requires some extra memory for optimizations. During optimization, optimized segments are fully loaded into RAM, so it is important to leave enough headroom. The larger max_segment_size is, the more headroom is needed.

When to put HNSW index on disk

Putting frequently used components (such as HNSW index) on disk might cause significant performance degradation. On Qdrant 1.19 or newer this is set with memory: cold in hnsw_config; on 1.18 or older with hnsw_config.on_disk: true. There are some scenarios, however, when it can be a good option:

  • Deployments with low latency disks - local NVMe or similar.
  • Multi-tenant deployments, where only a subset of tenants is frequently accessed, so that only a fraction of data & index is loaded in RAM at a time.
  • For deployments with inline storage enabled.
Show full SKILL.md (319 more words)Show less

How to minimize memory footprint

The main challenge is to put on disk those parts of data, which are rarely accessed. Here are the main techniques to achieve that:

  • Use quantization to store only compressed vectors in RAM Quantization docs

  • Use float16 or uint8 datatypes to reduce memory usage of vectors by 2x or 4x respectively, with some tradeoff in precision. On Qdrant 1.19 or newer, the turbo4 datatype (TurboQuant-based, 4 bits/dimension, dense vectors only) reduces memory by ~8x, and can be paired with 1-bit quantization for cheaper rescoring than pairing 1-bit quantization with full-precision vectors. Read more about vector datatypes in documentation

  • Leverage Matryoshka Representation Learning (MRL) to store only small vectors in RAM while keeping large vectors on disk. Examples of how to use MRL with Qdrant Cloud inference: MRL docs

  • For multi-tenant deployments with small tenants, vectors might be stored on disk because the same tenant's data is stored together Multitenancy docs

  • For deployments with fast local storage and relatively low requirements for search throughput, it may be possible to store all components of vector store on disk. Read more about the performance implications of on-disk storage in the article

  • For low RAM environments, enable async I/O (io_uring) for concurrent disk reads, which can significantly improve performance of on-disk storage: storage.performance.io_uring: auto on Qdrant 1.19 or newer (applies to every cold structure), async_scorer: true on 1.18 or older (vector rescoring only). Requires Linux with a kernel that supports io_uring Async I/O

  • Keep payloads on disk: memory: cold is the default on Qdrant 1.19 or newer; on 1.18 or older, set on_disk_payload: true Default tiers

  • Configure payload indexes to be stored on disk: memory: cold on Qdrant 1.19 or newer, on_disk: true on 1.18 or older docs

  • Configure sparse vectors to be stored on disk: memory: cold on the sparse vector index on Qdrant 1.19 or newer (defaults to pinned), on_disk: true on 1.18 or older docs

© qdrant, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/qdrant-performance-optimization/memory-usage-optimization of qdrant/skills.

Open the folder on GitHubat commit 476a18d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in qdrant/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Qdrant Memory Usage Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Qdrant Memory Usage Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Qdrant Memory Usage Optimization this skillqdrant/skills2532 repos~1.6kAutomated safety check: PassApache-2.0
Qdrant Performance Optimizationgithub/awesome-copilot40k1 repos~461Automated safety check: PassMIT
Codebase Explorationgiancarloerra/SocratiCode3.3k1 repos~1.5kAutomated safety check: PassAGPL-3.0
Qdrant Vector SearchOrchestra-Research/AI-Research-SKILLs13k5 repos~3.4kAutomated safety check: PassMIT
Using Vector Databasesancoleman/ai-design-components5261 repos~3.5kAutomated safety check: PassMIT
Qdrant Search Strategiesgithub/awesome-copilot40k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Qdrant Performance Optimization

    github/awesome-copilot

    Official

    Different techniques to optimize the performance of Qdrant, including indexing strategies, query optimization, and hardware considerations.

    40k GitHub starsUsed in 1 repo~461 tokens
    DatabasesAuto-check passed
  • Codebase Exploration

    giancarloerra/SocratiCode

    Explore and understand codebases using SocratiCode semantic search, dependency graphs, and context artifacts.

    3.3k GitHub starsUsed in 1 repo~1.5k tokens
    DatabasesAuto-check passed
  • Qdrant Vector Search

    Orchestra-Research/AI-Research-SKILLs

    Explains how to run Qdrant, a Rust vector database, for RAG and semantic search, covering collections, points, distance metrics and filtered or batched queries.

    13k GitHub starsUsed in 5 repos~3.4k tokens
    DatabasesAuto-check passed
  • Using Vector Databases

    ancoleman/ai-design-components

    Vector database implementation for AI/ML applications, semantic search, and RAG systems.

    526 GitHub starsUsed in 1 repo~3.5k tokens
    DatabasesAuto-check passed
  • Qdrant Search Strategies

    github/awesome-copilot

    Official

    Guides Qdrant search strategy selection. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~1.7k tokens
    DatabasesAuto-check passed
  • Qdrant

    giuseppe-trisciuoglio/developer-kit

    Provides Qdrant vector database integration patterns with LangChain4j.

    355 GitHub starsUsed in 1 repo~1.6k tokens
    DatabasesAuto-check: notes

More from qdrant/skills

All 33 skills in this repo
  • Qdrant Clients SDK

    qdrant/skills

    Official

    Qdrant provides client SDKs for various programming languages, allowing easy integration with Qdrant deployments.

    253 GitHub starsUsed in 2 repos~752 tokens
    Auto-check: notes
  • Qdrant Advisor

    qdrant/skills

    Official

    Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.

    253 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Official

    Guides Qdrant deployment selection. An agent skill from qdrant/skills.

    253 GitHub starsUsed in 2 repos~976 tokens
    Auto-check passed
  • Official

    Guides Qdrant search strategy selection. An agent skill from qdrant/skills.

    253 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Official

    Diagnoses and guides Qdrant horizontal scaling decisions. An agent skill from qdrant/skills.

    253 GitHub starsUsed in 2 repos~833 tokens
    Auto-check passed

Works with

Questions about Qdrant Memory Usage Optimization

What does Qdrant Memory Usage Optimization do?

Diagnoses and reduces Qdrant memory usage. An agent skill from qdrant/skills. Qdrant Memory Usage Optimization is an agent skill from qdrant/skills, published by the product's own GitHub organization. Diagnoses and reduces Qdrant memory usage.

When should I use Qdrant Memory Usage Optimization?

Qdrant Memory Usage Optimization fits situations like: someone reports memory too high; RAM keeps growing; asks why is memory usage so high?; how to reduce RAM?.

How do I install Qdrant Memory Usage Optimization in Claude Code?

Run `npx skills add qdrant/skills --skill qdrant-memory-usage-optimization -a claude-code`. Or copy the skill folder (skills/qdrant-performance-optimization/memory-usage-optimization in qdrant/skills) into .claude/skills/qdrant-memory-usage-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Qdrant Memory Usage Optimization in Codex?

Run `npx skills add qdrant/skills --skill qdrant-memory-usage-optimization -a codex`. Or copy the skill folder (skills/qdrant-performance-optimization/memory-usage-optimization in qdrant/skills) into .agents/skills/qdrant-memory-usage-optimization in your project. Codex loads it when a task matches its description.

Can I use Qdrant Memory Usage Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add qdrant/skills --skill qdrant-memory-usage-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qdrant-memory-usage-optimization, .gemini/skills/qdrant-memory-usage-optimization, .github/skills/qdrant-memory-usage-optimization and .opencode/skills/qdrant-memory-usage-optimization in your project.

What does Qdrant Memory Usage Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Qdrant Memory Usage Optimization is instructions for the agent only.

Does Qdrant Memory Usage Optimization access the network?

SKILL.md names 1 domain. As links in the text: skills.qdrant.tech. This is read from the text; nothing was executed.

Is Qdrant Memory Usage Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Qdrant Memory Usage Optimization use?

Qdrant Memory Usage Optimization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Qdrant Memory Usage Optimization use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Qdrant Memory Usage Optimization?

Skills that share tags, products or a category with Qdrant Memory Usage Optimization: Qdrant Performance Optimization (github/awesome-copilot, 40k stars), Codebase Exploration (giancarloerra/SocratiCode, 3.3k stars), Qdrant Vector Search (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Using Vector Databases (ancoleman/ai-design-components, 526 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Qdrant Memory Usage Optimization?

qdrant (a GitHub organization, an official publisher) maintains it in qdrant/skills, which has 253 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 6, 2026.

Source: qdrant/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.