Agent skill

Vector Index Tuning

by wshobson in wshobson/agents

Tune vector indexes for latency, recall and memory: pick an index type by data size, adjust HNSW parameters and choose a quantization level.

MITAuto-check passedAI & LLM Engineering

Install Vector Index Tuning

skills CLI
$ npx skills add wshobson/agents --skill vector-index-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents vector-index-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-application-dev/skills/vector-index-tuning .claude/skills/vector-index-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vector-index-tuning
GitHub stars
40k
Used in
9 other repos
Token cost
~557 tokens
SKILL.md length
165 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Tune vector indexes for latency, recall and memory: pick an index type by data size, adjust HNSW parameters and choose a quantization level.

  • Works in 3 steps: Index Type Selection → HNSW Parameters → Quantization Types
  • Reducing search latency on a large vector index
  • SKILL.md covers When to Use This Skill, Core Concepts, Templates and detailed worked… and Best Practices
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill is a tuning guide for vector search at production scale. It maps data size to a recommended index type, starting with flat exact search for fewer than 10K vectors, and tabulates the HNSW parameters M, efConstruction and efSearch with their defaults of 16, 100 and 50 and how raising each trades recall against memory, build time or search speed.

A quantization section compares full precision FP32, half precision FP16 and INT8 scalar storage by bytes per dimension. Advice includes benchmarking with real queries, watching recall for drift, starting with defaults, using tiered storage, warming cold indexes and planning for reindexing. Concrete templates are kept in references/details.md.

When your agent uses it

  • Reducing search latency on a large vector index
  • Tuning HNSW parameters to balance recall and speed
  • Cutting memory use with quantization
  • Planning an index for a collection that will grow to billions of vectors

Example prompts

  • “Our vector search misses relevant results under load. Suggest efSearch and M settings.”
  • “Estimate the memory savings if we move our embeddings from FP32 to INT8.”
  • “Pick an index type for a collection of a few thousand vectors and tell me when to change it.”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Index Type Selection
  2. HNSW Parameters
  3. Quantization Types

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vector Index Tuning loads about 557 tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 165 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~557
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 165 words, ~557 tokens.

Download SKILL.mdSave it as .claude/skills/vector-index-tuning/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vector-index-tuning
description
Optimize vector index performance for latency, recall, and memory. Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.

Vector Index Tuning

Guide to optimizing vector indexes for production performance.

When to Use This Skill

  • Tuning HNSW parameters
  • Implementing quantization
  • Optimizing memory usage
  • Reducing search latency
  • Balancing recall vs speed
  • Scaling to billions of vectors

Core Concepts

1. Index Type Selection
Data Size           Recommended Index
────────────────────────────────────────
< 10K vectors  →    Flat (exact search)
10K - 1M       →    HNSW
1M - 100M      →    HNSW + Quantization
> 100M         →    IVF + PQ or DiskANN
2. HNSW Parameters
ParameterDefaultEffect
M16Connections per node, ↑ = better recall, more memory
efConstruction100Build quality, ↑ = better index, slower build
efSearch50Search quality, ↑ = better recall, slower search
3. Quantization Types
Full Precision (FP32): 4 bytes × dimensions
Half Precision (FP16): 2 bytes × dimensions
INT8 Scalar:           1 byte × dimensions
Product Quantization:  ~32-64 bytes total
Binary:                dimensions/8 bytes

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's
  • Benchmark with real queries - Synthetic may not represent production
  • Monitor recall continuously - Can degrade with data drift
  • Start with defaults - Tune only when needed
  • Use quantization - Significant memory savings
  • Consider tiered storage - Hot/cold data separation
Don'ts
  • Don't over-optimize early - Profile first
  • Don't ignore build time - Index updates have cost
  • Don't forget reindexing - Plan for maintenance
  • Don't skip warming - Cold indexes are slow

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/llm-application-dev/skills/vector-index-tuning of wshobson/agents.

  • SKILL.md
  • references/details.md

Open the folder on GitHubat commit 46891e7

Used in 9 other repositories

We found 19 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 9 other GitHub owners. This page covers the copy in wshobson/agents, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Vector Index Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vector Index Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vector Index Tuning this skillwshobson/agents40k9 repos~557Automated safety check: PassMIT
Agentdb Performance Optimizationaiskillstore/marketplace4306 repos~3kAutomated safety check: PassNone
Qdrant Indexing Performance Optimizationqdrant/skills2542 repos~1.2kAutomated safety check: PassApache-2.0
Qdrant Memory Usage Optimizationqdrant/skills2542 repos~1.6kAutomated safety check: PassApache-2.0
Qdrant Performance Optimizationgithub/awesome-copilot40k1 repos~461Automated safety check: PassMIT
Qdrant Performance Optimizationqdrant/skills254—~456Automated safety check: PassApache-2.0

Similar skills

  • Agentdb Performance Optimization

    aiskillstore/marketplace

    Optimize AgentDB performance with quantization (4-32x memory reduction), HNSW indexing (150x faster search), caching, and batch operations.

    430 GitHub starsUsed in 6 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Diagnoses and reduces Qdrant memory usage. An agent skill from qdrant/skills.

    254 GitHub starsUsed in 2 repos~1.6k tokens
    DatabasesAuto-check passed
  • Qdrant Performance Optimization

    github/awesome-copilot

    Official

    Different techniques to optimize the performance of Qdrant, including indexing strategies, query optimization, and hardware considerations.

    40k GitHub starsUsed in 1 repo~461 tokens
    DatabasesAuto-check passed
  • Official

    Navigation hub linking sub-skills for proactive Qdrant tuning: search speed, indexing performance, and memory usage optimization.

    254 GitHub stars~456 tokensUpdated yesterday
    DatabasesAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed

Questions about Vector Index Tuning

What does Vector Index Tuning do?

Tune vector indexes for latency, recall and memory: pick an index type by data size, adjust HNSW parameters and choose a quantization level. The skill is a tuning guide for vector search at production scale. It maps data size to a recommended index type, starting with flat exact search for fewer than 10K vectors, and tabulates the HNSW parameters M, efConstruction and efSearch with their defaults of 16, 100 and 50 and how raising each trades recall against memory, build time or search speed.

When should I use Vector Index Tuning?

Vector Index Tuning fits situations like: reducing search latency on a large vector index; tuning HNSW parameters to balance recall and speed; cutting memory use with quantization; planning an index for a collection that will grow to billions of vectors.

How do I install Vector Index Tuning in Claude Code?

Run `npx skills add wshobson/agents --skill vector-index-tuning -a claude-code`. Or copy the skill folder (plugins/llm-application-dev/skills/vector-index-tuning in wshobson/agents) into .claude/skills/vector-index-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Vector Index Tuning in Codex?

Run `npx skills add wshobson/agents --skill vector-index-tuning -a codex`. Or copy the skill folder (plugins/llm-application-dev/skills/vector-index-tuning in wshobson/agents) into .agents/skills/vector-index-tuning in your project. Codex loads it when a task matches its description.

Can I use Vector Index Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill vector-index-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vector-index-tuning, .gemini/skills/vector-index-tuning, .github/skills/vector-index-tuning and .opencode/skills/vector-index-tuning in your project.

What does Vector Index Tuning need to run?

SKILL.md names no scripts, command-line tools or credentials: Vector Index Tuning is instructions for the agent only.

Does Vector Index Tuning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vector Index Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vector Index Tuning use?

Vector Index Tuning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vector Index Tuning use?

About 557 tokens (SKILL.md is roughly 2.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Vector Index Tuning?

Skills that share tags, products or a category with Vector Index Tuning: Agentdb Performance Optimization (aiskillstore/marketplace, 430 stars), Qdrant Indexing Performance Optimization (qdrant/skills, 254 stars), Qdrant Memory Usage Optimization (qdrant/skills, 254 stars) and Qdrant Performance Optimization (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vector Index Tuning?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,305 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.