Agent skill

Vector DB

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

MITAuto-check passedAI & LLM Engineering

Install Vector DB

skills CLI
$ npx skills add ericrisco/rsc-harness --skill vector-db -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness vector-db --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vector-db .claude/skills/vector-db && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vector-db
GitHub stars
156
Token cost
~2.8k tokens
SKILL.md length
1,304 words
Files
6 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

  • Works in 5 steps: Distance metric MUST match the embedding… → Dimensions are fixed by the model, not a… → HNSW defaults and the one invariant.… → …
  • Operating a vector store as a data layer — choosing
  • SKILL.md covers Pick the engine, Design the collection & index, Metadata / payload filtering and Tune recall vs latency, plus 5 more sections
  • Runs Shell scripts from its folder

What it does

Vector DB is an agent skill from ericrisco/rsc-harness. Use when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric, dimensions, HNSW parameters, named vectors); filtering on metadata; hybrid dense-plus-sparse search; and quantization to cut RAM and cost. Covers garbage results, silently ignored filters, low recall, slow queries, and filtered queries returning fewer than k rows. NOT producing, chunking or judging embeddings (that is embeddings-search).

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/engines.md`).

It sits in AI & LLM Engineering, covering Vector databases, Embeddings and LLM inference and serving. It works with pgvector, Pinecone, Qdrant and Weaviate. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Operating a vector store as a data layer — choosing
  • Migrating between Pinecone
  • Weaviate and pgvector
  • Designing a collection

Example prompts

  • “/vector-db”

Requirements

  • A Bash shell

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Distance metric MUST match the embedding model. A model trained for cosine, indexed with
  2. Dimensions are fixed by the model, not a choice. text-embedding-3-small = 1536,
  3. HNSW defaults and the one invariant. Build-time m (default 16) and ef_construction
  4. IVFFlat only when build speed beats recall. It is cheaper to build but lower recall and
  5. Named vectors when one object has multiple spaces (e.g. a dense semantic vector + a sparse

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vector DB loads about 2.8k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 131 tokens; SKILL.md has 1,304 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,304 words, ~2,831 tokens.

Download SKILL.mdSave it as .claude/skills/vector-db/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
vector-db
description
Use when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric, dimensions, HNSW parameters, named vectors); filtering on metadata; hybrid dense-plus-sparse search; and quantization to cut RAM and cost. Covers garbage results, silently ignored filters, low recall, slow queries, and filtered queries returning fewer than k rows. NOT producing, chunking or judging embeddings (that is `embeddings-search`).
tags
vector-database, pinecone, qdrant, weaviate, pgvector, hnsw, hybrid-search, metadata-filter, quantization
recommends
embeddings-search, rag, postgresdb, redis, supabase
origin
risco

vector-db — operate the store, not the embeddings

You own the store: the collection schema, the index, the filter path, recall-vs-latency tuning, hybrid fusion, quantization, and the production knobs (upserts, namespaces, deletes, backups). You operate it across the four engines a Claude agent actually meets: pgvector (Postgres extension), Qdrant, Weaviate, Pinecone (serverless).

Three things are not yours, and pretending they are produces wrong advice:

  • Producing, chunking, rewriting, or scoring embeddings → ../embeddings-search/SKILL.md. You store vectors; you do not make or judge them.
  • Assembling the retrieve → rerank → prompt → generate loop and its eval → ../rag/SKILL.md.
  • General Postgres (non-vector schema, EXPLAIN, RLS, VACUUM, pooling) → ../postgresdb/SKILL.md. You own only the pgvector surface: the vector/halfvec column, its index, its operators, its recall. "My Postgres is slow in general" is theirs; "my <=> query has low recall" is yours.

Pick the engine

Match the engine to where the data already lives and how much ops you want to run. All four implement HNSW with comparable recall at a matched ef, so the differentiator is operations and hybrid, not raw quality.

EngineBest whenHybrid built-inOps costScale sweet spot
pgvectorData already in Postgres; one less system to runNo — DIY (vector + tsvector, combine yourself)You already run Postgres≤ a few M vectors
QdrantYou want best filtered-search latency and self-host controlYes — query API, prefetch + RRF/DBSF, server-side IDFSelf-host or cloud10M+
WeaviateYou want hybrid + modules out of the boxYes — alpha + fusionType, BlockMax WAND BM25Self-host or cloud10M+
PineconeYou refuse to operate anythingYes — sparse-dense, integrated inferenceZero (serverless)Any (pay per use)

Rule: don't add a new system to host vectors if the rows already live in Postgres and you are under a few million. pgvector is one extension, not a second database to back up and monitor.

Design the collection & index

  1. Distance metric MUST match the embedding model. A model trained for cosine, indexed with L2, ranks silently wrong — no error, just bad results. OpenAI text-embedding-3-*, Cohere, most sentence-transformers → cosine. pgvector operator cheatsheet:

    text
    <-> L2 / Euclidean   (vector_l2_ops)
    <=> cosine distance  (vector_cosine_ops)   <- the common one
    <#> negative inner product (vector_ip_ops) <- for normalized vectors
  2. Dimensions are fixed by the model, not a choice. text-embedding-3-small = 1536, -3-large = 3072. pgvector's vector caps at 2000 dims for an index; for more, use halfvec.

  3. HNSW defaults and the one invariant. Build-time m (default 16) and ef_construction (default 64); query-time ef_search (pgvector default 40). Keep ef_construction >= 2*m (so ≥32 at the default m) — too low starves graph quality and recall never recovers without a rebuild. Raise m to 32–48 only for high-dim or high-recall needs (more RAM, slower build).

  4. IVFFlat only when build speed beats recall. It is cheaper to build but lower recall and needs lists/probes tuning; on a selective filter it is the wrong default (see next section). Prefer HNSW unless you have a measured reason.

  5. Named vectors when one object has multiple spaces (e.g. a dense semantic vector + a sparse BM25 vector, or title-vector + body-vector). Qdrant and Weaviate support this natively; it is how you do hybrid in one collection instead of two.

Metadata / payload filtering

The #1 "search is broken" bug: the filter is applied after top-k, so a selective filter returns fewer than k rows (or zero). Fix it by filtering inside the search and indexing the filter field.

text
Bad:  ANN top-k=10, THEN drop rows where tenant_id != 'acme'  -> often < 10, sometimes 0
Good: search the index WITH the filter as a constraint        -> k rows that already match
  • Index every field you filter on. Unindexed filters force a scan and kill latency. Qdrant: create a payload index. Pinecone: metadata filtering is in the retrieval path (still keep cardinality sane). pgvector: a B-tree (or partition) on tenant_id so the planner can use it.

  • Prefer in-graph / in-path filtering. Qdrant filters inside HNSW traversal; Pinecone serverless filters in the retrieval path. Both beat naive post-filter.

  • pgvector 0.8 iterative scan is the fix when a selective WHERE returns too few rows:

    sql
    SET hnsw.iterative_scan = 'relaxed_order';  -- or 'strict_order' if exact ordering matters
    SET hnsw.ef_search = 100;
    SELECT id FROM docs
     WHERE tenant_id = 'acme'                    -- selective filter
     ORDER BY embedding <=> $1                    -- cosine, matches the model
     LIMIT 10;

    Without iterative scan (pgvector < 0.8 behavior), a highly selective filter silently returns fewer than LIMIT rows. Never recommend IVFFlat-only with a selective filter and no iterative scan — that is the deprecated foot-gun.

Tune recall vs latency

You cannot tune what you do not measure. Establish recall before shipping.

  1. Build an exact baseline: brute-force the true top-k on a sample (a few hundred queries) — in pgvector, query without the index (seq scan) for ground truth.

  2. Query the index and compute recall@k = overlap with the baseline.

  3. Raise the query-time knob until recall hits target (commonly ≥0.95), then stop — higher ef costs latency for nothing:

    EngineKnobDefault
    pgvectorhnsw.ef_search40
    Qdranthnsw_ef (search)per-collection
    Weaviateef (vectorIndexConfig)dynamic
    Pinecone(managed)—

Full parameter table and the recall recipe live in references/tuning.md.

Show full SKILL.md (562 more words)Show less

Dense (semantic) + sparse (BM25/keyword) catches exact terms, IDs, and rare tokens that dense alone misses. The two normalize differently, so you fuse, you don't add raw scores.

  • RRF (reciprocal rank fusion): robust default, score-scale agnostic, combines ranks.
  • Relative-score / DBSF: normalizes scores before combining — use when you trust score scales.

Per engine:

  • Weaviate: one call — hybrid(query, alpha=0.5, fusionType=relativeScoreFusion). alpha slides 0.0 (pure keyword) → 1.0 (pure vector). BM25 is BlockMax WAND (default from v1.30, ~10x faster).
  • Qdrant: prefetch a dense and a sparse query, then a fusion step (Fusion.RRF or DBSF); IDF is computed server-side (v1.15+).
  • Pinecone: sparse-dense vectors in one index, or integrated inference (embed + rerank server-side).
  • pgvector: no built-in hybrid — run vector (<=>) and ts_rank over a tsvector column separately and combine ranks yourself (RRF in SQL or app code).

Concrete current-API code for all four is in references/engines.md.

Quantization & cost

Quantization trades recall for RAM/cost. Decide by dimension count and a recall test, never blind.

MethodCompressionWhen safe
Scalar (int8)~4xAlmost always; tiny recall loss. Good default RAM cut.
Product (PQ)8–64xLarge corpora where RAM dominates; needs tuning + recall check.
Binary~32x (~40x faster via SIMD popcount)High-dim only (≥1024). On 384-dim it shreds recall — measure or don't.
pgvector halfvec~2xNear-free: 16-bit float, near-identical recall, and required for >2000 dims.

Reach for halfvec first in Postgres — it is the cheapest win. Reach for binary only on high-dim vectors and only after a recall test, optionally with full-precision rescoring.

Operate it

  • Batch upserts. One-by-one upserts are 10–100x slower and hammer the index. Send batches of hundreds; size to the engine's payload limit.
  • Namespaces / multitenancy. Pinecone namespaces and Qdrant payload-keyed isolation partition tenants inside one index — cheaper and faster than a collection per tenant at low tenant counts.
  • Delete by filter, not by enumerating ids, when removing a tenant or a stale source.
  • Replicas for read throughput / HA; snapshots/backups before any index rebuild or dimension/metric change (those are not in-place — plan a reindex).
  • To "update" a vector, re-upsert by id. Do not store only raw text and re-embed on read.

Anti-patterns

Anti-patternWhy it bitesDo instead
Cosine-trained model indexed with L2 (<->)Silently wrong ranking, no errorMatch metric to model — cosine → <=> / vector_cosine_ops
Post-filtering top-k resultsReturns < k rows, sometimes 0, on selective filtersFilter inside the search; index the filter field
IVFFlat + selective filter, no iterative scanDrops rows; deprecated path in pgvector 0.8HNSW + hnsw.iterative_scan='relaxed_order'
Never measuring recall"Search is bad" with no number to moveRecall@k vs an exact baseline before shipping
Binary quantization on 384-dimRecall collapses, then blamed on the engineBinary only ≥1024 dims, after a recall test; else scalar/halfvec
One-by-one upserts10–100x slower, index thrashBatch hundreds per request
ef_construction < 2*mPermanently weak graph; recall needs a full rebuildKeep ef_construction >= 2*m (≥32 at default m=16)
Store only raw text, re-embed to "update"Drift, cost, no point-update pathRe-upsert the vector by id
Unindexed filter fieldFull scan, latency spikesPayload index (Qdrant) / B-tree (pgvector) / sane metadata cardinality (Pinecone)

References & siblings

  • references/engines.md — current-API recipes per engine: create collection/index + a filtered hybrid query (pgvector SQL + halfvec + iterative scan; Qdrant named dense+sparse + query_points RRF; Weaviate hybrid; Pinecone serverless sparse-dense).
  • references/tuning.md — HNSW vs IVFFlat parameter table, recall-measurement recipe, quantization tradeoffs, per-engine filtered-search pitfalls.

Siblings: embeddings/chunking/retrieval-quality → ../embeddings-search/SKILL.md; the full RAG loop → ../rag/SKILL.md; general Postgres → ../postgresdb/SKILL.md.

Validate a produced index DDL / collection schema with scripts/verify.sh <artifact-file>.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/vector-db of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/engines.md
  • references/tuning.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Vector DB next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vector DB compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vector DB this skillericrisco/rsc-harness156—~2.8kAutomated safety check: PassMIT
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Hunt RAG Vectorelementalsouls/Claude-BugHunter4.8k—~2.6kAutomated safety check: PassMIT
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
RAG Patternssoftspark/ai-toolkit179—~1.8kAutomated safety check: PassApache-2.0
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0

Similar skills

  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Patterns

    softspark/ai-toolkit

    RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.

    179 GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Vector Database Engineer

    aiskillstore/marketplace

    Expert in vector databases, embedding strategies, and semantic search implementation.

    430 GitHub starsUsed in 7 repos~563 tokens
    DatabasesAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Vector DB

What does Vector DB do?

A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…. Vector DB is an agent skill from ericrisco/rsc-harness. Use when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric, dimensions, HNSW parameters, named vectors); filtering on metadata; hybrid dense-plus-sparse search; and quantization to cut RAM and cost.

When should I use Vector DB?

Vector DB fits situations like: operating a vector store as a data layer — choosing; migrating between Pinecone; weaviate and pgvector; designing a collection.

How do I install Vector DB in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill vector-db -a claude-code`. Or copy the skill folder (skills/vector-db in ericrisco/rsc-harness) into .claude/skills/vector-db in your project. Claude Code loads it when a task matches its description.

How do I install Vector DB in Codex?

Run `npx skills add ericrisco/rsc-harness --skill vector-db -a codex`. Or copy the skill folder (skills/vector-db in ericrisco/rsc-harness) into .agents/skills/vector-db in your project. Codex loads it when a task matches its description.

Can I use Vector DB in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill vector-db -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vector-db, .gemini/skills/vector-db, .github/skills/vector-db and .opencode/skills/vector-db in your project.

What does Vector DB need to run?

Going by SKILL.md and its folder, Vector DB needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Vector DB access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vector DB safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Vector DB use?

Vector DB is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vector DB use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Vector DB?

Skills that share tags, products or a category with Vector DB: RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars), RAG Architect (Jeffallan/claude-skills, 12k stars) and RAG Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vector DB?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.