RAG Implementation
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…
$ npx skills add ericrisco/rsc-harness --skill vector-db -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness vector-db --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vector-db .claude/skills/vector-db && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vector-db" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vector-db into .claude/skills/vector-db/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vector-db", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/vector-dbType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill vector-db -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness vector-db --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/vector-db .agents/skills/vector-db && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vector-db" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vector-db into .agents/skills/vector-db/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vector-db", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill vector-db -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness vector-db --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/vector-db .cursor/skills/vector-db && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vector-db" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vector-db into .cursor/skills/vector-db/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vector-db", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/vector-db--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill vector-db -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness vector-db --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/vector-db .gemini/skills/vector-db && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vector-db" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vector-db into .gemini/skills/vector-db/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vector-db", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness vector-dbInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill vector-db -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/vector-db .github/skills/vector-db && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vector-db" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vector-db into .github/skills/vector-db/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vector-db", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill vector-db -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness vector-db --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/vector-db .opencode/skills/vector-db && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vector-db" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vector-db into .opencode/skills/vector-db/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vector-db", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vector-dbA skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…
Vector DB is an agent skill from ericrisco/rsc-harness. Use when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric, dimensions, HNSW parameters, named vectors); filtering on metadata; hybrid dense-plus-sparse search; and quantization to cut RAM and cost. Covers garbage results, silently ignored filters, low recall, slow queries, and filtered queries returning fewer than k rows. NOT producing, chunking or judging embeddings (that is embeddings-search).
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/engines.md`).
It sits in AI & LLM Engineering, covering Vector databases, Embeddings and LLM inference and serving. It works with pgvector, Pinecone, Qdrant and Weaviate. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vector DB loads about 2.8k tokens when it runs, and up to ~4.9k if it reads all its reference files. Until then it costs about 131 tokens; SKILL.md has 1,304 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,304 words, ~2,831 tokens.
.claude/skills/vector-db/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.You own the store: the collection schema, the index, the filter path, recall-vs-latency tuning, hybrid fusion, quantization, and the production knobs (upserts, namespaces, deletes, backups). You operate it across the four engines a Claude agent actually meets: pgvector (Postgres extension), Qdrant, Weaviate, Pinecone (serverless).
Three things are not yours, and pretending they are produces wrong advice:
../embeddings-search/SKILL.md. You
store vectors; you do not make or judge them.../rag/SKILL.md.../postgresdb/SKILL.md.
You own only the pgvector surface: the vector/halfvec column, its index, its operators,
its recall. "My Postgres is slow in general" is theirs; "my <=> query has low recall" is yours.Match the engine to where the data already lives and how much ops you want to run. All four
implement HNSW with comparable recall at a matched ef, so the differentiator is operations and
hybrid, not raw quality.
| Engine | Best when | Hybrid built-in | Ops cost | Scale sweet spot |
|---|---|---|---|---|
| pgvector | Data already in Postgres; one less system to run | No — DIY (vector + tsvector, combine yourself) | You already run Postgres | ≤ a few M vectors |
| Qdrant | You want best filtered-search latency and self-host control | Yes — query API, prefetch + RRF/DBSF, server-side IDF | Self-host or cloud | 10M+ |
| Weaviate | You want hybrid + modules out of the box | Yes — alpha + fusionType, BlockMax WAND BM25 | Self-host or cloud | 10M+ |
| Pinecone | You refuse to operate anything | Yes — sparse-dense, integrated inference | Zero (serverless) | Any (pay per use) |
Rule: don't add a new system to host vectors if the rows already live in Postgres and you are under a few million. pgvector is one extension, not a second database to back up and monitor.
Distance metric MUST match the embedding model. A model trained for cosine, indexed with
L2, ranks silently wrong — no error, just bad results. OpenAI text-embedding-3-*, Cohere,
most sentence-transformers → cosine. pgvector operator cheatsheet:
<-> L2 / Euclidean (vector_l2_ops)
<=> cosine distance (vector_cosine_ops) <- the common one
<#> negative inner product (vector_ip_ops) <- for normalized vectorsDimensions are fixed by the model, not a choice. text-embedding-3-small = 1536,
-3-large = 3072. pgvector's vector caps at 2000 dims for an index; for more, use halfvec.
HNSW defaults and the one invariant. Build-time m (default 16) and ef_construction
(default 64); query-time ef_search (pgvector default 40). Keep ef_construction >= 2*m
(so ≥32 at the default m) — too low starves graph quality and recall never recovers without
a rebuild. Raise m to 32–48 only for high-dim or high-recall needs (more RAM, slower build).
IVFFlat only when build speed beats recall. It is cheaper to build but lower recall and
needs lists/probes tuning; on a selective filter it is the wrong default (see next section).
Prefer HNSW unless you have a measured reason.
Named vectors when one object has multiple spaces (e.g. a dense semantic vector + a sparse BM25 vector, or title-vector + body-vector). Qdrant and Weaviate support this natively; it is how you do hybrid in one collection instead of two.
The #1 "search is broken" bug: the filter is applied after top-k, so a selective filter
returns fewer than k rows (or zero). Fix it by filtering inside the search and indexing the
filter field.
Bad: ANN top-k=10, THEN drop rows where tenant_id != 'acme' -> often < 10, sometimes 0
Good: search the index WITH the filter as a constraint -> k rows that already matchIndex every field you filter on. Unindexed filters force a scan and kill latency. Qdrant:
create a payload index. Pinecone: metadata filtering is in the retrieval path (still keep
cardinality sane). pgvector: a B-tree (or partition) on tenant_id so the planner can use it.
Prefer in-graph / in-path filtering. Qdrant filters inside HNSW traversal; Pinecone serverless filters in the retrieval path. Both beat naive post-filter.
pgvector 0.8 iterative scan is the fix when a selective WHERE returns too few rows:
SET hnsw.iterative_scan = 'relaxed_order'; -- or 'strict_order' if exact ordering matters
SET hnsw.ef_search = 100;
SELECT id FROM docs
WHERE tenant_id = 'acme' -- selective filter
ORDER BY embedding <=> $1 -- cosine, matches the model
LIMIT 10;Without iterative scan (pgvector < 0.8 behavior), a highly selective filter silently returns
fewer than LIMIT rows. Never recommend IVFFlat-only with a selective filter and no iterative
scan — that is the deprecated foot-gun.
You cannot tune what you do not measure. Establish recall before shipping.
Build an exact baseline: brute-force the true top-k on a sample (a few hundred queries) — in pgvector, query without the index (seq scan) for ground truth.
Query the index and compute recall@k = overlap with the baseline.
Raise the query-time knob until recall hits target (commonly ≥0.95), then stop — higher ef
costs latency for nothing:
| Engine | Knob | Default |
|---|---|---|
| pgvector | hnsw.ef_search | 40 |
| Qdrant | hnsw_ef (search) | per-collection |
| Weaviate | ef (vectorIndexConfig) | dynamic |
| Pinecone | (managed) | — |
Full parameter table and the recall recipe live in references/tuning.md.
Dense (semantic) + sparse (BM25/keyword) catches exact terms, IDs, and rare tokens that dense alone misses. The two normalize differently, so you fuse, you don't add raw scores.
Per engine:
hybrid(query, alpha=0.5, fusionType=relativeScoreFusion). alpha
slides 0.0 (pure keyword) → 1.0 (pure vector). BM25 is BlockMax WAND (default from v1.30, ~10x faster).prefetch a dense and a sparse query, then a fusion step (Fusion.RRF or DBSF);
IDF is computed server-side (v1.15+).<=>) and ts_rank over a tsvector column
separately and combine ranks yourself (RRF in SQL or app code).Concrete current-API code for all four is in references/engines.md.
Quantization trades recall for RAM/cost. Decide by dimension count and a recall test, never blind.
| Method | Compression | When safe |
|---|---|---|
| Scalar (int8) | ~4x | Almost always; tiny recall loss. Good default RAM cut. |
| Product (PQ) | 8–64x | Large corpora where RAM dominates; needs tuning + recall check. |
| Binary | ~32x (~40x faster via SIMD popcount) | High-dim only (≥1024). On 384-dim it shreds recall — measure or don't. |
pgvector halfvec | ~2x | Near-free: 16-bit float, near-identical recall, and required for >2000 dims. |
Reach for halfvec first in Postgres — it is the cheapest win. Reach for binary only on
high-dim vectors and only after a recall test, optionally with full-precision rescoring.
| Anti-pattern | Why it bites | Do instead |
|---|---|---|
Cosine-trained model indexed with L2 (<->) | Silently wrong ranking, no error | Match metric to model — cosine → <=> / vector_cosine_ops |
| Post-filtering top-k results | Returns < k rows, sometimes 0, on selective filters | Filter inside the search; index the filter field |
| IVFFlat + selective filter, no iterative scan | Drops rows; deprecated path in pgvector 0.8 | HNSW + hnsw.iterative_scan='relaxed_order' |
| Never measuring recall | "Search is bad" with no number to move | Recall@k vs an exact baseline before shipping |
| Binary quantization on 384-dim | Recall collapses, then blamed on the engine | Binary only ≥1024 dims, after a recall test; else scalar/halfvec |
| One-by-one upserts | 10–100x slower, index thrash | Batch hundreds per request |
ef_construction < 2*m | Permanently weak graph; recall needs a full rebuild | Keep ef_construction >= 2*m (≥32 at default m=16) |
| Store only raw text, re-embed to "update" | Drift, cost, no point-update path | Re-upsert the vector by id |
| Unindexed filter field | Full scan, latency spikes | Payload index (Qdrant) / B-tree (pgvector) / sane metadata cardinality (Pinecone) |
query_points RRF; Weaviate hybrid; Pinecone serverless sparse-dense).Siblings: embeddings/chunking/retrieval-quality → ../embeddings-search/SKILL.md; the full RAG
loop → ../rag/SKILL.md; general Postgres → ../postgresdb/SKILL.md.
Validate a produced index DDL / collection schema with scripts/verify.sh <artifact-file>.
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/vector-db of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Vector DB next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vector DB this skillericrisco/rsc-harness | 156 | — | ~2.8k | Automated safety check: Pass | MIT | |
| RAG Implementationwshobson/agents | 40k | 9 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Hunt RAG Vectorelementalsouls/Claude-BugHunter | 4.8k | — | ~2.6k | Automated safety check: Pass | MIT | |
| RAG ArchitectJeffallan/claude-skills | 12k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| RAG Patternssoftspark/ai-toolkit | 179 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Pgvector Semantic Searchtimescale/pg-aiguide | 1.9k | 1 repos | ~3.8k | Automated safety check: Pass | Apache-2.0 |
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
elementalsouls/Claude-BugHunter
Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
softspark/ai-toolkit
RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
aiskillstore/marketplace
Expert in vector databases, embedding strategies, and semantic search implementation.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…. Vector DB is an agent skill from ericrisco/rsc-harness. Use when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric, dimensions, HNSW parameters, named vectors); filtering on metadata; hybrid dense-plus-sparse search; and quantization to cut RAM and cost.
Vector DB fits situations like: operating a vector store as a data layer — choosing; migrating between Pinecone; weaviate and pgvector; designing a collection.
Run `npx skills add ericrisco/rsc-harness --skill vector-db -a claude-code`. Or copy the skill folder (skills/vector-db in ericrisco/rsc-harness) into .claude/skills/vector-db in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill vector-db -a codex`. Or copy the skill folder (skills/vector-db in ericrisco/rsc-harness) into .agents/skills/vector-db in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill vector-db -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vector-db, .gemini/skills/vector-db, .github/skills/vector-db and .opencode/skills/vector-db in your project.
Going by SKILL.md and its folder, Vector DB needs a shell for the scripts in its folder. Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Vector DB is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Vector DB: RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars), RAG Architect (Jeffallan/claude-skills, 12k stars) and RAG Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.