Pgvector Semantic Search
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
Retrieval-Augmented Generation patterns for grounded LLM responses.
$ npx skills add yonatangross/orchestkit --skill rag-retrieval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yonatangross/orchestkit rag-retrieval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/rag-retrieval .claude/skills/rag-retrieval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rag-retrieval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/rag-retrieval into .claude/skills/rag-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-retrieval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yonatangross/orchestkit/tree/main/src/skills/rag-retrievalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yonatangross/orchestkit --skill rag-retrieval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yonatangross/orchestkit rag-retrieval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skills/rag-retrieval .agents/skills/rag-retrieval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rag-retrieval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/rag-retrieval into .agents/skills/rag-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-retrieval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill rag-retrieval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yonatangross/orchestkit rag-retrieval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skills/rag-retrieval .cursor/skills/rag-retrieval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rag-retrieval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/rag-retrieval into .cursor/skills/rag-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-retrieval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yonatangross/orchestkit.git --path src/skills/rag-retrieval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yonatangross/orchestkit --skill rag-retrieval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yonatangross/orchestkit rag-retrieval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skills/rag-retrieval .gemini/skills/rag-retrieval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rag-retrieval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/rag-retrieval into .gemini/skills/rag-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-retrieval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yonatangross/orchestkit rag-retrievalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yonatangross/orchestkit --skill rag-retrieval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skills/rag-retrieval .github/skills/rag-retrieval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rag-retrieval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/rag-retrieval into .github/skills/rag-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-retrieval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill rag-retrieval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yonatangross/orchestkit rag-retrieval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skills/rag-retrieval .opencode/skills/rag-retrieval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rag-retrieval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/rag-retrieval into .opencode/skills/rag-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-retrieval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rag-retrievalRetrieval-Augmented Generation patterns for grounded LLM responses.
RAG Retrieval is an agent skill from yonatangross/orchestkit. Retrieval-Augmented Generation patterns for grounded LLM responses. Use when building RAG pipelines, embedding documents, implementing hybrid search, contextual retrieval, HyDE, agentic RAG, multimodal RAG, query decomposition, reranking, or pgvector search.
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 47 other files, including scripts (for example `checklists/rag-quality.md`, `checklists/search-implementation-checklist.md` and `examples/chatbot-with-rag-example.ts`). Compatibility notes: Claude Code 2.1.277+.
It sits in AI & LLM Engineering, covering Retrieval-augmented generation. It works with pgvector. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.
10 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 02bbf9a. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGlobGrepWebFetchWebSearchFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (TypeScript, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
postgresql.orgelastic.colearn.microsoft.comdocs.voyageai.comgithub.complatform.openai.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Claude Code 2.1.277+.
From compatibility in the SKILL.md frontmatter.
RAG Retrieval loads about 4.4k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 1,800 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from yonatangross/orchestkit at commit 02bbf9a, republished under its MIT licence (© yonatangross). 1,800 words, ~4,396 tokens.
.claude/skills/rag-retrieval/SKILL.md (or your agent's skills folder). This skill also uses 44 other files; get the full folder from GitHub.Comprehensive patterns for building production RAG systems. Each category has individual rule files in rules/ loaded on-demand.
House thresholds, fusion ordering, and the latency and quality budgets we assert live in House delta below. Vendor documentation is linked, not restated (see Upstream coverage).
| Category | Rules | Impact | When to Use |
|---|---|---|---|
| Core RAG | 4 | CRITICAL | Basic RAG, citations, hybrid search, context management |
| Embeddings | 3 | HIGH | Model selection, chunking, batch/cache optimization |
| Contextual Retrieval | 3 | HIGH | Context-prepending, hybrid BM25+vector, pipeline |
| HyDE | 3 | HIGH | Vocabulary mismatch, hypothetical document generation |
| Agentic RAG | 4 | HIGH | Self-RAG, CRAG, knowledge graphs, adaptive routing |
| Multimodal RAG | 3 | MEDIUM | Image+text retrieval, PDF chunking, cross-modal search |
| Query Decomposition | 3 | MEDIUM | Multi-concept queries, parallel retrieval, RRF fusion |
| Reranking | 3 | MEDIUM | Cross-encoder, LLM scoring, combined signals |
| PGVector | 4 | HIGH | PostgreSQL hybrid search, HNSW indexes, schema design |
Total: 30 rules across 9 categories
Fundamental patterns for retrieval, generation, and pipeline composition.
| Rule | File | Key Pattern |
|---|---|---|
| Basic RAG | rules/core-basic-rag.md | Retrieve + context + generate with citations |
| Hybrid Search | rules/core-hybrid-search.md | RRF fusion (k=60) for semantic + keyword |
| Context Management | rules/core-context-management.md | Token budgeting + sufficiency check |
| Pipeline Composition | rules/core-pipeline-composition.md | Composable Decompose → HyDE → Retrieve → Rerank |
Embedding models, chunking strategies, and production optimization.
| Rule | File | Key Pattern |
|---|---|---|
| Models & API | rules/embeddings-models.md | Model selection, batch API, similarity |
| Chunking | rules/embeddings-chunking.md | Semantic boundary splitting, 512 token sweet spot |
| Advanced | rules/embeddings-advanced.md | Redis cache, Matryoshka dims, batch processing |
Anthropic's context-prepending technique — 67% fewer retrieval failures.
| Rule | File | Key Pattern |
|---|---|---|
| Context Prepending | rules/contextual-prepend.md | LLM-generated context + prompt caching |
| Hybrid Search | rules/contextual-hybrid.md | 40% BM25 / 60% vector weight split |
| Complete Pipeline | rules/contextual-pipeline.md | End-to-end indexing + hybrid retrieval |
Hypothetical Document Embeddings for bridging vocabulary gaps.
| Rule | File | Key Pattern |
|---|---|---|
| Generation | rules/hyde-generation.md | Embed hypothetical doc, not query |
| Per-Concept | rules/hyde-per-concept.md | Parallel HyDE for multi-topic queries |
| Fallback | rules/hyde-fallback.md | 2-3s timeout → direct embedding fallback |
Self-correcting retrieval with LLM-driven decision making.
| Rule | File | Key Pattern |
|---|---|---|
| Self-RAG | rules/agentic-self-rag.md | Binary document grading for relevance |
| Corrective RAG | rules/agentic-corrective-rag.md | CRAG workflow with web fallback |
| Knowledge Graph | rules/agentic-knowledge-graph.md | KG + vector hybrid for entity-rich domains |
| Adaptive Retrieval | rules/agentic-adaptive-retrieval.md | Query routing to optimal strategy |
Image + text retrieval with cross-modal search.
| Rule | File | Key Pattern |
|---|---|---|
| Embeddings | rules/multimodal-embeddings.md | CLIP, SigLIP 2, Voyage multimodal-3 |
| Chunking | rules/multimodal-chunking.md | PDF extraction preserving images |
| Pipeline | rules/multimodal-pipeline.md | Dedup + hybrid retrieval + generation |
Breaking complex queries into concepts for parallel retrieval.
| Rule | File | Key Pattern |
|---|---|---|
| Detection | rules/query-detection.md | Heuristic indicators (<1ms fast path) |
| Decompose + RRF | rules/query-decompose.md | LLM concept extraction + parallel retrieval |
| HyDE Combo | rules/query-hyde-combo.md | Decompose + HyDE for maximum coverage |
Post-retrieval re-scoring for higher precision.
| Rule | File | Key Pattern |
|---|---|---|
| Cross-Encoder | rules/reranking-cross-encoder.md | ms-marco-MiniLM (~50ms, free) |
| LLM Reranking | rules/reranking-llm.md | Batch scoring + Cohere API |
| Combined | rules/reranking-combined.md | Multi-signal weighted scoring |
Production hybrid search with PostgreSQL.
| Rule | File | Key Pattern |
|---|---|---|
| Schema | rules/pgvector-schema.md | HNSW index + pre-computed tsvector |
| Hybrid Search | rules/pgvector-hybrid-search.md | SQLAlchemy RRF with FULL OUTER JOIN |
| Indexing | rules/pgvector-indexing.md | HNSW (17x faster) vs IVFFlat |
| Metadata | rules/pgvector-metadata.md | Filtering, boosting, Redis 8 comparison |
from openai import OpenAI
client = OpenAI()
async def rag_query(question: str, top_k: int = 5) -> dict:
"""Basic RAG with citations."""
docs = await vector_db.search(question, limit=top_k)
context = "\n\n".join([f"[{i+1}] {doc.text}" for i, doc in enumerate(docs)])
response = await llm.chat([
{"role": "system", "content": "Answer with inline citations [1], [2]. Use ONLY provided context."},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"}
])
return {"answer": response.content, "sources": [d.metadata['source'] for d in docs]}| Decision | Recommendation |
|---|---|
| Embedding model | text-embedding-3-small (general), voyage-3.5 (production) |
| Chunk size | 256-1024 tokens (512 typical) |
| Hybrid weight | 40% BM25 / 60% vector |
| Top-k | 3-10 documents |
| Temperature | 0.1-0.3 (factual) |
| Context budget | 4K-8K tokens |
| Reranking | Retrieve 50, rerank to 10 |
| Vector index | HNSW (production), IVFFlat (high-volume) |
| HyDE timeout | 2-3 seconds with fallback |
| Query decomposition | Heuristic first, LLM only if multi-concept |
Fetch these from the source instead of restating them here. Where a row says the house subset stays, that named file carries only the tuned values and the reason, not a tutorial.
| Topic | Source |
|---|---|
| pgvector install, operators, index build syntax, "why isn't my index used" | https://github.com/pgvector/pgvector; house subset (HNSW m=16, ef_construction=64, halfvec, binary quantize) stays in rules/pgvector-indexing.md and rules/pgvector-schema.md |
Postgres full-text search: tsvector, tsquery, GIN, ts_rank_cd | https://www.postgresql.org/docs/current/textsearch.html; house subset (generated STORED column) stays in rules/pgvector-schema.md |
| Reading query plans, confirming an index is used, VACUUM/ANALYZE | https://www.postgresql.org/docs/current/using-explain.html |
| RRF mechanics: rank constant, rank window size, tie handling | https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion; house subset (k=60, 3x fetch) stays in rules/core-hybrid-search.md and rules/pgvector-hybrid-search.md |
| Field-weighted boosting and scoring profiles as a concept | https://learn.microsoft.com/en-us/azure/search/index-add-scoring-profiles; our boost factors and their ordering are in House delta and rules/pgvector-metadata.md |
Embedding API mechanics: batch limits, input_type, dimensions, pricing | https://docs.voyageai.com/docs/embeddings and https://platform.openai.com/docs/guides/embeddings; model choice stays in rules/embeddings-models.md |
| Retrieval metric definitions (precision@k, recall@k, MRR, nDCG) and building the query set | ork skill golden-dataset |
| Search endpoint shape, pagination, error bodies | ork skill api-design |
| Tracing and dashboarding search calls | ork skill monitoring-observability |
| Writing the integration test that runs these assertions | ork skill testing-integration |
See test-cases.json for 30 test cases across all categories.
There is no pass-rate, precision, recall, or MRR floor for this skill. Earlier
wording here promised those numbers; they were never measured and are written down
nowhere. They are left unset rather than filled in with plausible-looking values,
because a threshold nobody measured is worse than an absent one: it gets cited in
review as though it means something, and passes or fails changes for no defensible
reason. test-cases.json is the right substrate for establishing real ones, and a
baseline run over those 30 cases would produce numbers worth asserting.
The one budget this skill does assert is latency, in House delta.
ork:langgraph - LangGraph workflow patterns (for agentic RAG workflows)ork:golden-dataset - Evaluate retrieval qualityork:llm-integration - Local embeddings with nomic-embed-textork:multimodal-llm - Image analysis for multimodal RAGork:database-patterns - Schema design for vector searchork:performance - Caching repeated RAG responsesKeywords: retrieval, context, chunks, relevance, rag Solves:
Keywords: hybrid, bm25, vector, fusion, rrf Solves:
Keywords: embedding, text to vector, vectorize, chunk, similarity Solves:
Keywords: contextual, anthropic, context-prepend, bm25 Solves:
Keywords: hyde, hypothetical, vocabulary mismatch Solves:
Keywords: self-rag, crag, corrective, adaptive, grading Solves:
Keywords: multimodal, image, clip, vision, pdf Solves:
Keywords: decompose, multi-concept, complex query Solves:
Keywords: rerank, cross-encoder, precision, scoring Solves:
Keywords: pgvector, postgresql, hnsw, tsvector, hybrid Solves:
Inlined rather than placed in references/: this skill carries its house knowledge
in rules/ (32 files) and has never had a references/ directory, so a lone
delta file there would be the only occupant.
Embedding-model choice, chunking algorithms, vector-index tuning, reranker APIs and the query-rewriting literature are vendor and upstream territory; see Upstream coverage for where each lives. What follows is only what OrchestKit adds, contradicts, or has to warn about.
Source: checklists/rag-quality.md:57, the only retrieval budget this skill has ever
actually asserted. It covers the whole path a user waits on (embed the query, search,
rerank, assemble context), not one stage in isolation. A pipeline that clears 500ms
per stage and blows it in aggregate has failed this budget.
Measure at p95, not mean. Retrieval latency is dominated by tail behaviour (cold index shards, reranker queueing), and a mean hides exactly the requests users complain about.
Fusion weights, rerank depth and hybrid alpha are properties of a corpus, not of the technique. A weighting that lifts recall on prose documentation regularly hurts it on code or tabular data, so values copied between projects are noise rather than a starting point. Re-derive them per corpus against that corpus's own eval set.
Why: boost factors multiply onto the RRF score of an already-fused result set and the list is then re-sorted, using section-title match 1.5x, section-path match 1.15x, and code block on a technical query 1.2x; boosting per-method scores before fusion destroys the property RRF depends on (it fuses by rank, not by score) and lets a lone strong keyword hit outrank a document both methods agreed on, and this exact order is what produced the +6% MRR credited to boosting in rules/pgvector-metadata.md (distilled from the retired checklists/search-implementation-checklist.md and examples/examples/orchestkit-retrieval.md; no traced incident).
Upstream: https://learn.microsoft.com/en-us/azure/search/index-add-scoring-profiles
Why: measured pass rate on the reference corpus went 87.2% at 1x, 89.5% at 2x, 91.1% at 3x, 91.3% at 4x, so the fourth multiple bought 0.2 points for a third more rows scanned per method, which is why 3x is the house default hard-coded in rules/core-hybrid-search.md and rules/pgvector-hybrid-search.md; raise it only with a golden-set number that beats this curve (distilled from the retired examples/examples/orchestkit-retrieval.md; no traced incident).
Upstream: https://www.elastic.co/docs/reference/elasticsearch/rest-apis/reciprocal-rank-fusion
Why: the house bar before a retrieval change ships is pass rate at least 90% on the golden query set, precision@10 at least 0.70, recall@10 at least 0.85, and MRR at least 0.65, floors set just under the reference pipeline's observed 91.6% pass rate and 0.777 overall MRR (0.695 on the hard slice) so that a real regression trips them instead of a noisy run (distilled from the retired checklists/search-implementation-checklist.md and examples/examples/orchestkit-retrieval.md; no traced incident).
Upstream: ork skill golden-dataset
Why: the integration test asserts vector search under 100ms, keyword search under 50ms, and fused hybrid under 150ms separately, against an observed 415-chunk baseline of P50 15ms / P95 32ms / P99 62ms, because a single total-latency assertion stays green while one stage silently doubles, which is precisely the shape of an index that quietly stopped being used (distilled from the retired checklists/search-implementation-checklist.md and examples/examples/orchestkit-retrieval.md; no traced incident). Upstream: https://www.postgresql.org/docs/current/using-explain.html
Why: in the reference measurement the query embedding was P50 8ms of a 15ms end-to-end search against 2ms for vector search and 3ms for keyword search, and it was the throughput ceiling at roughly 120 requests per second, so cache and batch embeddings first because index tuning below that ceiling moves the total by single-digit percent (distilled from the retired examples/examples/orchestkit-retrieval.md; no traced incident). Upstream: https://docs.voyageai.com/docs/embeddings
© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 44 other files (scripts) in src/skills/rag-retrieval of yonatangross/orchestkit.
Open the folder on GitHubat commit 02bbf9a
RAG Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| RAG Retrieval this skillyonatangross/orchestkit | 290 | — | ~4.4k | Automated safety check: Pass | MIT | |
| Pgvector Semantic Searchtimescale/pg-aiguide | 1.9k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Postgres Hybrid Text Searchtimescale/pg-aiguide | 1.9k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| RAG Implementationwshobson/agents | 40k | 9 repos | ~1.1k | Automated safety check: Pass | MIT | |
| RAG ArchitectJeffallan/claude-skills | 12k | — | ~2k | Automated safety check: Pass | MIT | |
| Hunt RAG Vectorelementalsouls/Claude-BugHunter | 4.8k | — | ~2.6k | Automated safety check: Pass | MIT |
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
timescale/pg-aiguide
A skill your agent uses to implement hybrid search combining BM25 keyword search with semantic vector search using Reciprocal Rank Fusion (RRF).
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
elementalsouls/Claude-BugHunter
Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…
github/awesome-copilot
Build production RAG pipelines and persistent agent memory using Pinecone as the vector database backend.
yonatangross/orchestkit
API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.
yonatangross/orchestkit
ADR templates in the Nygard format with context, decision, consequences, and alternatives.
yonatangross/orchestkit
Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.
yonatangross/orchestkit
Structured review processes, conventional comments, language-specific checklists, and feedback templates.
yonatangross/orchestkit
Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.
yonatangross/orchestkit
Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.
Works with
Categories
Retrieval-Augmented Generation patterns for grounded LLM responses. RAG Retrieval is an agent skill from yonatangross/orchestkit. Retrieval-Augmented Generation patterns for grounded LLM responses.
RAG Retrieval fits situations like: building RAG pipelines; embedding documents; implementing hybrid search; contextual retrieval.
Run `npx skills add yonatangross/orchestkit --skill rag-retrieval -a claude-code`. Or copy the skill folder (src/skills/rag-retrieval in yonatangross/orchestkit) into .claude/skills/rag-retrieval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yonatangross/orchestkit --skill rag-retrieval -a codex`. Or copy the skill folder (src/skills/rag-retrieval in yonatangross/orchestkit) into .agents/skills/rag-retrieval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill rag-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-retrieval, .gemini/skills/rag-retrieval, .github/skills/rag-retrieval and .opencode/skills/rag-retrieval in your project.
Going by SKILL.md and its folder, RAG Retrieval needs TypeScript for the scripts in its folder. Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..
SKILL.md names 6 domains. As links in the text: postgresql.org, elastic.co, learn.microsoft.com, docs.voyageai.com, github.com and platform.openai.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
RAG Retrieval is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with RAG Retrieval: Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Postgres Hybrid Text Search (timescale/pg-aiguide, 1.9k stars), RAG Implementation (wshobson/agents, 40k stars) and RAG Architect (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 290 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.
Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.