RAG Implementation
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
$ npx skills add Jeffallan/claude-skills --skill rag-architect -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Jeffallan/claude-skills rag-architect --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-architect .claude/skills/rag-architect && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rag-architect" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/rag-architect into .claude/skills/rag-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-architect", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Jeffallan/claude-skills/tree/main/skills/rag-architectType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Jeffallan/claude-skills --skill rag-architect -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Jeffallan/claude-skills rag-architect --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/rag-architect .agents/skills/rag-architect && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rag-architect" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/rag-architect into .agents/skills/rag-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-architect", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Jeffallan/claude-skills --skill rag-architect -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Jeffallan/claude-skills rag-architect --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/rag-architect .cursor/skills/rag-architect && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rag-architect" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/rag-architect into .cursor/skills/rag-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-architect", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Jeffallan/claude-skills.git --path skills/rag-architect--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Jeffallan/claude-skills --skill rag-architect -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Jeffallan/claude-skills rag-architect --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/rag-architect .gemini/skills/rag-architect && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rag-architect" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/rag-architect into .gemini/skills/rag-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-architect", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Jeffallan/claude-skills rag-architectInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Jeffallan/claude-skills --skill rag-architect -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/rag-architect .github/skills/rag-architect && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rag-architect" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/rag-architect into .github/skills/rag-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-architect", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Jeffallan/claude-skills --skill rag-architect -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Jeffallan/claude-skills rag-architect --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/rag-architect .opencode/skills/rag-architect && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rag-architect" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/rag-architect into .opencode/skills/rag-architect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-architect", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rag-architectDesigns retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
The agent works through requirements for latency, accuracy and scale, vector store design, chunking strategy, the retrieval pipeline and evaluation. The examples come with assertion checkpoints: chunks must carry source metadata, the indexed count must match the number of unique point IDs, and a hybrid search for a test query must return results. The chunking example advises tuning chunk size against your own documents instead of copying a default.
Code examples use LangChain text splitters, OpenAI embeddings stored in Qdrant, hybrid vector plus BM25 search with tenant filtering, and reranking with Cohere. Reference files compare Pinecone, Weaviate, Chroma, pgvector and Qdrant, and cover embedding model choice, chunking, retrieval optimization with query expansion and filtering, and evaluation. Provider API keys are meant to come from environment variables or a secrets manager.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comsynergetic.solutionsjeffallan.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
COHERE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
RAG Architect loads about 2k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 390 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 390 words, ~2,009 tokens.
.claude/skills/rag-architect/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.For each step, validate before moving on (see checkpoints below).
Load detailed guidance based on context:
| Topic | Reference | Load When |
|---|---|---|
| Vector Databases | references/vector-databases.md | Comparing Pinecone, Weaviate, Chroma, pgvector, Qdrant |
| Embedding Models | references/embedding-models.md | Selecting embeddings, fine-tuning, dimension trade-offs |
| Chunking Strategies | references/chunking-strategies.md | Document splitting, overlap, semantic chunking |
| Retrieval Optimization | references/retrieval-optimization.md | Hybrid search, reranking, query expansion, filtering |
| RAG Evaluation | references/rag-evaluation.md | Metrics, evaluation frameworks, debugging retrieval |
from langchain.text_splitter import RecursiveCharacterTextSplitter
# Evaluate chunk_size on your domain data — never use 512 blindly
splitter = RecursiveCharacterTextSplitter(
chunk_size=800,
chunk_overlap=100,
separators=["\n\n", "\n", ". ", " "],
)
chunks = splitter.create_documents(
texts=[doc.page_content for doc in raw_docs],
metadatas=[{"source": doc.metadata["source"], "timestamp": doc.metadata.get("timestamp")} for doc in raw_docs],
)Checkpoint: assert all(c.metadata.get("source") for c in chunks), "Missing source metadata"
from openai import OpenAI
import qdrant_client
from qdrant_client.models import VectorParams, Distance, PointStruct
client = OpenAI()
qdrant = qdrant_client.QdrantClient("localhost", port=6333)
# Create collection
qdrant.recreate_collection(
collection_name="knowledge_base",
vectors_config=VectorParams(size=1536, distance=Distance.COSINE),
)
def embed_chunks(chunks: list[str], model: str = "text-embedding-3-small") -> list[list[float]]:
response = client.embeddings.create(input=chunks, model=model)
return [r.embedding for r in response.data]
# Idempotent upsert with deduplication via deterministic IDs
import hashlib, uuid
points = []
for i, chunk in enumerate(chunks):
doc_id = str(uuid.UUID(hashlib.md5(chunk.page_content.encode()).hexdigest()))
embedding = embed_chunks([chunk.page_content])[0]
points.append(PointStruct(id=doc_id, vector=embedding, payload=chunk.metadata))
qdrant.upsert(collection_name="knowledge_base", points=points)Checkpoint: assert qdrant.count("knowledge_base").count == len(set(p.id for p in points)), "Deduplication failed"
from qdrant_client.models import Filter, FieldCondition, MatchValue, SparseVector
from rank_bm25 import BM25Okapi
def hybrid_search(query: str, tenant_id: str, top_k: int = 20) -> list:
# Dense retrieval
query_embedding = embed_chunks([query])[0]
tenant_filter = Filter(must=[FieldCondition(key="tenant_id", match=MatchValue(value=tenant_id))])
dense_results = qdrant.search(
collection_name="knowledge_base",
query_vector=query_embedding,
query_filter=tenant_filter,
limit=top_k,
)
# Sparse retrieval (BM25)
corpus = [r.payload.get("text", "") for r in dense_results]
bm25 = BM25Okapi([doc.split() for doc in corpus])
bm25_scores = bm25.get_scores(query.split())
# Reciprocal Rank Fusion
ranked = sorted(
zip(dense_results, bm25_scores),
key=lambda x: 0.6 * x[0].score + 0.4 * x[1],
reverse=True,
)
return [r for r, _ in ranked[:top_k]]Checkpoint: assert len(hybrid_search("test query", tenant_id="demo")) > 0, "Hybrid search returned no results"
Load provider API keys from environment variables or a secrets manager; never commit them to source code.
import os
import cohere
co = cohere.Client(os.environ["COHERE_API_KEY"])
def rerank(query: str, results: list, top_n: int = 5) -> list:
docs = [r.payload.get("text", "") for r in results]
reranked = co.rerank(query=query, documents=docs, top_n=top_n, model="rerank-english-v3.0")
return [results[r.index] for r in reranked.results]# Run precision@k and recall@k against a labeled evaluation set
# python evaluate.py --metrics precision@10 recall@10 mrr --collection knowledge_base
from ragas import evaluate
from ragas.metrics import context_precision, context_recall, faithfulness, answer_relevancy
from datasets import Dataset
eval_dataset = Dataset.from_dict({
"question": questions,
"contexts": retrieved_contexts,
"answer": generated_answers,
"ground_truth": ground_truth_answers,
})
results = evaluate(eval_dataset, metrics=[context_precision, context_recall, faithfulness, answer_relevancy])
print(results)Checkpoint: Target context_precision >= 0.7 and context_recall >= 0.6 before moving to LLM integration.
When designing RAG architecture, deliver:
Maintained by @jeffallan, Principal Consultant at Synergetic Solutions
© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/rag-architect of Jeffallan/claude-skills.
Open the folder on GitHubat commit 1be15d8
RAG Architect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| RAG Architect this skillJeffallan/claude-skills | 12k | — | ~2k | Automated safety check: Pass | MIT | |
| RAG Implementationwshobson/agents | 40k | 9 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Hunt RAG Vectorelementalsouls/Claude-BugHunter | 4.8k | — | ~2.6k | Automated safety check: Pass | MIT | |
| Agentsop Multi Tenant RAGagentsope/SkillAlchemy | 466 | — | ~9.8k | Automated safety check: Pass | MIT | |
| RAG Patternssoftspark/ai-toolkit | 179 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Vector DBericrisco/rsc-harness | 174 | — | ~2.8k | Automated safety check: Pass | MIT |
wshobson/agents
Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.
elementalsouls/Claude-BugHunter
Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…
agentsope/SkillAlchemy
Security-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy.
softspark/ai-toolkit
RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.
ericrisco/rsc-harness
A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…
langchain-ai/langchain-skills
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.
Jeffallan/claude-skills
Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.
Jeffallan/claude-skills
Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.
Jeffallan/claude-skills
Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.
Jeffallan/claude-skills
Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.
Jeffallan/claude-skills
Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
Categories
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step. The agent works through requirements for latency, accuracy and scale, vector store design, chunking strategy, the retrieval pipeline and evaluation. The examples come with assertion checkpoints: chunks must carry source metadata, the indexed count must match the number of unique point IDs, and a hybrid search for a test query must return results.
RAG Architect fits situations like: building a knowledge-grounded assistant over internal documents; choosing a vector database and indexing strategy; tuning chunk size, overlap and metadata for retrieval quality; adding hybrid search and reranking to improve results.
Run `npx skills add Jeffallan/claude-skills --skill rag-architect -a claude-code`. Or copy the skill folder (skills/rag-architect in Jeffallan/claude-skills) into .claude/skills/rag-architect in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Jeffallan/claude-skills --skill rag-architect -a codex`. Or copy the skill folder (skills/rag-architect in Jeffallan/claude-skills) into .agents/skills/rag-architect in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill rag-architect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-architect, .gemini/skills/rag-architect, .github/skills/rag-architect and .opencode/skills/rag-architect in your project.
Going by SKILL.md and its folder, RAG Architect needs credentials named COHERE_API_KEY. Our summary lists: API keys for the embedding and reranking providers you choose.
SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
RAG Architect is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 25k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with RAG Architect: RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars), Agentsop Multi Tenant RAG (agentsope/SkillAlchemy, 466 stars) and RAG Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,788 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.
Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.