Agent skill

RAG Architect

by Jeffallan in Jeffallan/claude-skills

Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

MITAuto-check passedAI & LLM Engineering

Install RAG Architect

skills CLI
$ npx skills add Jeffallan/claude-skills --skill rag-architect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jeffallan/claude-skills rag-architect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-architect .claude/skills/rag-architect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-architect
GitHub stars
12k
Token cost
~2k tokens
SKILL.md length
390 words
Files
6 (incl. references)
Skills in repo
58
Repo updated
First seen
Licence
MIT

At a glance

Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

  • Works in 5 steps: Chunking Documents → Generating Embeddings & Indexing → Hybrid Search (Vector + BM25) → …
  • Building a knowledge-grounded assistant over internal documents
  • SKILL.md covers Core Workflow, Reference Guide, Implementation Examples and Constraints, plus 1 more section
  • Needs COHERE_API_KEY

What it does

The agent works through requirements for latency, accuracy and scale, vector store design, chunking strategy, the retrieval pipeline and evaluation. The examples come with assertion checkpoints: chunks must carry source metadata, the indexed count must match the number of unique point IDs, and a hybrid search for a test query must return results. The chunking example advises tuning chunk size against your own documents instead of copying a default.

Code examples use LangChain text splitters, OpenAI embeddings stored in Qdrant, hybrid vector plus BM25 search with tenant filtering, and reranking with Cohere. Reference files compare Pinecone, Weaviate, Chroma, pgvector and Qdrant, and cover embedding model choice, chunking, retrieval optimization with query expansion and filtering, and evaluation. Provider API keys are meant to come from environment variables or a secrets manager.

When your agent uses it

  • Building a knowledge-grounded assistant over internal documents
  • Choosing a vector database and indexing strategy
  • Tuning chunk size, overlap and metadata for retrieval quality
  • Adding hybrid search and reranking to improve results
  • Measuring and debugging retrieval quality

Example prompts

  • “Design a RAG pipeline for our help center articles with chunking, embeddings and Qdrant.”
  • “Retrieval misses obvious matches for product codes, so add BM25 hybrid search with reranking.”
  • “Compare Pinecone, Weaviate and pgvector for a document set that keeps growing.”
  • “Set up retrieval metrics so we can compare two chunking strategies.”

Requirements

  • API keys for the embedding and reranking providers you choose

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Chunking Documents
  2. Generating Embeddings & Indexing
  3. Hybrid Search (Vector + BM25)
  4. Reranking Top-K Results
  5. Retrieval Evaluation

What it can do on your machine

Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • synergetic.solutions
    • jeffallan.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • COHERE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Architect loads about 2k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 390 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~27k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 390 words, ~2,009 tokens.

Download SKILL.mdSave it as .claude/skills/rag-architect/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
rag-architect
description
Designs and implements production-grade RAG systems by chunking documents, generating embeddings, configuring vector stores, building hybrid search pipelines, applying reranking, and evaluating retrieval quality. Use when building RAG systems, vector databases, or knowledge-grounded AI applications requiring semantic search, document retrieval, context augmentation, similarity search, or embedding-based indexing.
license
MIT
metadata.author
https://github.com/Jeffallan
metadata.company
https://synergetic.solutions
metadata.version
1.1.0
metadata.domain
data-ml
metadata.triggers
RAG, retrieval-augmented generation, vector search, embeddings, semantic search, vector database, document retrieval, knowledge base, context retrieval…
metadata.role
architect
metadata.scope
system-design
metadata.output-format
architecture
metadata.related-skills
python-pro, database-optimizer, monitoring-expert, api-designer

RAG Architect

Core Workflow

  1. Requirements Analysis — Identify retrieval needs, latency constraints, accuracy requirements, and scale
  2. Vector Store Design — Select database, schema design, indexing strategy, sharding approach
  3. Chunking Strategy — Document splitting, overlap, semantic boundaries, metadata enrichment
  4. Retrieval Pipeline — Embedding selection, query transformation, hybrid search, reranking
  5. Evaluation & Iteration — Metrics tracking, retrieval debugging, continuous optimization

For each step, validate before moving on (see checkpoints below).

Reference Guide

Load detailed guidance based on context:

TopicReferenceLoad When
Vector Databasesreferences/vector-databases.mdComparing Pinecone, Weaviate, Chroma, pgvector, Qdrant
Embedding Modelsreferences/embedding-models.mdSelecting embeddings, fine-tuning, dimension trade-offs
Chunking Strategiesreferences/chunking-strategies.mdDocument splitting, overlap, semantic chunking
Retrieval Optimizationreferences/retrieval-optimization.mdHybrid search, reranking, query expansion, filtering
RAG Evaluationreferences/rag-evaluation.mdMetrics, evaluation frameworks, debugging retrieval

Implementation Examples

1. Chunking Documents
python
from langchain.text_splitter import RecursiveCharacterTextSplitter

# Evaluate chunk_size on your domain data — never use 512 blindly
splitter = RecursiveCharacterTextSplitter(
    chunk_size=800,
    chunk_overlap=100,
    separators=["\n\n", "\n", ". ", " "],
)

chunks = splitter.create_documents(
    texts=[doc.page_content for doc in raw_docs],
    metadatas=[{"source": doc.metadata["source"], "timestamp": doc.metadata.get("timestamp")} for doc in raw_docs],
)

Checkpoint: assert all(c.metadata.get("source") for c in chunks), "Missing source metadata"

2. Generating Embeddings & Indexing
python
from openai import OpenAI
import qdrant_client
from qdrant_client.models import VectorParams, Distance, PointStruct

client = OpenAI()
qdrant = qdrant_client.QdrantClient("localhost", port=6333)

# Create collection
qdrant.recreate_collection(
    collection_name="knowledge_base",
    vectors_config=VectorParams(size=1536, distance=Distance.COSINE),
)

def embed_chunks(chunks: list[str], model: str = "text-embedding-3-small") -> list[list[float]]:
    response = client.embeddings.create(input=chunks, model=model)
    return [r.embedding for r in response.data]

# Idempotent upsert with deduplication via deterministic IDs
import hashlib, uuid

points = []
for i, chunk in enumerate(chunks):
    doc_id = str(uuid.UUID(hashlib.md5(chunk.page_content.encode()).hexdigest()))
    embedding = embed_chunks([chunk.page_content])[0]
    points.append(PointStruct(id=doc_id, vector=embedding, payload=chunk.metadata))

qdrant.upsert(collection_name="knowledge_base", points=points)

Checkpoint: assert qdrant.count("knowledge_base").count == len(set(p.id for p in points)), "Deduplication failed"

3. Hybrid Search (Vector + BM25)
python
from qdrant_client.models import Filter, FieldCondition, MatchValue, SparseVector
from rank_bm25 import BM25Okapi

def hybrid_search(query: str, tenant_id: str, top_k: int = 20) -> list:
    # Dense retrieval
    query_embedding = embed_chunks([query])[0]
    tenant_filter = Filter(must=[FieldCondition(key="tenant_id", match=MatchValue(value=tenant_id))])
    dense_results = qdrant.search(
        collection_name="knowledge_base",
        query_vector=query_embedding,
        query_filter=tenant_filter,
        limit=top_k,
    )

    # Sparse retrieval (BM25)
    corpus = [r.payload.get("text", "") for r in dense_results]
    bm25 = BM25Okapi([doc.split() for doc in corpus])
    bm25_scores = bm25.get_scores(query.split())

    # Reciprocal Rank Fusion
    ranked = sorted(
        zip(dense_results, bm25_scores),
        key=lambda x: 0.6 * x[0].score + 0.4 * x[1],
        reverse=True,
    )
    return [r for r, _ in ranked[:top_k]]

Checkpoint: assert len(hybrid_search("test query", tenant_id="demo")) > 0, "Hybrid search returned no results"

4. Reranking Top-K Results

Load provider API keys from environment variables or a secrets manager; never commit them to source code.

python
import os

import cohere

co = cohere.Client(os.environ["COHERE_API_KEY"])

def rerank(query: str, results: list, top_n: int = 5) -> list:
    docs = [r.payload.get("text", "") for r in results]
    reranked = co.rerank(query=query, documents=docs, top_n=top_n, model="rerank-english-v3.0")
    return [results[r.index] for r in reranked.results]
5. Retrieval Evaluation
python
# Run precision@k and recall@k against a labeled evaluation set
# python evaluate.py --metrics precision@10 recall@10 mrr --collection knowledge_base

from ragas import evaluate
from ragas.metrics import context_precision, context_recall, faithfulness, answer_relevancy
from datasets import Dataset

eval_dataset = Dataset.from_dict({
    "question": questions,
    "contexts": retrieved_contexts,
    "answer": generated_answers,
    "ground_truth": ground_truth_answers,
})

results = evaluate(eval_dataset, metrics=[context_precision, context_recall, faithfulness, answer_relevancy])
print(results)

Checkpoint: Target context_precision >= 0.7 and context_recall >= 0.6 before moving to LLM integration.

Constraints

Show full SKILL.md (187 more words)Show less
MUST DO
  • Evaluate multiple embedding models on your domain data before committing
  • Implement hybrid search (vector + keyword) for production systems
  • Add metadata filters for multi-tenant or domain-specific retrieval
  • Measure retrieval metrics (precision@k, recall@k, MRR, NDCG)
  • Use reranking for top-k results before passing context to LLM
  • Implement idempotent ingestion with deduplication (deterministic IDs)
  • Monitor retrieval latency and quality over time
  • Version embeddings and plan for model migration
MUST NOT DO
  • Use default chunk size (512) without evaluation on your domain data
  • Skip metadata enrichment (source, timestamp, section)
  • Ignore retrieval quality metrics in favor of only LLM output quality
  • Store raw documents without preprocessing/cleaning
  • Use cosine similarity alone for complex multi-domain retrieval
  • Deploy without testing on production-like data volumes
  • Forget to handle edge cases (empty results, malformed docs)
  • Couple the embedding model tightly to application code

Output Templates

When designing RAG architecture, deliver:

  1. System architecture diagram (ingestion + retrieval pipelines)
  2. Vector database selection with trade-off analysis
  3. Chunking strategy with examples and rationale
  4. Retrieval pipeline design (query → results flow)
  5. Evaluation plan with metrics, benchmarks, and pass/fail thresholds

Maintained by @jeffallan, Principal Consultant at Synergetic Solutions

Documentation

© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/rag-architect of Jeffallan/claude-skills.

  • SKILL.md
  • references/chunking-strategies.md
  • references/embedding-models.md
  • references/rag-evaluation.md
  • references/retrieval-optimization.md
  • references/vector-databases.md

Open the folder on GitHubat commit 1be15d8

Compare with similar skills

RAG Architect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Architect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Architect this skillJeffallan/claude-skills12k—~2kAutomated safety check: PassMIT
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Hunt RAG Vectorelementalsouls/Claude-BugHunter4.8k—~2.6kAutomated safety check: PassMIT
Agentsop Multi Tenant RAGagentsope/SkillAlchemy466—~9.8kAutomated safety check: PassMIT
RAG Patternssoftspark/ai-toolkit179—~1.8kAutomated safety check: PassApache-2.0
Vector DBericrisco/rsc-harness174—~2.8kAutomated safety check: PassMIT

Similar skills

  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Agentsop Multi Tenant RAG

    agentsope/SkillAlchemy

    Security-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy.

    466 GitHub stars~9.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • RAG Patterns

    softspark/ai-toolkit

    RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.

    179 GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Vector DB

    ericrisco/rsc-harness

    A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

    174 GitHub stars~2.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub stars~3.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from Jeffallan/claude-skills

All 58 skills in this repo
  • API Designer

    Jeffallan/claude-skills

    Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.

    12k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • CLI Developer

    Jeffallan/claude-skills

    Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.

    12k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Kubernetes Specialist

    Jeffallan/claude-skills

    Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Laravel Specialist

    Jeffallan/claude-skills

    Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Questions about RAG Architect

What does RAG Architect do?

Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step. The agent works through requirements for latency, accuracy and scale, vector store design, chunking strategy, the retrieval pipeline and evaluation. The examples come with assertion checkpoints: chunks must carry source metadata, the indexed count must match the number of unique point IDs, and a hybrid search for a test query must return results.

When should I use RAG Architect?

RAG Architect fits situations like: building a knowledge-grounded assistant over internal documents; choosing a vector database and indexing strategy; tuning chunk size, overlap and metadata for retrieval quality; adding hybrid search and reranking to improve results.

How do I install RAG Architect in Claude Code?

Run `npx skills add Jeffallan/claude-skills --skill rag-architect -a claude-code`. Or copy the skill folder (skills/rag-architect in Jeffallan/claude-skills) into .claude/skills/rag-architect in your project. Claude Code loads it when a task matches its description.

How do I install RAG Architect in Codex?

Run `npx skills add Jeffallan/claude-skills --skill rag-architect -a codex`. Or copy the skill folder (skills/rag-architect in Jeffallan/claude-skills) into .agents/skills/rag-architect in your project. Codex loads it when a task matches its description.

Can I use RAG Architect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill rag-architect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-architect, .gemini/skills/rag-architect, .github/skills/rag-architect and .opencode/skills/rag-architect in your project.

What does RAG Architect need to run?

Going by SKILL.md and its folder, RAG Architect needs credentials named COHERE_API_KEY. Our summary lists: API keys for the embedding and reranking providers you choose.

Does RAG Architect access the network?

SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.

Is RAG Architect safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Architect use?

RAG Architect is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Architect use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 25k tokens, read only when the agent opens those files.

What are the alternatives to RAG Architect?

Skills that share tags, products or a category with RAG Architect: RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars), Agentsop Multi Tenant RAG (agentsope/SkillAlchemy, 466 stars) and RAG Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Architect?

Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,788 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.

Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.