Agent skill

RAG Implementation

by wshobson in wshobson/agents

Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

MITAuto-check passedAI & LLM Engineering

Install RAG Implementation

skills CLI
$ npx skills add wshobson/agents --skill rag-implementation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents rag-implementation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-application-dev/skills/rag-implementation .claude/skills/rag-implementation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-implementation
GitHub stars
40k
Used in
9 other repos
Token cost
~1.1k tokens
SKILL.md length
245 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

  • Works in 4 steps: Vector Databases → Embeddings → Retrieval Strategies → …
  • Building question answering over a company's private documents
  • SKILL.md covers When to Use This Skill, Core Components, Quick Start with LangGraph and Detailed patterns and worked…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill maps out the parts of a RAG system. It compares vector database options (Pinecone, Weaviate, Milvus, Chroma, Qdrant and pgvector), lists embedding models with their dimensions and intended use, and describes retrieval approaches from dense and sparse retrieval to hybrid search, multi-query generation and HyDE.

Reranking options include cross-encoders, Cohere Rerank, maximal marginal relevance and scoring by an LLM. A Python quick start assembles the pipeline with LangGraph, using Claude through langchain_anthropic and Voyage AI embeddings. Typical goals are question answering over private documents, documentation assistants and research tools that cite sources, with answers grounded to cut down hallucinations.

When your agent uses it

  • Building question answering over a company's private documents
  • Choosing a vector database and embedding model for a new RAG project
  • Reducing hallucinations by grounding answers in retrieved sources
  • Adding reranking or multi-query retrieval to weak search results

Example prompts

  • “Build a document Q&A bot over the PDFs in ./handbook that cites the source pages.”
  • “Compare Qdrant and pgvector for our RAG store and recommend one.”
  • “Add a reranking step to our retriever so the top results are more relevant.”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Vector Databases
  2. Embeddings
  3. Retrieval Strategies
  4. Reranking

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Implementation loads about 1.1k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 245 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 245 words, ~1,139 tokens.

Download SKILL.mdSave it as .claude/skills/rag-implementation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
rag-implementation
description
Build Retrieval-Augmented Generation (RAG) systems for LLM applications with vector databases and semantic search. Use when implementing knowledge-grounded AI, building document Q&A systems, or integrating LLMs with external knowledge bases.

RAG Implementation

Master Retrieval-Augmented Generation (RAG) to build LLM applications that provide accurate, grounded responses using external knowledge sources.

When to Use This Skill

  • Building Q&A systems over proprietary documents
  • Creating chatbots with current, factual information
  • Implementing semantic search with natural language queries
  • Reducing hallucinations with grounded responses
  • Enabling LLMs to access domain-specific knowledge
  • Building documentation assistants
  • Creating research tools with source citation

Core Components

1. Vector Databases

Purpose: Store and retrieve document embeddings efficiently

Options:

  • Pinecone: Managed, scalable, serverless
  • Weaviate: Open-source, hybrid search, GraphQL
  • Milvus: High performance, on-premise
  • Chroma: Lightweight, easy to use, local development
  • Qdrant: Fast, filtered search, Rust-based
  • pgvector: PostgreSQL extension, SQL integration
2. Embeddings

Purpose: Convert text to numerical vectors for similarity search

Models (2026):

ModelDimensionsBest For
voyage-3-large1024Claude apps (Anthropic recommended)
voyage-code-31024Code search
text-embedding-3-large3072OpenAI apps, high accuracy
text-embedding-3-small1536OpenAI apps, cost-effective
bge-large-en-v1.51024Open source, local deployment
multilingual-e5-large1024Multi-language support
3. Retrieval Strategies

Approaches:

  • Dense Retrieval: Semantic similarity via embeddings
  • Sparse Retrieval: Keyword matching (BM25, TF-IDF)
  • Hybrid Search: Combine dense + sparse with weighted fusion
  • Multi-Query: Generate multiple query variations
  • HyDE: Generate hypothetical documents for better retrieval
4. Reranking

Purpose: Improve retrieval quality by reordering results

Methods:

  • Cross-Encoders: BERT-based reranking (ms-marco-MiniLM)
  • Cohere Rerank: API-based reranking
  • Maximal Marginal Relevance (MMR): Diversity + relevance
  • LLM-based: Use LLM to score relevance

Quick Start with LangGraph

python
from langgraph.graph import StateGraph, START, END
from langchain_anthropic import ChatAnthropic
from langchain_voyageai import VoyageAIEmbeddings
from langchain_pinecone import PineconeVectorStore
from langchain_core.documents import Document
from langchain_core.prompts import ChatPromptTemplate
from langchain_text_splitters import RecursiveCharacterTextSplitter
from typing import TypedDict, Annotated

class RAGState(TypedDict):
    question: str
    context: list[Document]
    answer: str

# Initialize components
llm = ChatAnthropic(model="claude-sonnet-5")
embeddings = VoyageAIEmbeddings(model="voyage-3-large")
vectorstore = PineconeVectorStore(index_name="docs", embedding=embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 4})

# RAG prompt
rag_prompt = ChatPromptTemplate.from_template(
    """Answer based on the context below. If you cannot answer, say so.

    Context:
    {context}

    Question: {question}

    Answer:"""
)

async def retrieve(state: RAGState) -> RAGState:
    """Retrieve relevant documents."""
    docs = await retriever.ainvoke(state["question"])
    return {"context": docs}

async def generate(state: RAGState) -> RAGState:
    """Generate answer from context."""
    context_text = "\n\n".join(doc.page_content for doc in state["context"])
    messages = rag_prompt.format_messages(
        context=context_text,
        question=state["question"]
    )
    response = await llm.ainvoke(messages)
    return {"answer": response.content}

# Build RAG graph
builder = StateGraph(RAGState)
builder.add_node("retrieve", retrieve)
builder.add_node("generate", generate)
builder.add_edge(START, "retrieve")
builder.add_edge("retrieve", "generate")
builder.add_edge("generate", END)

rag_chain = builder.compile()

# Use
result = await rag_chain.ainvoke({"question": "What are the main features?"})
print(result["answer"])

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/llm-application-dev/skills/rag-implementation of wshobson/agents.

  • SKILL.md
  • references/details.md

Open the folder on GitHubat commit 46891e7

Used in 9 other repositories

We found 36 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 9 other GitHub owners. This page covers the copy in wshobson/agents, which our catalogue first saw on October 7, 2026.

Compare with similar skills

RAG Implementation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Implementation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Implementation this skillwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Hunt RAG Vectorelementalsouls/Claude-BugHunter4.8k—~2.6kAutomated safety check: PassMIT
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
RAG Patternssoftspark/ai-toolkit179—~1.8kAutomated safety check: PassApache-2.0
Vector Database Engineeraiskillstore/marketplace4307 repos~563Automated safety check: PassNone
Vector DBericrisco/rsc-harness156—~2.8kAutomated safety check: PassMIT

Similar skills

  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Patterns

    softspark/ai-toolkit

    RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.

    179 GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vector Database Engineer

    aiskillstore/marketplace

    Expert in vector databases, embedding strategies, and semantic search implementation.

    430 GitHub starsUsed in 7 repos~563 tokens
    DatabasesAuto-check passed
  • Vector DB

    ericrisco/rsc-harness

    A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

    156 GitHub stars~2.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    FerroxLabs/wayland

    RAG system design covering document chunking strategies, embedding model selection, vector database selection (Pinecone, Weaviate, Chroma, pgvector), retrieval strategies (hybrid search…

    608 GitHub stars~4.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 13 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 12 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 12 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 12 repos~502 tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 11 repos~1.3k tokens
    Auto-check passed

Questions about RAG Implementation

What does RAG Implementation do?

Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline. The skill maps out the parts of a RAG system. It compares vector database options (Pinecone, Weaviate, Milvus, Chroma, Qdrant and pgvector), lists embedding models with their dimensions and intended use, and describes retrieval approaches from dense and sparse retrieval to hybrid search, multi-query generation and HyDE.

When should I use RAG Implementation?

RAG Implementation fits situations like: building question answering over a company's private documents; choosing a vector database and embedding model for a new RAG project; reducing hallucinations by grounding answers in retrieved sources; adding reranking or multi-query retrieval to weak search results.

How do I install RAG Implementation in Claude Code?

Run `npx skills add wshobson/agents --skill rag-implementation -a claude-code`. Or copy the skill folder (plugins/llm-application-dev/skills/rag-implementation in wshobson/agents) into .claude/skills/rag-implementation in your project. Claude Code loads it when a task matches its description.

How do I install RAG Implementation in Codex?

Run `npx skills add wshobson/agents --skill rag-implementation -a codex`. Or copy the skill folder (plugins/llm-application-dev/skills/rag-implementation in wshobson/agents) into .agents/skills/rag-implementation in your project. Codex loads it when a task matches its description.

Can I use RAG Implementation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill rag-implementation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-implementation, .gemini/skills/rag-implementation, .github/skills/rag-implementation and .opencode/skills/rag-implementation in your project.

What does RAG Implementation need to run?

SKILL.md names no scripts, command-line tools or credentials: RAG Implementation is instructions for the agent only.

Does RAG Implementation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Implementation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Implementation use?

RAG Implementation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Implementation use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to RAG Implementation?

Skills that share tags, products or a category with RAG Implementation: Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars), RAG Architect (Jeffallan/claude-skills, 12k stars), RAG Patterns (softspark/ai-toolkit, 179 stars) and Vector Database Engineer (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Implementation?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,254 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.