Agent skill

RAG Patterns

by softspark in softspark/ai-toolkit

RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.

Apache-2.0Auto-check passedAI & LLM Engineering

Install RAG Patterns

skills CLI
$ npx skills add softspark/ai-toolkit --skill rag-patterns -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install softspark/ai-toolkit rag-patterns --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/rag-patterns .claude/skills/rag-patterns && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-patterns
GitHub stars
179
Token cost
~1.8k tokens
SKILL.md length
603 words
Files
1
Skills in repo
112
Repo updated
First seen
Licence
Apache-2.0

At a glance

RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.

  • Works in 4 steps: Hybrid Search → Corrective RAG (CRAG) → HyDE (Hypothetical Document Embeddings) → …
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Core Patterns, Indexing Best Practices, MCP Tools Reference (v5.5.0) and Quality Metrics, plus 6 more sections
  • Calls docker, python and make

What it does

RAG Patterns is an agent skill from softspark/ai-toolkit. RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop. Triggers: RAG, embedding, pgvector, Qdrant, Pinecone, Weaviate, reranker, semantic search.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Embeddings and Vector databases. It works with pgvector, Pinecone, Qdrant and Weaviate. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Embeddings
  • Tasks that involve Vector databases

Example prompts

  • “/rag-patterns”

Requirements

  • Python 3
  • Docker
  • Pre-approved tools (allowed-tools): Read

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Hybrid Search
  2. Corrective RAG (CRAG)
  3. HyDE (Hypothetical Document Embeddings)
  4. Multi-hop Retrieval

What it can do on your machine

Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • python
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Patterns loads about 1.8k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 603 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 603 words, ~1,812 tokens.

Download SKILL.mdSave it as .claude/skills/rag-patterns/SKILL.md (or your agent's skills folder).
name
rag-patterns
description
RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop. Triggers: RAG, embedding, pgvector, Qdrant, Pinecone, Weaviate, reranker, semantic search.
allowed-tools
Read
effort
medium
user-invocable
false

RAG Patterns Skill

Core Patterns

Combine dense (vector) and sparse (BM25) retrieval with RRF fusion:

python
# RAG-MCP hybrid search
result = await hybrid_search_kb(
    query="rate limiting configuration",
    service="nginx",
    limit=10
)
2. Corrective RAG (CRAG)

Self-correcting retrieval with relevance validation:

python
result = await crag_search(
    query="fuzzy query",
    relevance_threshold=0.4,
    max_retries=2
)

# Or via smart_query
result = await smart_query(query="...", use_crag=True)
3. HyDE (Hypothetical Document Embeddings)

Generate hypothetical answers for better retrieval on conceptual queries:

python
result = await smart_query(
    query="conceptual question about design patterns",
    use_hyde=True
)
4. Multi-hop Retrieval

Complex queries requiring multiple retrieval steps:

python
result = await multi_hop_search(
    query="Compare nginx with varnish for Magento cache",
    max_hops=3
)

# Or via smart_query
result = await smart_query(query="compare A vs B", use_multi_hop=True)

Indexing Best Practices

AspectRecommendation
Chunk size512-1024 tokens
Overlap10-20% of chunk
StructurePreserve headers, sections
MetadataInclude title, path, date, category, tags
FrontmatterYAML with standardized fields
Frontmatter Template
yaml
---
title: "Document Title"
service: {project-name}
category: reference|howto|procedures|troubleshooting|decisions|best-practices
tags: [tag1, tag2, tag3]
last_updated: "YYYY-MM-DD"
---

MCP Tools Reference (v5.5.0)

ToolUse CaseSpeed
smart_query ⭐Default for 90% of queries2-4s
hybrid_search_kbRaw vector + text search<1s
get_documentFull document content<1s
crag_searchVague/fuzzy queries1-3s
multi_hop_searchComplex reasoning20-30s
Tool Selection Guide
python
# Default - auto-routing
smart_query("specific technical question")

# Vague query - self-correcting
crag_search("jak to skonfigurować")

# Complex comparison
multi_hop_search("nginx vs varnish performance comparison")

# Known document
get_document(path="kb/reference/architecture.md")

Quality Metrics

MetricDescriptionTarget
FaithfulnessAnswer based on context>70%
RelevancyAnswer addresses question>70%
Context PrecisionFound context is accurate>60%
Latency (p95)Response time<2s
Precision@kRelevant results in top-k>80%

Retrieval Optimization

Reranking
python
# Retrieve more, rerank to top-k
initial_results = await hybrid_search_kb(query, limit=20)
reranked = rerank_results(initial_results, query)
final_results = reranked[:5]
Context Window Management
  • Place critical info at start/end (serial position effect)
  • Summarize long documents before insertion
  • Use tiered context: critical → supporting → background
Query Enhancement
  • Query expansion with synonyms
  • Query decomposition for complex questions
  • Entity extraction for filtering

Anti-Patterns

❌ Don't:

  • Skip reranking for final results
  • Use very large chunks (>2000 tokens)
  • Ignore metadata in retrieval
  • Trust LLM output without citation
  • Use latest for model versions

✅ Do:

  • Use top-k=20 → rerank → top-5
  • Chunk semantically (by section)
  • Enrich metadata at indexing time
  • Require source attribution in answers
  • Pin model versions for reproducibility

RAG System Implementation

Key Files (Typical Structure)
scripts/
├── search_core.py           # Core search
├── query_enhancements.py    # HyDE, query expansion
├── corrective_rag.py        # CRAG
├── multi_hop.py             # Multi-hop
├── unified_indexer.py       # Indexing
└── rag_evaluator.py         # Evaluation
Running RAG Commands

Direct execution:

bash
# Index KB
make index

# Evaluate RAG
python scripts/evaluate_rag.py

# Detect gaps
python scripts/knowledge_gaps.py --detect

Docker execution (if containerized):

bash
# Index KB
docker exec {app-container} make index

# Evaluate RAG
docker exec {api-container} python3 scripts/evaluate_rag.py

# Detect gaps
docker exec {api-container} python3 scripts/knowledge_gaps.py --detect

Rules

  • MUST chunk by document structure (headers, lists, code fences), not by fixed byte/token count — structure-aware chunking recovers 20-40% of retrieval quality on technical docs
  • MUST always use hybrid search (BM25 + vector) for keyword-heavy queries — pure vector search misses exact identifiers (function names, config keys)
  • NEVER trust a single embedding model on multilingual corpora; pair with a bilingual model or translate queries at the edge
  • NEVER index without content-hash change detection — full rebuilds on every change waste embedding budget and corrupt orphan tracking
  • CRITICAL: every response includes verifiable citations (source path + exact chunk). A RAG answer without traceable sources is a hallucination wearing a badge.
  • MANDATORY: evaluate with a golden dataset (faithfulness, relevancy, context precision) before promoting any pipeline change to production
Show full SKILL.md (222 more words)Show less

Gotchas

  • Top-k cosine similarity is not relevance — semantically close chunks may be topically wrong. Always compare hybrid vs pure-vector scores on a held-out set before committing to one.
  • Default embedding models (e.g., text-embedding-ada-002) underperform on long technical docs (>8k tokens). For long-form content consider chunking before embedding, not embedding then slicing.
  • Chunk overlap (10-20%) helps narrative text but duplicates storage and token cost. Code and structured tables do not benefit from overlap — disable per content type.
  • Cross-encoder rerankers (e.g., bge-reranker) add 100-300ms per query. For real-time UX, rerank only the top-20 candidates, not the top-100.
  • RAG failure modes are structural (retrieval, routing, chunking), not prompt-level. Before "tuning the prompt", check retrieval metrics — a prompt fix on top of broken retrieval is theater.
  • Query rewriting (HyDE, hypothetical doc generation) improves some queries and degrades others. A/B test before enabling globally; a blanket "always rewrite" often regresses simple lookups.

When NOT to Load

  • For executing a reindex — use /index (task skill)
  • For measuring RAG quality — use /evaluate (task skill)
  • For chunking documentation strategy without an index — this skill assumes you already have a vector store; use /architecture-decision for pipeline choice
  • For MCP-specific retrieval via smart_query() — the tool is already built; reach for this skill only when tuning the underlying index
  • For prompt engineering alone without retrieval concerns — use /prompt-caching-patterns or the relevant language skill

© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in app/skills/rag-patterns of softspark/ai-toolkit.

Open the folder on GitHubat commit d64db2b

Compare with similar skills

RAG Patterns next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Patterns compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Patterns this skillsoftspark/ai-toolkit179—~1.8kAutomated safety check: PassApache-2.0
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
RAG Implementationwshobson/agents40k10 repos~1.1kAutomated safety check: PassMIT
Hunt RAG Vectorelementalsouls/Claude-BugHunter4.8k—~2.6kAutomated safety check: PassMIT
Agentsop Multi Tenant RAGagentsope/SkillAlchemy459—~9.8kAutomated safety check: PassMIT
Vector DBericrisco/rsc-harness167—~2.8kAutomated safety check: PassMIT

Similar skills

  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 10 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agentsop Multi Tenant RAG

    agentsope/SkillAlchemy

    Security-first SOP for multi-tenant RAG systems. An agent skill from agentsope/SkillAlchemy.

    459 GitHub stars~9.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vector DB

    ericrisco/rsc-harness

    A skill your agent uses when operating a vector store as a data layer — choosing or migrating between Pinecone, Qdrant, Weaviate and pgvector; designing a collection or index (distance metric…

    167 GitHub stars~2.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vector Database Engineer

    aiskillstore/marketplace

    Expert in vector databases, embedding strategies, and semantic search implementation.

    430 GitHub starsUsed in 8 repos~563 tokens
    DatabasesAuto-check passed

More from softspark/ai-toolkit

All 112 skills in this repo
  • Prepare Test Env

    softspark/ai-toolkit

    Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.

    179 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check: notes
  • A11y Validate

    softspark/ai-toolkit

    Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check: notes
  • Analyze

    softspark/ai-toolkit

    Analyzes code quality, complexity, patterns across codebase.

    179 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Autonomous Dev

    softspark/ai-toolkit

    Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.

    179 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check: notes
  • Brand Voice

    softspark/ai-toolkit

    Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • CI

    softspark/ai-toolkit

    Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).

    179 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes

Questions about RAG Patterns

What does RAG Patterns do?

RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop. RAG Patterns is an agent skill from softspark/ai-toolkit. RAG: embeddings, chunking, hybrid search (BM25+vector), reranking, CRAG, multi-hop.

When should I use RAG Patterns?

RAG Patterns fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Embeddings; tasks that involve Vector databases.

How do I install RAG Patterns in Claude Code?

Run `npx skills add softspark/ai-toolkit --skill rag-patterns -a claude-code`. Or copy the skill folder (app/skills/rag-patterns in softspark/ai-toolkit) into .claude/skills/rag-patterns in your project. Claude Code loads it when a task matches its description.

How do I install RAG Patterns in Codex?

Run `npx skills add softspark/ai-toolkit --skill rag-patterns -a codex`. Or copy the skill folder (app/skills/rag-patterns in softspark/ai-toolkit) into .agents/skills/rag-patterns in your project. Codex loads it when a task matches its description.

Can I use RAG Patterns in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill rag-patterns -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-patterns, .gemini/skills/rag-patterns, .github/skills/rag-patterns and .opencode/skills/rag-patterns in your project.

What does RAG Patterns need to run?

Going by SKILL.md and its folder, RAG Patterns needs the command-line tools its instructions call (docker, python and make). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read.

Does RAG Patterns access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is RAG Patterns safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Patterns use?

RAG Patterns is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Patterns use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Patterns?

Skills that share tags, products or a category with RAG Patterns: RAG Architect (Jeffallan/claude-skills, 12k stars), RAG Implementation (wshobson/agents, 40k stars), Hunt RAG Vector (elementalsouls/Claude-BugHunter, 4.8k stars) and Agentsop Multi Tenant RAG (agentsope/SkillAlchemy, 459 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Patterns?

softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.

Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.