Agent skill

RAG

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is…

MITAuto-check passedAI & LLM Engineering

Install RAG

skills CLI
$ npx skills add ericrisco/rsc-harness --skill rag -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness rag --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag .claude/skills/rag && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag
GitHub stars
156
Token cost
~2.9k tokens
SKILL.md length
1,206 words
Files
6 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is…

  • Building grounded Q&A over your own corpus — chunk
  • SKILL.md covers The pipeline, and where each…, Retrieval is the bottleneck —…, Chunking — the… and Contextual retrieval — prepend…, plus 5 more sections
  • Runs Shell scripts from its folder
  • Retrieve hybrid

What it does

RAG is an agent skill from ericrisco/rsc-harness. Use when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is retrieved but the answer is still wrong, invented, or unmeasured. NOT operating the store itself — collection schema, HNSW efsearch, quantization (that is vector-db).

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/evaluation.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Vector databases and LLM inference and serving. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Building grounded Q&A over your own corpus — chunk
  • Retrieve hybrid
  • Refuse when the sources fall short —
  • The right document is retrieved but the answer is still wrong

Example prompts

  • “/rag”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG loads about 2.9k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 88 tokens; SKILL.md has 1,206 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,206 words, ~2,896 tokens.

Download SKILL.mdSave it as .claude/skills/rag/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
rag
description
Use when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is retrieved but the answer is still wrong, invented, or unmeasured. NOT operating the store itself — collection schema, HNSW ef_search, quantization (that is `vector-db`).
tags
rag, retrieval-augmented-generation, chunking, hybrid-search, reranking, grounding, citations, faithfulness, contextual-retrieval
recommends
vector-db, embeddings-search, document-processing, chatbot, agent-eval
origin
risco

rag — own the retrieve → rerank → ground → cite → refuse pipeline

You own the pipeline that turns a corpus plus a question into a grounded, cited answer: chunk, optionally contextualize, index, retrieve hybrid, rerank, assemble a grounded prompt, cite the sources, and refuse when the context does not contain the answer.

You are judged by retrieval quality and answer faithfulness, not by raw vector math. If you find yourself tuning HNSW parameters, you wandered into the store underneath you (../vector-db/SKILL.md). If you are comparing embedding models or chunk sizes, that is the science beside you (embeddings-search).

The pipeline, and where each stage hands off

Each stage is a real branch — most failures live in one specific stage, and several stages delegate to a sibling skill rather than living here.

StageWhat you doHands off to
IngestGet clean text out of PDFs/DOCX/HTML/OCRyou assume text exists → ../document-processing/SKILL.md
ChunkHeading/semantic-aware splits with overlap, stable idsmodel, dims + chunk-size science → embeddings-search
ContextualizePrepend an LLM-written context blurb per chunk (optional)stays here
IndexEmbed + write dense vectors and a BM25/keyword indexyou upsert, it owns the knobs → ../vector-db/SKILL.md
RetrieveHybrid dense + BM25, fuse with RRF, top ~150hybrid query mechanics → ../vector-db/SKILL.md
RerankCross-encoder over the 150, keep top ~20stays here
Ground + citeSystem prompt: answer only from context, cite chunk idsstays here
RefuseOutput "I don't have enough information" on weak contextstays here
EvaluateFaithfulness, answer relevancy, context precision/recallgeneral harness → agent-eval

Three neighbors are not stages at all. Surfacing this answer inside a chat product (sessions, channels, UI) is ../chatbot/SKILL.md — it calls you, not the reverse. A multi-step tool loop with state, where retrieval is one tool among many, is ../building-agents/SKILL.md. Pulling schema-constrained fields out of text instead of a grounded prose answer is structured-extraction. rag is the retrieval brain those products call.

Retrieval is the bottleneck — measure recall before touching the prompt

Naive RAG pipelines fail at the retrieval step in up to ~40% of cases even when the correct document is in the corpus (StackAI / Lushbinary, 2026-06-02). Why this matters: if the right passage never reaches the model, no prompt wording can save the answer. So your first move on a broken pipeline is never the prompt — it is measuring whether retrieval delivered the goods.

text
Bad   → "answers are wrong, let me lower temperature and reword the system prompt."
Good  → measure context recall on a golden set; if the right chunk isn't in the top-K,
         fix chunking + hybrid + rerank first. Only then touch grounding.

Order of attack when answers are wrong: context recall → context precision → grounding prompt → generation params. The last item almost never moves the needle.

Chunking — the highest-leverage single fix

Chunk on structure, not on a blind character count. Why: too-small loses the context a passage needs to be interpretable; too-large dilutes the embedding so the relevant sentence gets averaged away (StackAI; EdenAI 2025, accessed 2026-06-02).

  • Split on headings/sections first, then sub-split long sections to a target window.
  • Keep overlap (~10–20% of the window) so a fact spanning a boundary survives in one chunk.
  • Attach a stable chunk_id and source at creation — you will need them end to end for citations (see below). Never embed text and discard where it came from.
python
# Heading-aware split sketch; real size tuning belongs in embeddings-search.
def chunk_markdown(doc_id, text, target=800, overlap=120):
    sections, buf, head = [], [], None
    for line in text.splitlines():
        if line.startswith("#"):
            if buf: sections.append((head, "\n".join(buf))); buf = []
            head = line.lstrip("# ").strip()
        else:
            buf.append(line)
    if buf: sections.append((head, "\n".join(buf)))
    out, i = [], 0
    for head, body in sections:
        for start in range(0, max(1, len(body)), target - overlap):
            piece = body[start:start + target]
            out.append({"chunk_id": f"{doc_id}#{i}", "source": doc_id,
                        "heading": head, "text": piece})
            i += 1
    return out

The "what window/overlap maximizes recall for this corpus" study is embeddings-search; here you just need structurally sane chunks that keep their ids.

Contextual retrieval — prepend context before you embed

Anthropic's Contextual Retrieval (Sept 2024) prepends a short LLM-generated blurb to each chunk before embedding and before BM25 indexing, so an isolated chunk knows what document and section it belongs to. Why it matters: it cuts failed retrievals by ~35% (contextual embeddings alone), ~49% (contextual embeddings + contextual BM25), and ~67% once reranking is added (anthropic.com/news/contextual-retrieval, 2024-09; accessed 2026-06-02).

text
<document>{{WHOLE_DOC}}</document>
Here is the chunk we want to situate within the whole document:
<chunk>{{CHUNK}}</chunk>
Give a short, standalone context (1–2 sentences) that situates this chunk within the
document for search retrieval. Answer only with the context, nothing else.

Embed context + "\n" + chunk_text (not the bare chunk). The full implementation — caching the document prompt, batching, the BM25 side, and a runnable retrieve → rerank → answer skeleton — lives in references/pipeline.md.

Hybrid retrieval + RRF + rerank

Dense vectors miss exact terms (codes, names, error strings); BM25 keyword search catches them but misses paraphrase. Combine both, fuse with Reciprocal Rank Fusion (RRF), then rerank with a cross-encoder. The default-best quality/cost funnel (Microsoft Cloud Blog 2025-02-04; StackAI, accessed 2026-06-02):

text
retrieve ~150 candidates (dense + BM25, fused with RRF)
   → rerank all 150 with a cross-encoder
   → keep top ~20 for the prompt

The actual hybrid query (named vectors, sparse-dense, server-side fusion) is vector-db. Here you own the funnel and the reranker choice:

  • Cohere Rerank 3.5 — managed cross-encoder, context length 4096, SOTA on BEIR and multilingual, available via Cohere API, Bedrock, Pinecone, Azure (docs.cohere.com/changelog/ rerank-v3.5, accessed 2026-06-02). Use when you want quality without hosting a model.
  • A local cross-encoder (e.g. a bge-reranker) — use when data cannot leave your network or you need zero per-call cost; you pay in GPU/latency instead.
python
# Rerank the fused candidates down to the prompt set; keep ids intact.
import cohere
co = cohere.ClientV2()
ranked = co.rerank(model="rerank-v3.5", query=q,
                   documents=[c["text"] for c in candidates], top_n=20)
top = [candidates[r.index] for r in ranked.results]  # each still carries chunk_id + source
Show full SKILL.md (470 more words)Show less

Ground the prompt — answer only from context, cite, refuse

Properly grounded RAG reduces hallucination rates by up to ~71%; poorly grounded pipelines still hallucinate in up to ~40% of responses even with the right doc retrieved (Confident AI / Maxim 2025, accessed 2026-06-02). The prompt must do three things: bind the answer to the context, force inline citations, and provide an explicit refusal path.

text
You answer ONLY using the information inside <context>. Do not use prior knowledge.
Cite every claim with the chunk id it came from, like [chunk_id]. Multiple ids are fine.
If the context does not contain enough information to answer, reply exactly:
"I don't have enough information in the provided sources to answer that."
(Spanish corpora: "No tengo suficiente información en las fuentes para responder.")

<context>
[doc12#3] {chunk text...}
[doc12#4] {chunk text...}
</context>

Question: {{question}}

A grounding prompt without a refusal path is a bug — it converts "missing context" into a confident fabrication. The refusal clause is what turns retrieval failures into honest non-answers.

Citations — carry ids the whole way

The answer can only link back if a stable id survives every stage: chunk → retrieve → rerank → prompt → answer. Why: if you embed text and drop the id at index time, there is nothing for a citation to point at, and you cannot debug which passage produced a wrong claim.

  • Put chunk_id and source on the chunk at creation; keep them on the object through rerank.
  • Render them into the <context> block ([chunk_id] text) so the model can quote them.
  • Map cited ids back to source URLs/pages when you render the answer to the user.

Evaluating it — metrics, not vibes

Faithfulness is not correctness: an answer can be faithful to a wrong chunk. Measure four RAGAS metrics on a small golden Q/A set (RAGAS docs; Cohorte 2025; Confident AI, accessed 2026-06-02):

Failure symptomMetric that catches itFirst fix
Answer states things not in the sourcesfaithfulness (claims supported by context)grounding prompt + refusal
Answer is on-topic but doesn't address the questionanswer relevancyprompt / query rewriting
Relevant chunks exist but rank below junkcontext precisionreranker, RRF weights
The needed chunk never gets retrievedcontext recallchunking, contextual retrieval, hybrid

Build a 30–50 question golden set with known-good answers, score with RAGAS, and gate CI below a threshold (e.g. faithfulness ≥ 0.90, context recall ≥ 0.85). Full formulas, thresholds, and the CI snippet are in references/evaluation.md. The general-purpose eval harness is agent-eval; the RAG-specific metrics live here.

Anti-patterns

Anti-patternWhy it breaksDo instead
Fixed char-count chunking, blind to structureSplits mid-sentence; dilutes embeddingsHeading/semantic chunks with overlap
Rerank disabled — stuff top-50 raw into promptNoise drowns the right passage; cost balloonsRetrieve ~150 → rerank → keep ~20
No refusal path in the grounding promptMissing context becomes confident fabricationExplicit "I don't have enough information"
Embed text, drop the chunk idNothing to cite or debugCarry chunk_id+source end to end
"It looks good" eval on vibesRegressions ship silentlyGolden set + RAGAS + CI threshold gate
Embedding the query differently from the corpusQuery and chunks land in different spacesSame model + same preprocessing both sides
Dense-only, ignoring BM25/keywordMisses exact codes/names/error stringsHybrid dense + BM25 fused with RRF
Tuning temperature to fix wrong answersGeneration is rarely the bottleneckMeasure context recall first

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/rag of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/evaluation.md
  • references/pipeline.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

RAG next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG this skillericrisco/rsc-harness156—~2.9kAutomated safety check: PassMIT
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0
Qdrant Search Qualitygithub/awesome-copilot40k1 repos~336Automated safety check: PassMIT
Neo4j Vector Index Skillneo4j-contrib/neo4j-skills114—~5.6kAutomated safety check: NotesMIT
Scholar RAGjoshzyj/open-scholar-skill167—~7.4kAutomated safety check: NotesCustom licence
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Qdrant Search Quality

    github/awesome-copilot

    Official

    Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~336 tokens
    AI & LLM EngineeringAuto-check passed
  • Neo4j Vector Index Skill

    neo4j-contrib/neo4j-skills

    Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN), store embeddings on nodes or relationships, use SEARCH clause (Neo4j 2026.01+, preferred) or…

    114 GitHub stars~5.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    167 GitHub stars~7.4k tokensUpdated 19 days ago
    Research & ScienceAuto-check: notes
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed

Questions about RAG

What does RAG do?

A skill your agent uses when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is…. RAG is an agent skill from ericrisco/rsc-harness. Use when building grounded Q&A over your own corpus — chunk, retrieve hybrid, rerank, ground, cite chunk ids, refuse when the sources fall short — or when the right document is retrieved but the answer is still wrong, invented, or unmeasured.

When should I use RAG?

RAG fits situations like: building grounded Q&A over your own corpus — chunk; retrieve hybrid; refuse when the sources fall short —; the right document is retrieved but the answer is still wrong.

How do I install RAG in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill rag -a claude-code`. Or copy the skill folder (skills/rag in ericrisco/rsc-harness) into .claude/skills/rag in your project. Claude Code loads it when a task matches its description.

How do I install RAG in Codex?

Run `npx skills add ericrisco/rsc-harness --skill rag -a codex`. Or copy the skill folder (skills/rag in ericrisco/rsc-harness) into .agents/skills/rag in your project. Codex loads it when a task matches its description.

Can I use RAG in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill rag -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag, .gemini/skills/rag, .github/skills/rag and .opencode/skills/rag in your project.

What does RAG need to run?

Going by SKILL.md and its folder, RAG needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does RAG access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does RAG use?

RAG is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to RAG?

Skills that share tags, products or a category with RAG: Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Qdrant Search Quality (github/awesome-copilot, 40k stars), Neo4j Vector Index Skill (neo4j-contrib/neo4j-skills, 114 stars) and Scholar RAG (joshzyj/open-scholar-skill, 167 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.