Evaluate RAG
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
Iterate on RAG systems with structured evals instead of eyeballing.
$ npx skills add glebis/claude-skills --skill rag-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install glebis/claude-skills rag-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/rag-eval .claude/skills/rag-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rag-eval" agent skill from https://github.com/glebis/claude-skills/tree/main/rag-eval into .claude/skills/rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/glebis/claude-skills/tree/main/rag-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add glebis/claude-skills --skill rag-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install glebis/claude-skills rag-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/rag-eval .agents/skills/rag-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rag-eval" agent skill from https://github.com/glebis/claude-skills/tree/main/rag-eval into .agents/skills/rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill rag-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install glebis/claude-skills rag-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/rag-eval .cursor/skills/rag-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rag-eval" agent skill from https://github.com/glebis/claude-skills/tree/main/rag-eval into .cursor/skills/rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/glebis/claude-skills.git --path rag-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add glebis/claude-skills --skill rag-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install glebis/claude-skills rag-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/rag-eval .gemini/skills/rag-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rag-eval" agent skill from https://github.com/glebis/claude-skills/tree/main/rag-eval into .gemini/skills/rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install glebis/claude-skills rag-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add glebis/claude-skills --skill rag-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/rag-eval .github/skills/rag-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rag-eval" agent skill from https://github.com/glebis/claude-skills/tree/main/rag-eval into .github/skills/rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill rag-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install glebis/claude-skills rag-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/rag-eval .opencode/skills/rag-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rag-eval" agent skill from https://github.com/glebis/claude-skills/tree/main/rag-eval into .opencode/skills/rag-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rag-evalIterate on RAG systems with structured evals instead of eyeballing.
RAG Eval is an agent skill from glebis/claude-skills. Iterate on RAG systems with structured evals instead of eyeballing. This skill should be used when the user is tuning a RAG pipeline — changing retrieval prompts, swapping models, adjusting chunking, or debugging poor answers — and wants a cheap, ranked set of experiments with cost tracking and structured feedback on the stack. Also use when the user asks "how do I know if my RAG is working?", "this RAG eval is burning money", or "what should I try next on retrieval?".
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/session_ingest.py`).
It sits in AI & LLM Engineering, covering Retrieval-augmented generation, LLM evaluation and LLM cost and token optimization. The repository describes itself as: Collection of Claude Code skills for enhanced AI workflows. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7524dff. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENROUTER_API_KEYOPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
RAG Eval loads about 1.5k tokens when it runs. Until then it costs about 121 tokens; SKILL.md has 755 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from glebis/claude-skills at commit 7524dff, republished under its MIT licence (© glebis). 755 words, ~1,522 tokens.
.claude/skills/rag-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Replace the "tweak → squint → swap model → burn credits" loop with a single command that runs a grid of eval variants on the user's gold-set, ranks them by a cost-aware score, and returns structured feedback on architecture, stack, and likely-issues. Draws on evidence-based RAG practices and learns from the user's past runs.
Trigger on: "help me test a RAG", "tune my RAG", "my RAG is bad", "compare retrieval prompts", "how do I eval this", "what's the best embedding model for X", "my RAG eval is expensive". Also trigger when the user reports burning OpenRouter / OpenAI credits with no clear signal of improvement.
Collect these from the user before the first sweep. Many are optional with sensible defaults; always confirm the ones that gate cost.
references/best-practices.md.OPENROUTER_API_KEY or OPENAI_API_KEY (read from env)..rag-eval/history.jsonl in the repo root.Follow this order. Refer to references/best-practices.md for the canonical checklist and references/evidence-base.md for the research-backed defaults.
When the user provides a session ID (Claude Code transcript, skill-studio session, or a Fathom meeting), run the deterministic ingest first — no LLM calls. This extracts only the useful signals (models tried, prompt variants, cost events, eval results) as compact JSON, so the rest of the skill works off a tiny structured bundle instead of a long raw transcript.
python scripts/session_ingest.py <session_id> > /tmp/rag-eval-bundle.json
# or with a direct path:
python scripts/session_ingest.py --path /path/to/transcript.jsonl > /tmp/rag-eval-bundle.jsonThe bundle includes: models_tried, prompts_tried (hashes only), iterations, total_cost_usd, summary_stats. Feed this into Step 1 — do not paste the raw transcript.
Why this matters: transcripts can be 100k+ tokens of noise. The ingest script does regex extraction only, keeping the LLM budget for the actual audit + sweep planning. This is a hard requirement, not an optimization.
Read references/best-practices.md and inspect the user's repo + vector-store config. Produce a structured report covering:
Present the report to the user and ask which issues to address first.
Based on the audit, propose 3–8 variants to test. Keep the grid small on the first run (default: 2 prompts × 2 models × 1 retrieval variant = 4 cells). Estimate cost using gold-set size × variants × avg tokens × provider pricing. Present the cost estimate and wait for user confirmation before running.
Use scripts/eval_sweep.py (see the script header for invocation). It reads a config YAML, runs each variant against the gold-set, records per-variant cost and answer quality, and appends to history.jsonl.
Guardrails:
.rag-eval/ (gitignore it).After the sweep, rank variants by a cost-aware score: quality × (1 / log(1 + cost)). Present:
Write the full report to .rag-eval/reports/<timestamp>.md.
Before each subsequent run, read history.jsonl and factor in what the user has already tried. Avoid re-testing rejected variants. Surface patterns ("models A, B, C all underperformed on multi-hop queries — next try a reranker").
scripts/eval_sweep.py — grid-search runner. Reads eval_config.yaml, writes results to history.jsonl.references/best-practices.md — evidence-based RAG checklist the agent uses as an anchor.references/evidence-base.md — pointers to recent RAG research and when each technique helps.assets/eval_config.template.yaml — starter config to copy into the user's repo.assets/gold_set.template.jsonl — 3 example Q&A pairs to show the gold-set format..rag-eval/ in the target repo.tavily-search or firecrawl-research to pull current evidence, then synthesize into the audit report.© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in rag-eval of glebis/claude-skills.
Open the folder on GitHubat commit 7524dff
RAG Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| RAG Eval this skillglebis/claude-skills | 389 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Evaluate RAGai-evals-course/evals-skills | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Context Auditundefined-ui/second-brain-os | 1k | — | ~802 | Automated safety check: Pass | MIT | |
| Jd Gap Analysisstarkyru/learn-ai | 107 | — | ~1.9k | Automated safety check: Pass | MIT | |
| RAG ArchitectJeffallan/claude-skills | 12k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| GAIA Agent Benchmarkingamd/gaia | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT |
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
undefined-ui/second-brain-os
Audit an agent's context layout against the four places: system prompt, tools, history, tail.
starkyru/learn-ai
Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
glebis/claude-skills
Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.
glebis/claude-skills
Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.
glebis/claude-skills
This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.
glebis/claude-skills
This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…
glebis/claude-skills
Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.
glebis/claude-skills
Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.
Categories
Iterate on RAG systems with structured evals instead of eyeballing. RAG Eval is an agent skill from glebis/claude-skills. Iterate on RAG systems with structured evals instead of eyeballing.
RAG Eval fits situations like: is tuning a RAG pipeline — changing retrieval prompts; swapping models; adjusting chunking; debugging poor answers — and wants a cheap.
Run `npx skills add glebis/claude-skills --skill rag-eval -a claude-code`. Or copy the skill folder (rag-eval in glebis/claude-skills) into .claude/skills/rag-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add glebis/claude-skills --skill rag-eval -a codex`. Or copy the skill folder (rag-eval in glebis/claude-skills) into .agents/skills/rag-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill rag-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-eval, .gemini/skills/rag-eval, .github/skills/rag-eval and .opencode/skills/rag-eval in your project.
Going by SKILL.md and its folder, RAG Eval needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named OPENROUTER_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENROUTER_API_KEY; A credential in OPENAI_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
RAG Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with RAG Eval: Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), Context Audit (undefined-ui/second-brain-os, 1k stars), Jd Gap Analysis (starkyru/learn-ai, 107 stars) and RAG Architect (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
glebis (a GitHub user) maintains it in glebis/claude-skills, which has 389 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on September 26, 2026.
Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.