Agent skill

RAG Eval

by glebis in glebis/claude-skills

Iterate on RAG systems with structured evals instead of eyeballing.

MITAuto-check passedAI & LLM Engineering

Install RAG Eval

skills CLI
$ npx skills add glebis/claude-skills --skill rag-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install glebis/claude-skills rag-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/rag-eval .claude/skills/rag-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-eval
GitHub stars
389
Token cost
~1.5k tokens
SKILL.md length
755 words
Files
2 (incl. scripts)
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Iterate on RAG systems with structured evals instead of eyeballing.

  • Works in 6 steps: (Optional) Ingest a prior iteration… → Audit the stack → Propose a sweep plan → …
  • Is tuning a RAG pipeline — changing retrieval prompts
  • SKILL.md covers Purpose, When to use, Prerequisites — gather before… and Workflow, plus 2 more sections
  • Runs Python scripts from its folder; calls python; needs OPENROUTER_API_KEY and OPENAI_API_KEY

What it does

RAG Eval is an agent skill from glebis/claude-skills. Iterate on RAG systems with structured evals instead of eyeballing. This skill should be used when the user is tuning a RAG pipeline — changing retrieval prompts, swapping models, adjusting chunking, or debugging poor answers — and wants a cheap, ranked set of experiments with cost tracking and structured feedback on the stack. Also use when the user asks "how do I know if my RAG is working?", "this RAG eval is burning money", or "what should I try next on retrieval?".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/session_ingest.py`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, LLM evaluation and LLM cost and token optimization. The repository describes itself as: Collection of Claude Code skills for enhanced AI workflows. The licence is MIT.

When your agent uses it

  • Is tuning a RAG pipeline — changing retrieval prompts
  • Swapping models
  • Adjusting chunking
  • Debugging poor answers — and wants a cheap

Example prompts

  • “how do I know if my RAG is working?”
  • “this RAG eval is burning money”
  • “what should I try next on retrieval?”
  • “/rag-eval”

Requirements

  • Python 3
  • A credential in OPENROUTER_API_KEY
  • A credential in OPENAI_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. (Optional) Ingest a prior iteration session
  2. Audit the stack
  3. Propose a sweep plan
  4. Run the sweep
  5. Rank and report
  6. Self-improve

What it can do on your machine

Read from SKILL.md and the folder at commit 7524dff. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Eval loads about 1.5k tokens when it runs. Until then it costs about 121 tokens; SKILL.md has 755 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~121
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from glebis/claude-skills at commit 7524dff, republished under its MIT licence (© glebis). 755 words, ~1,522 tokens.

Download SKILL.mdSave it as .claude/skills/rag-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
rag-eval
description
Iterate on RAG systems with structured evals instead of eyeballing. This skill should be used when the user is tuning a RAG pipeline — changing retrieval prompts, swapping models, adjusting chunking, or debugging poor answers — and wants a cheap, ranked set of experiments with cost tracking and structured feedback on the stack. Also use when the user asks "how do I know if my RAG is working?", "this RAG eval is burning money", or "what should I try next on retrieval?".

rag-eval

Purpose

Replace the "tweak → squint → swap model → burn credits" loop with a single command that runs a grid of eval variants on the user's gold-set, ranks them by a cost-aware score, and returns structured feedback on architecture, stack, and likely-issues. Draws on evidence-based RAG practices and learns from the user's past runs.

When to use

Trigger on: "help me test a RAG", "tune my RAG", "my RAG is bad", "compare retrieval prompts", "how do I eval this", "what's the best embedding model for X", "my RAG eval is expensive". Also trigger when the user reports burning OpenRouter / OpenAI credits with no clear signal of improvement.

Prerequisites — gather before running

Collect these from the user before the first sweep. Many are optional with sensible defaults; always confirm the ones that gate cost.

  1. RAG codebase root — path to the repo/module under test.
  2. Gold-set — at least 10 Q&A pairs. If missing, offer to generate a starter gold-set from the user's dataset (LLM-synthesized, human-reviewed). See references/best-practices.md.
  3. Dataset — the corpus the RAG retrieves over.
  4. Budget cap — hard dollar limit per run (default: $2 if user doesn't specify). Always confirm before any sweep.
  5. Provider keys — OPENROUTER_API_KEY or OPENAI_API_KEY (read from env).
  6. Vector-store config — collection name, embedding model, chunk size (read from repo; confirm if ambiguous).
  7. Eval history path (optional) — defaults to .rag-eval/history.jsonl in the repo root.

Workflow

Follow this order. Refer to references/best-practices.md for the canonical checklist and references/evidence-base.md for the research-backed defaults.

Step 0 — (Optional) Ingest a prior iteration session

When the user provides a session ID (Claude Code transcript, skill-studio session, or a Fathom meeting), run the deterministic ingest first — no LLM calls. This extracts only the useful signals (models tried, prompt variants, cost events, eval results) as compact JSON, so the rest of the skill works off a tiny structured bundle instead of a long raw transcript.

bash
python scripts/session_ingest.py <session_id> > /tmp/rag-eval-bundle.json
# or with a direct path:
python scripts/session_ingest.py --path /path/to/transcript.jsonl > /tmp/rag-eval-bundle.json

The bundle includes: models_tried, prompts_tried (hashes only), iterations, total_cost_usd, summary_stats. Feed this into Step 1 — do not paste the raw transcript.

Why this matters: transcripts can be 100k+ tokens of noise. The ingest script does regex extraction only, keeping the LLM budget for the actual audit + sweep planning. This is a hard requirement, not an optimization.

Step 1 — Audit the stack

Read references/best-practices.md and inspect the user's repo + vector-store config. Produce a structured report covering:

  • Architecture (retrieval type: dense / hybrid / rerank; chunking strategy; prompt structure)
  • Tech stack (embedding model, LLM, vector store)
  • Resources (dataset size, gold-set size, prior eval runs)
  • Risks (known anti-patterns, missing pieces)

Present the report to the user and ask which issues to address first.

Show full SKILL.md (325 more words)Show less
Step 2 — Propose a sweep plan

Based on the audit, propose 3–8 variants to test. Keep the grid small on the first run (default: 2 prompts × 2 models × 1 retrieval variant = 4 cells). Estimate cost using gold-set size × variants × avg tokens × provider pricing. Present the cost estimate and wait for user confirmation before running.

Step 3 — Run the sweep

Use scripts/eval_sweep.py (see the script header for invocation). It reads a config YAML, runs each variant against the gold-set, records per-variant cost and answer quality, and appends to history.jsonl.

Guardrails:

  • Never exceed the budget cap — halt mid-sweep if reached.
  • Never mutate the user's repo. Write all artifacts under .rag-eval/ (gitignore it).
  • Confirm before any sweep estimated to exceed the user's cap.
Step 4 — Rank and report

After the sweep, rank variants by a cost-aware score: quality × (1 / log(1 + cost)). Present:

  • Top 3 variants with quality metrics and cost
  • What changed vs the previous best
  • Concrete next experiment to try

Write the full report to .rag-eval/reports/<timestamp>.md.

Step 5 — Self-improve

Before each subsequent run, read history.jsonl and factor in what the user has already tried. Avoid re-testing rejected variants. Surface patterns ("models A, B, C all underperformed on multi-hop queries — next try a reranker").

Reusable resources

  • scripts/eval_sweep.py — grid-search runner. Reads eval_config.yaml, writes results to history.jsonl.
  • references/best-practices.md — evidence-based RAG checklist the agent uses as an anchor.
  • references/evidence-base.md — pointers to recent RAG research and when each technique helps.
  • assets/eval_config.template.yaml — starter config to copy into the user's repo.
  • assets/gold_set.template.jsonl — 3 example Q&A pairs to show the gold-set format.

Notes

  • Cost is the main failure mode. Never run without a confirmed budget. Err on the side of smaller sweeps; users can always run again.
  • No repo mutation. All outputs go under .rag-eval/ in the target repo.
  • When uncertain about best practices, do web research. Use tavily-search or firecrawl-research to pull current evidence, then synthesize into the audit report.
  • Defer to the user. Before changing any file in the target repo, always confirm.

© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in rag-eval of glebis/claude-skills.

  • SKILL.md
  • scripts/session_ingest.py

Open the folder on GitHubat commit 7524dff

Compare with similar skills

RAG Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Eval this skillglebis/claude-skills389—~1.5kAutomated safety check: PassMIT
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
Context Auditundefined-ui/second-brain-os1k—~802Automated safety check: PassMIT
Jd Gap Analysisstarkyru/learn-ai107—~1.9kAutomated safety check: PassMIT
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
GAIA Agent Benchmarkingamd/gaia1.6k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check passed
  • Context Audit

    undefined-ui/second-brain-os

    Audit an agent's context layout against the four places: system prompt, tools, history, tail.

    1k GitHub stars~802 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Jd Gap Analysis

    starkyru/learn-ai

    Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.

    107 GitHub stars~1.9k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    167 GitHub stars~3.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from glebis/claude-skills

All 91 skills in this repo
  • Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.

    389 GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.

    389 GitHub stars~973 tokensUpdated 12 days ago
    Auto-check passed
  • Deep Research

    glebis/claude-skills

    This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.

    389 GitHub stars~2.6k tokensUpdated 12 days ago
    Auto-check: notes
  • Elimination Research

    glebis/claude-skills

    This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…

    389 GitHub stars~1.6k tokensUpdated 12 days ago
    Auto-check passed
  • Narrated HTML Presentations

    glebis/claude-skills

    Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.

    389 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check: notes
  • Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

    389 GitHub stars~2.9k tokensUpdated 12 days ago
    Auto-check passed

Questions about RAG Eval

What does RAG Eval do?

Iterate on RAG systems with structured evals instead of eyeballing. RAG Eval is an agent skill from glebis/claude-skills. Iterate on RAG systems with structured evals instead of eyeballing.

When should I use RAG Eval?

RAG Eval fits situations like: is tuning a RAG pipeline — changing retrieval prompts; swapping models; adjusting chunking; debugging poor answers — and wants a cheap.

How do I install RAG Eval in Claude Code?

Run `npx skills add glebis/claude-skills --skill rag-eval -a claude-code`. Or copy the skill folder (rag-eval in glebis/claude-skills) into .claude/skills/rag-eval in your project. Claude Code loads it when a task matches its description.

How do I install RAG Eval in Codex?

Run `npx skills add glebis/claude-skills --skill rag-eval -a codex`. Or copy the skill folder (rag-eval in glebis/claude-skills) into .agents/skills/rag-eval in your project. Codex loads it when a task matches its description.

Can I use RAG Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill rag-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-eval, .gemini/skills/rag-eval, .github/skills/rag-eval and .opencode/skills/rag-eval in your project.

What does RAG Eval need to run?

Going by SKILL.md and its folder, RAG Eval needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named OPENROUTER_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENROUTER_API_KEY; A credential in OPENAI_API_KEY.

Does RAG Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does RAG Eval use?

RAG Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Eval use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Eval?

Skills that share tags, products or a category with RAG Eval: Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), Context Audit (undefined-ui/second-brain-os, 1k stars), Jd Gap Analysis (starkyru/learn-ai, 107 stars) and RAG Architect (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Eval?

glebis (a GitHub user) maintains it in glebis/claude-skills, which has 389 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on September 26, 2026.

Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.