Agent skill

RAG Evaluate Quality

by lyonzin in lyonzin/knowledge-rag

Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats.

MITAuto-check passedAI & LLM Engineering

Install RAG Evaluate Quality

skills CLI
$ npx skills add lyonzin/knowledge-rag --skill rag-evaluate-quality -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lyonzin/knowledge-rag rag-evaluate-quality --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lyonzin/knowledge-rag.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/maintenance/rag-evaluate-quality .claude/skills/rag-evaluate-quality && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-evaluate-quality
GitHub stars
292
Token cost
~1.4k tokens
SKILL.md length
518 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats.

  • Works in 7 steps: Call get_index_stats(). The response… → Reuse an independently selected… → Serialize the array as a JSON string and… → …
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Workflow, Example report, Metrics endpoint and Related skills
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

RAG Evaluate Quality is an agent skill from lyonzin/knowledge-rag. Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats. Use after ingestion, a model or configuration change, or a reported search regression. Compare representative questions against a recorded baseline.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation. It works with Model Context Protocol. The repository describes itself as: Local RAG MCP server for Claude Code — hybrid search (semantic + BM25), cross-encoder reranking, 13 MCP tools, 20 format parsers. Zero external servers, zero API keys. The licence is MIT.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation

Example prompts

  • “/rag-evaluate-quality”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Call get_index_stats(). The response contains a stats object. Record stats.total_documents, stats.total_chunks, stats.embedding_model…
  2. Reuse an independently selected evaluation set, or prepare questions with known answer documents. Each case has one non-empty question and…
  3. Serialize the array as a JSON string and pass the test_cases parameter
  4. Inspect individual misses and rank changes before interpreting aggregates
  5. Compare with a previous run only after recording corpus revision, model, dimensions, query/passage prefixes, search configuration, and…
  6. Investigate changes before recommending a rebuild
  7. Save an evaluation report only where the user has authorized writing. Keep evaluation questions and results outside the measured corpus…

What it can do on your machine

Read from SKILL.md and the folder at commit df9cccb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Evaluate Quality loads about 1.4k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 518 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lyonzin/knowledge-rag at commit df9cccb, republished under its MIT licence (© lyonzin). 518 words, ~1,372 tokens.

Download SKILL.mdSave it as .claude/skills/rag-evaluate-quality/SKILL.md (or your agent's skills folder).
name
rag-evaluate-quality
description
Measure retrieval quality using evaluate_retrieval (MRR@5 and Recall@5) and get_index_stats. Use after ingestion, a model or configuration change, or a reported search regression. Compare representative questions against a recorded baseline.
metadata.type
rag-workflow
metadata.kind
maintenance
metadata.target
any-mcp-client

rag-evaluate-quality — measure, do not guess

Use this workflow when the user asks to evaluate retrieval, after a relevant change, or as part of an already authorized recurring evaluation. Do not infer that every session needs a benchmark or that invoking this skill creates a recurring schedule.

Workflow

  1. Call get_index_stats(). The response contains a stats object. Record stats.total_documents, stats.total_chunks, stats.embedding_model, stats.embedding_dim, and stats.query_cache.hit_rate. The cache rate describes repeated-query reuse, not relevance.

  2. Reuse an independently selected evaluation set, or prepare questions with known answer documents. Each case has one non-empty question and one non-empty expected path:

    json
    [
      {"query": "authentication design", "expected_filepath": "docs/adr/0018-auth.md"},
      {"query": "retry policy", "expected_filepath": "docs/adr/0031-retries.md"}
    ]

    These are illustrative paths. Verify that expected documents exist in the actual corpus. Do not derive the expected answer from whichever source the current search happens to rank first.

  3. Serialize the array as a JSON string and pass the test_cases parameter:

    python
    import json
    
    cases = [
        {"query": "authentication design", "expected_filepath": "docs/adr/0018-auth.md"},
        {"query": "retry policy", "expected_filepath": "docs/adr/0031-retries.md"},
    ]
    evaluate_retrieval(test_cases=json.dumps(cases))

    The MCP tool returns mrr_at_5, recall_at_5, total_queries, and per_query. It does not return Precision@5. Invalid cases are rejected before searches execute.

  4. Inspect individual misses and rank changes before interpreting aggregates:

    MetricMeaning
    MRR@5Mean reciprocal rank of the expected document; a miss contributes zero
    Recall@5Fraction of cases whose expected document occurs in the first five results
    found_at_rankRank for each case, or null if the expected document was not found

    A small corpus can still be evaluated. Report the number and coverage of questions; do not infer a universal quality threshold or declare a fixed delta statistically significant. With five cases, one changed result has a large effect.

  5. Compare with a previous run only after recording corpus revision, model, dimensions, query/passage prefixes, search configuration, and question set. The tool uses the server's default query settings; it does not accept hybrid_alpha, search_method, or min_score arguments. To compare those options, execute a separate controlled search workload with the same questions.

  6. Investigate changes before recommending a rebuild:

    ObservationNext check
    Expected document missingVerify file discovery/exclusions, parse errors, indexed chunks, and category
    Rank changed after new ingestionInspect competing hits and per-query evidence
    Poor results in a non-English corpusEvaluate an appropriate multilingual model and its required prefixes
    Model or passage prefix changedFollow the full model migration procedure; incremental indexing cannot convert old vectors
    Cache hit rate is zeroCheck whether queries actually repeat; no hit-rate target establishes retrieval quality
  7. Save an evaluation report only where the user has authorized writing. Keep evaluation questions and results outside the measured corpus unless deliberately testing their effect; indexing the answer key can contaminate later measurements.

Show full SKILL.md (104 more words)Show less

Example report

The following is a format example, not an observed benchmark:

text
Corpus revision: <commit or snapshot>
Configuration: <model, dimensions, prefixes, search settings>
Cases: 12 unchanged questions; 2 Portuguese, 10 English
MRR@5: <measured value> (previous <value>)
Recall@5: <measured value> (previous <value>)
Changed cases: <query IDs and before/after ranks>
Latency: <separately measured; specify warm/cold, sample count and units>
Next action: inspect <specific missed document or configuration change>

Do not say that nothing broke based only on a small retrieval set. Persistence, indexing completeness, concurrent access, memory, and runtime compatibility require their own checks.

Metrics endpoint

The current generic tool-duration metric provides knowledge_rag_tool_duration_seconds_count and knowledge_rag_tool_duration_seconds_sum; these support a mean, not a p95. The optional FTS5 latency histogram has knowledge_rag_fast_path_latency_seconds_bucket buckets. Do not invent a knowledge_rag_search_latency_seconds metric or derive percentiles from a count and sum.

© lyonzin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/maintenance/rag-evaluate-quality of lyonzin/knowledge-rag.

Open the folder on GitHubat commit df9cccb

Compare with similar skills

RAG Evaluate Quality next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Evaluate Quality compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Evaluate Quality this skilllyonzin/knowledge-rag292—~1.4kAutomated safety check: PassMIT
MCP Local RAGshinpr/mcp-local-rag411—~4.4kAutomated safety check: PassMIT
Local RAG Searchnkapila6/mcp-local-rag1341 repos~1.6kAutomated safety check: PassMIT
AutoRAG Setup and RepairMarker-Inc-Korea/AutoRAG5.1k—~5.5kAutomated safety check: PassMIT
Sciverseopendatalab/Sciverse-Agent-Tools119—~3kAutomated safety check: PassCustom licence
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Local RAG

    shinpr/mcp-local-rag

    Searches, saves, and maintains a local document index through a local RAG MCP server.

    411 GitHub stars~4.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Local RAG Search

    nkapila6/mcp-local-rag

    Efficiently perform web searches using the mcp-local-rag server with semantic similarity ranking.

    134 GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • AutoRAG Setup and Repair

    Marker-Inc-Korea/AutoRAG

    Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server.

    5.1k GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Sciverse

    opendatalab/Sciverse-Agent-Tools

    A skill your agent uses when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original…

    119 GitHub stars~3k tokensUpdated 20 days ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Postgres Hybrid Text Search

    timescale/pg-aiguide

    A skill your agent uses to implement hybrid search combining BM25 keyword search with semantic vector search using Reciprocal Rank Fusion (RRF).

    1.9k GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from lyonzin/knowledge-rag

All 10 skills in this repo
  • RAG Check First

    lyonzin/knowledge-rag

    Before answering any technical question, code request, architecture decision, or factual claim, call searchknowledge to check the local corpus.

    292 GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check passed
  • RAG Cite Sources

    lyonzin/knowledge-rag

    Every technical claim drawn from the local corpus must ship with a source citation formatted as path:line or path:section.

    292 GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check passed
  • RAG Code Review

    lyonzin/knowledge-rag

    When performing code review on a PR, diff, snippet, or "look at this change" request, first consult the corpus for related ADRs, coding standards, prior patterns, and similar files.

    292 GitHub stars~1.8k tokensUpdated 5 days ago
    Auto-check passed
  • RAG Deep Dive

    lyonzin/knowledge-rag

    Three-step multi-tool workflow — search the corpus, fetch the most relevant document in full, then find similar documents.

    292 GitHub stars~1.6k tokensUpdated 5 days ago
    Auto-check passed
  • RAG Troubleshoot

    lyonzin/knowledge-rag

    When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks.

    292 GitHub stars~1.8k tokensUpdated 5 days ago
    Auto-check passed
  • RAG Index Decisions

    lyonzin/knowledge-rag

    After making a non-obvious architectural decision, solving a novel bug, agreeing on a coding standard, or reaching a conclusion worth remembering, index it back into the knowledge base so the next…

    292 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check passed

Questions about RAG Evaluate Quality

What does RAG Evaluate Quality do?

Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats. RAG Evaluate Quality is an agent skill from lyonzin/knowledge-rag. Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats.

When should I use RAG Evaluate Quality?

RAG Evaluate Quality fits situations like: tasks that involve Retrieval-augmented generation.

How do I install RAG Evaluate Quality in Claude Code?

Run `npx skills add lyonzin/knowledge-rag --skill rag-evaluate-quality -a claude-code`. Or copy the skill folder (skills/maintenance/rag-evaluate-quality in lyonzin/knowledge-rag) into .claude/skills/rag-evaluate-quality in your project. Claude Code loads it when a task matches its description.

How do I install RAG Evaluate Quality in Codex?

Run `npx skills add lyonzin/knowledge-rag --skill rag-evaluate-quality -a codex`. Or copy the skill folder (skills/maintenance/rag-evaluate-quality in lyonzin/knowledge-rag) into .agents/skills/rag-evaluate-quality in your project. Codex loads it when a task matches its description.

Can I use RAG Evaluate Quality in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lyonzin/knowledge-rag --skill rag-evaluate-quality -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-evaluate-quality, .gemini/skills/rag-evaluate-quality, .github/skills/rag-evaluate-quality and .opencode/skills/rag-evaluate-quality in your project.

What does RAG Evaluate Quality need to run?

SKILL.md names no scripts, command-line tools or credentials: RAG Evaluate Quality is instructions for the agent only. Our summary lists: Python 3.

Does RAG Evaluate Quality access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Evaluate Quality safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Evaluate Quality use?

RAG Evaluate Quality is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Evaluate Quality use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Evaluate Quality?

Skills that share tags, products or a category with RAG Evaluate Quality: MCP Local RAG (shinpr/mcp-local-rag, 411 stars), Local RAG Search (nkapila6/mcp-local-rag, 134 stars), AutoRAG Setup and Repair (Marker-Inc-Korea/AutoRAG, 5.1k stars) and Sciverse (opendatalab/Sciverse-Agent-Tools, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Evaluate Quality?

lyonzin (a GitHub user) maintains it in lyonzin/knowledge-rag, which has 292 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: lyonzin/knowledge-rag on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.