Evaluate RAG
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).
$ npx skills add softspark/ai-toolkit --skill evaluate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install softspark/ai-toolkit evaluate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/evaluate .claude/skills/evaluate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluate" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/evaluate into .claude/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/softspark/ai-toolkit/tree/main/app/skills/evaluateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add softspark/ai-toolkit --skill evaluate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install softspark/ai-toolkit evaluate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/app/skills/evaluate .agents/skills/evaluate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluate" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/evaluate into .agents/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add softspark/ai-toolkit --skill evaluate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install softspark/ai-toolkit evaluate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/app/skills/evaluate .cursor/skills/evaluate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluate" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/evaluate into .cursor/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/softspark/ai-toolkit.git --path app/skills/evaluate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add softspark/ai-toolkit --skill evaluate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install softspark/ai-toolkit evaluate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/app/skills/evaluate .gemini/skills/evaluate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluate" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/evaluate into .gemini/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install softspark/ai-toolkit evaluateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add softspark/ai-toolkit --skill evaluate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/app/skills/evaluate .github/skills/evaluate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluate" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/evaluate into .github/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add softspark/ai-toolkit --skill evaluate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install softspark/ai-toolkit evaluate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/app/skills/evaluate .opencode/skills/evaluate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluate" agent skill from https://github.com/softspark/ai-toolkit/tree/main/app/skills/evaluate into .opencode/skills/evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluateEvaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).
Evaluate is an agent skill from softspark/ai-toolkit. Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision). Triggers: measure RAG quality, knowledge gap, RAG eval, golden dataset.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Retrieval-augmented generation and LLM evaluation. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3dockerFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluate loads about 1.1k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 333 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, ReadAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 333 words, ~1,135 tokens.
.claude/skills/evaluate/SKILL.md (or your agent's skills folder).Evaluate RAG quality using LLM-as-a-Judge methodology.
/evaluate [--threshold 0.7]# Run RAG evaluation
python3 scripts/evaluate_rag.py
# With custom thresholds
python3 scripts/evaluate_rag.py \
--faithfulness 0.7 \
--relevancy 0.7 \
--context 0.6
# Detect knowledge gaps
python3 scripts/knowledge_gaps.py --detect
# Generate gap report
python3 scripts/knowledge_gaps.py --report# Replace {api-container} with your API server container name
docker exec {api-container} python3 scripts/evaluate_rag.py
# With custom thresholds
docker exec {api-container} python3 scripts/evaluate_rag.py \
--faithfulness 0.7 \
--relevancy 0.7 \
--context 0.6
# Detect knowledge gaps
docker exec {api-container} python3 scripts/knowledge_gaps.py --detect
# Generate gap report
docker exec {api-container} python3 scripts/knowledge_gaps.py --report| Metric | Description | Target |
|---|---|---|
| Faithfulness | Is answer based on context? | >70% |
| Relevancy | Does answer address question? | >70% |
| Context Precision | Is found context accurate? | >60% |
Located at: scripts/golden_dataset.json (or project-specific path)
{
"queries": [
{
"query": "How to configure rate limiting?",
"expected_topics": ["nginx", "rate-limiting"],
"expected_sources": ["kb/nginx/howto/rate-limiting.md"]
}
]
}RAG Evaluation Results
======================
Total Queries: 50
Average Faithfulness: 0.82
Average Relevancy: 0.78
Average Context Precision: 0.71
Quality: GOOD
Failed Queries (faithfulness < 0.7):
- Query: "How to backup PostgreSQL?"
Score: 0.45
Issue: No relevant documents foundAfter evaluation, check for gaps:
# Direct execution
python3 scripts/knowledge_gaps.py --detect
# Docker execution
docker exec {api-container} python3 scripts/knowledge_gaps.py --detectOutput:
Knowledge Gaps Detected:
1. PostgreSQL backup procedures (5 failed queries)
2. Redis caching configuration (3 failed queries)
3. Ollama model selection (2 failed queries)temperature=0. Always report the average and stddev over ≥3 runs, not a one-shot number.expected_sources may point at moved or deleted paths. A sudden drop in context_precision across unrelated queries usually means dataset rot, not RAG regression — validate the dataset paths first.scripts/evaluate_skills.py/review or a tailored prompt/test© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in app/skills/evaluate of softspark/ai-toolkit.
Open the folder on GitHubat commit d64db2b
Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluate this skillsoftspark/ai-toolkit | 179 | — | ~1.1k | Automated safety check: Notes | Apache-2.0 | |
| Evaluate RAGai-evals-course/evals-skills | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| RAG ArchitectJeffallan/claude-skills | 12k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Jd Gap Analysisstarkyru/learn-ai | 105 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Agent Evalericrisco/rsc-harness | 156 | — | ~3.2k | Automated safety check: Pass | MIT | |
| RAG Observability Evalssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.1k | Automated safety check: Pass | MIT |
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
starkyru/learn-ai
Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
sickn33/agentic-awesome-skills
Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.
davepoon/buildwithclaude
Evaluate retrieval and citation behavior for RAG pipelines from deterministic JSONL fixtures.
softspark/ai-toolkit
Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.
softspark/ai-toolkit
Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.
softspark/ai-toolkit
Analyzes code quality, complexity, patterns across codebase.
softspark/ai-toolkit
Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.
softspark/ai-toolkit
Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.
softspark/ai-toolkit
Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).
Categories
Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision). Evaluate is an agent skill from softspark/ai-toolkit. Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).
Evaluate fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve LLM evaluation.
Run `npx skills add softspark/ai-toolkit --skill evaluate -a claude-code`. Or copy the skill folder (app/skills/evaluate in softspark/ai-toolkit) into .claude/skills/evaluate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add softspark/ai-toolkit --skill evaluate -a codex`. Or copy the skill folder (app/skills/evaluate in softspark/ai-toolkit) into .agents/skills/evaluate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate, .gemini/skills/evaluate, .github/skills/evaluate and .opencode/skills/evaluate in your project.
Going by SKILL.md and its folder, Evaluate needs the command-line tools its instructions call (python3 and docker). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read.
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Evaluate is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluate: Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), RAG Architect (Jeffallan/claude-skills, 12k stars), Jd Gap Analysis (starkyru/learn-ai, 105 stars) and Agent Eval (ericrisco/rsc-harness, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.
Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.