Agent skill

Evaluate

by softspark in softspark/ai-toolkit

Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Evaluate

skills CLI
$ npx skills add softspark/ai-toolkit --skill evaluate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install softspark/ai-toolkit evaluate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/softspark/ai-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/app/skills/evaluate .claude/skills/evaluate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate
GitHub stars
179
Token cost
~1.1k tokens
SKILL.md length
333 words
Files
1
Skills in repo
112
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).

  • Works in 4 steps: Generate test queries from golden dataset → Execute RAG pipeline for each query → LLM judges each response on metrics → …
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Usage, Execution, Metrics and Evaluation Process, plus 7 more sections
  • Calls python3 and docker

What it does

Evaluate is an agent skill from softspark/ai-toolkit. Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision). Triggers: measure RAG quality, knowledge gap, RAG eval, golden dataset.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and LLM evaluation. The repository describes itself as: Professional-grade AI coding toolkit: 94 skills, 44 agents, multi-platform (Claude, Cursor, Windsurf, Copilot, Gemini, Cline, Roo Code, Aider, Augment, Antigravity, Codex CLI… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve LLM evaluation

Example prompts

  • “Use the evaluate skill to evaluate RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision)”
  • “/evaluate”

Requirements

  • Python 3
  • Docker
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Generate test queries from golden dataset
  2. Execute RAG pipeline for each query
  3. LLM judges each response on metrics
  4. Report aggregate scores

What it can do on your machine

Read from SKILL.md and the folder at commit d64db2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate loads about 1.1k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 333 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from softspark/ai-toolkit at commit d64db2b, republished under its Apache-2.0 licence (© softspark). 333 words, ~1,135 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate/SKILL.md (or your agent's skills folder).
name
evaluate
description
Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision). Triggers: measure RAG quality, knowledge gap, RAG eval, golden dataset.
allowed-tools
Bash, Read
effort
medium
disable-model-invocation
true
argument-hint
[--threshold N]

RAG Evaluation

Evaluate RAG quality using LLM-as-a-Judge methodology.

Usage

/evaluate [--threshold 0.7]

Execution

bash
# Run RAG evaluation
python3 scripts/evaluate_rag.py

# With custom thresholds
python3 scripts/evaluate_rag.py \
  --faithfulness 0.7 \
  --relevancy 0.7 \
  --context 0.6

# Detect knowledge gaps
python3 scripts/knowledge_gaps.py --detect

# Generate gap report
python3 scripts/knowledge_gaps.py --report
Docker Execution (containerized projects)
bash
# Replace {api-container} with your API server container name
docker exec {api-container} python3 scripts/evaluate_rag.py

# With custom thresholds
docker exec {api-container} python3 scripts/evaluate_rag.py \
  --faithfulness 0.7 \
  --relevancy 0.7 \
  --context 0.6

# Detect knowledge gaps
docker exec {api-container} python3 scripts/knowledge_gaps.py --detect

# Generate gap report
docker exec {api-container} python3 scripts/knowledge_gaps.py --report

Metrics

MetricDescriptionTarget
FaithfulnessIs answer based on context?>70%
RelevancyDoes answer address question?>70%
Context PrecisionIs found context accurate?>60%

Evaluation Process

  1. Generate test queries from golden dataset
  2. Execute RAG pipeline for each query
  3. LLM judges each response on metrics
  4. Report aggregate scores

Golden Dataset

Located at: scripts/golden_dataset.json (or project-specific path)

json
{
  "queries": [
    {
      "query": "How to configure rate limiting?",
      "expected_topics": ["nginx", "rate-limiting"],
      "expected_sources": ["kb/nginx/howto/rate-limiting.md"]
    }
  ]
}

Output Example

RAG Evaluation Results
======================
Total Queries: 50
Average Faithfulness: 0.82
Average Relevancy: 0.78
Average Context Precision: 0.71

Quality: GOOD

Failed Queries (faithfulness < 0.7):
- Query: "How to backup PostgreSQL?"
  Score: 0.45
  Issue: No relevant documents found

Knowledge Gaps

After evaluation, check for gaps:

bash
# Direct execution
python3 scripts/knowledge_gaps.py --detect

# Docker execution
docker exec {api-container} python3 scripts/knowledge_gaps.py --detect

Output:

Knowledge Gaps Detected:
1. PostgreSQL backup procedures (5 failed queries)
2. Redis caching configuration (3 failed queries)
3. Ollama model selection (2 failed queries)

Quality Gates

  • Faithfulness >70%
  • Relevancy >70%
  • Context Precision >60%
  • No critical knowledge gaps

Rules

  • MUST use a golden dataset — never evaluate on synthetic queries only
  • NEVER report a score without listing the failed queries alongside it
  • CRITICAL: if the golden dataset is missing, stop and ask the user to provide one
  • MANDATORY: thresholds come from project config, not hardcoded defaults, when available

Gotchas

  • LLM-as-a-judge scores are non-deterministic; a single run fluctuates by ±10 points even with temperature=0. Always report the average and stddev over ≥3 runs, not a one-shot number.
  • The default threshold trio (0.7 / 0.7 / 0.6) was calibrated on English KBs. Multilingual corpora (Polish + English in the same index) score systematically 5-15 points lower — recalibrate per language, or split the golden dataset by language.
  • Golden datasets drift: when the KB is reindexed or documents are renamed, expected_sources may point at moved or deleted paths. A sudden drop in context_precision across unrelated queries usually means dataset rot, not RAG regression — validate the dataset paths first.
  • Judges often reward verbose answers as "more faithful" because there is more text to ground. Tune the judge prompt to penalize padding, or cap answer length in the generator before evaluation.

When NOT to Use

  • For auditing skill quality (the 5-criteria check) — that lives in scripts/evaluate_skills.py
  • For general-purpose LLM output scoring without a KB — use /review or a tailored prompt
  • For unit tests or code correctness — use /test
  • For continuous evaluation without a golden dataset — build the dataset first

© softspark, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in app/skills/evaluate of softspark/ai-toolkit.

Open the folder on GitHubat commit d64db2b

Compare with similar skills

Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate this skillsoftspark/ai-toolkit179—~1.1kAutomated safety check: NotesApache-2.0
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
Jd Gap Analysisstarkyru/learn-ai105—~1.9kAutomated safety check: PassMIT
Agent Evalericrisco/rsc-harness156—~3.2kAutomated safety check: PassMIT
RAG Observability Evalssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT

Similar skills

  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Jd Gap Analysis

    starkyru/learn-ai

    Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.

    105 GitHub stars~1.9k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • RAG Observability Evals

    sickn33/agentic-awesome-skills

    Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Evaluation Harness

    davepoon/buildwithclaude

    Evaluate retrieval and citation behavior for RAG pipelines from deterministic JSONL fixtures.

    3.6k GitHub stars~813 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from softspark/ai-toolkit

All 112 skills in this repo
  • Prepare Test Env

    softspark/ai-toolkit

    Prepare or verify a project QA environment with source identity, readiness, browser access, evidence paths and owned cleanup.

    179 GitHub stars~1.8k tokensUpdated today
    Auto-check: notes
  • A11y Validate

    softspark/ai-toolkit

    Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~3.8k tokensUpdated today
    Auto-check: notes
  • Analyze

    softspark/ai-toolkit

    Analyzes code quality, complexity, patterns across codebase.

    179 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Autonomous Dev

    softspark/ai-toolkit

    Drives a brief, specification, issue or existing PR through implementation, review, tests and QA to a ready PR.

    179 GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Brand Voice

    softspark/ai-toolkit

    Direct technical voice for docs, README, user-facing text. An agent skill from softspark/ai-toolkit.

    179 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • CI

    softspark/ai-toolkit

    Detect/generate/debug CI pipeline config (GitHub Actions, GitLab CI).

    179 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes

Questions about Evaluate

What does Evaluate do?

Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision). Evaluate is an agent skill from softspark/ai-toolkit. Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).

When should I use Evaluate?

Evaluate fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve LLM evaluation.

How do I install Evaluate in Claude Code?

Run `npx skills add softspark/ai-toolkit --skill evaluate -a claude-code`. Or copy the skill folder (app/skills/evaluate in softspark/ai-toolkit) into .claude/skills/evaluate in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate in Codex?

Run `npx skills add softspark/ai-toolkit --skill evaluate -a codex`. Or copy the skill folder (app/skills/evaluate in softspark/ai-toolkit) into .agents/skills/evaluate in your project. Codex loads it when a task matches its description.

Can I use Evaluate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add softspark/ai-toolkit --skill evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate, .gemini/skills/evaluate, .github/skills/evaluate and .opencode/skills/evaluate in your project.

What does Evaluate need to run?

Going by SKILL.md and its folder, Evaluate needs the command-line tools its instructions call (python3 and docker). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read.

Does Evaluate access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evaluate safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Evaluate use?

Evaluate is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate?

Skills that share tags, products or a category with Evaluate: Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), RAG Architect (Jeffallan/claude-skills, 12k stars), Jd Gap Analysis (starkyru/learn-ai, 105 stars) and Agent Eval (ericrisco/rsc-harness, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate?

softspark (a GitHub user) maintains it in softspark/ai-toolkit, which has 179 GitHub stars. The repository holds 112 skills in this directory. The repository was last updated on October 7, 2026.

Source: softspark/ai-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.