A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
Install the "agent-eval" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-eval into .claude/skills/agent-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add ericrisco/rsc-harness --skill agent-eval -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "agent-eval" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-eval into .agents/skills/agent-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add ericrisco/rsc-harness --skill agent-eval -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "agent-eval" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-eval into .cursor/skills/agent-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add ericrisco/rsc-harness --skill agent-eval -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "agent-eval" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-eval into .gemini/skills/agent-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add ericrisco/rsc-harness --skill agent-eval -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "agent-eval" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-eval into .github/skills/agent-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add ericrisco/rsc-harness --skill agent-eval -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "agent-eval" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/agent-eval into .opencode/skills/agent-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-eval", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
agent-eval
GitHub stars
156
Token cost
~3.2k tokens
SKILL.md length
1,403 words
Files
6 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT
At a glance
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
Measuring whether an LLM
SKILL.md covers Do NOT use — route instead, The eval anatomy, Build the dataset first and Choose the scorer — the…, plus 7 more sections
Runs Shell scripts from its folder
Agent system actually got better and gating merges on it: golden sets
What it does
Agent Eval is an agent skill from ericrisco/rsc-harness. Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is building-agents).
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/judge-design.md`).
It sits in AI & LLM Engineering, covering LLM evaluation, Building AI agents and Autonomous loops. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
When your agent uses it
Measuring whether an LLM
Agent system actually got better and gating merges on it: golden sets
Fixing an inflated LLM-as-judge
Scoring RAG (faithfulness
Example prompts
“/agent-eval”
Requirements
Python 3
A Bash shell
What it can do on your machine
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Agent Eval loads about 3.2k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 1,403 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~90
When it runs· the whole SKILL.md, loaded when a task matches
~3.2k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~6k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Download SKILL.mdSave it as .claude/skills/agent-eval/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
agent-eval
description
Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`).
tags
evals, llm, agents, llm-as-judge, regression-gate, ai
Turn "the agent feels better" into a number you can put in a PR check. You own the eval dataset, the scorer mix, the LLM-as-judge calibration, and the block-on-regression CI gate — framework-neutral, provider-neutral.
Do NOT use — route instead
The ask
Route to
Why it is not this skill
Build the agent loop, tools, RAG plumbing
building-agents
It builds the system; you score it. They cross-link.
"Make the answers shorter / rewrite the prompt"
prompt-engineering
Evals say it is worse; that skill changes the words. You never edit the prompt.
pytest/jest on deterministic functions
testing-py / testing-web
Assert-equals on pure code, not stochastic outputs scored by a judge.
Dashboards / tracing of live production traffic
observability
Online monitoring; you are offline + pre-merge.
Red-team, jailbreak, prompt injection
agent-safety
Adversarial coverage, not quality measurement.
Per-token cost budgets and accounting
cost-tracking
You report cost-per-task as one metric; the discipline lives there.
A/B stats on product/funnel metrics
ab-testing
Web experiments, not offline model comparison on a fixed set.
The eval anatomy
Every framework instantiates the same five-stage pipeline. Learn it once; the tool is a detail.
text
dataset ──▶ runner ──▶ scorers ──▶ metrics ──▶ gate
(JSONL (calls the (det / judge (aggregate + (pass/fail
golden system per / human) bootstrap CI) exit code)
set) case)
DeepEval, Inspect AI, and promptfoo are all just opinionated wrappers around this. If you understand the stages you can switch tools without relearning the craft.
Build the dataset first
Build your own golden set — a public leaderboard number is not your number, because identical model weights swing SWE-Bench Verified by 10–20 points just by changing the harness. Measure your task on your data. The dataset is the asset; everything else is replaceable. Rules:
50–200 hand-labeled cases per failure mode, not per total. Coverage of how the system fails beats raw volume. 80 real failure cases > 1000 generic ones.
Never synthetic-only. A set the model wrote will not surface the model's blind spots. Mine real traffic / tickets / transcripts and hand-label.
Version it in git as JSONL, a first-class reviewed asset — same as code. Diffs are reviewable; relabels are auditable.
Decontaminate. The eval set must not appear in training data or few-shot examples, or the score is a memorization artifact, not a capability.
Case schema — one JSON object per line:
jsonl
{"id":"refund-001","input":"Where is my refund for order 4821?","expected":"States refunds take 5-7 business days and asks for nothing already on file","context":["policy: refunds 5-7 business days"],"meta":{"failure_mode":"hallucinated_policy","source":"ticket#4821"}}
{"id":"refund-002","input":"Cancel my subscription and refund this month","expected":"Cancels, refunds prorated amount, confirms no future charge","context":["policy: prorated refund on cancel"],"meta":{"failure_mode":"missed_tool_call","source":"ticket#5190"}}
failure_mode in meta is what lets you slice metrics by mode and find which kind of bug regressed — not just that the aggregate dropped.
Bad: "generate 1000 test questions with GPT and use those." Good: "80 real failure-mode cases pulled from support tickets, hand-labeled, tagged by failure mode."
Choose the scorer — the 60/30/10 mix
Reach for the cheapest scorer that correlates with human judgment. Default mix:
Anything with a checkable shape: format, required fields, a known string, a budget
Free, instant, zero drift. Never spend a judge call on something a regex settles.
~30%
LLM-as-judge — G-Eval, DAG, custom Python scorer
Meaning: is this answer faithful, relevant, helpful
Only where correctness is semantic. Costs money and can drift — so calibrate it.
~10%
Human-in-the-loop
Genuinely ambiguous cases the judge disagrees on
The ground truth you calibrate the judge against.
One Scorer protocol, two implementations behind it — deterministic and judge are interchangeable to the runner:
python
from typing import Protocol
class Scorer(Protocol):
name: str
def score(self, case: dict, output: str) -> float: ... # 0.0–1.0
class JsonSchemaScorer:
name = "schema_valid"
def score(self, case, output): # deterministic, free, no drift
import json
try:
json.loads(output)
return 1.0
except ValueError:
return 0.0
class FaithfulnessJudge:
name = "faithfulness"
def __init__(self, judge_model): self.judge = judge_model
def score(self, case, output): # judge only where meaning matters
return self.judge.rate(case["context"], output) # see judge-design.md
LLM-as-judge you can trust
A score you do not trust is worse than no score: an uncalibrated judge gives false confidence, which is more dangerous than admitted ignorance. Each rule, with its why:
Judge model ≥ system under test. A weaker judge cannot reliably rank a stronger system — it scores noise.
The rubric must force a written rationale before the score. Rationale-first judging is what pushes judge–human agreement to ~85% — higher than two humans agree with each other. A bare number is a vibe with a decimal point.
Pairwise beats pointwise for stability. "Is A or B better?" is more reproducible than "rate A from 1–10," which inflates and clusters at 8–9.
Swap positions and average. Judges favor whichever answer came first; run A-then-B and B-then-A to cancel position bias.
Calibrate against human gold and report the agreement before you gate anything on the judge. Not a formality — this is the step that makes every number downstream defensible.
Bad judge prompt: "Rate this answer 1–10." → everything lands 8–9, useless.
Good: "Compare answer A and answer B against the reference. First write one sentence on each per the rubric, then output the better label." → forces reasoning, gives a stable signal.
Full rubric templates (pointwise + pairwise), the position-swap harness, the calibration script (agreement / Cohen's kappa vs human gold), G-Eval vs DAG, and the judge bias catalog (length, position, self-preference) with mitigations live in references/judge-design.md.
Agent and RAG scorers
Score the path, not only the destination. Beyond exact/judge:
RAG (DeepEval / RAGAS names):
Faithfulness — does the answer only claim what the retrieved context supports? Catches hallucination.
Answer relevancy — does it actually address the question, or drift?
Contextual recall / precision — did retrieval fetch the right chunks, and not bury them in noise? Separates a retrieval bug from a generation bug.
Agent:
Tool correctness — right tool, right arguments, right order.
Task completion / goal accuracy — did it finish the job, not just produce plausible text.
Trajectory scoring — grade the sequence of steps. A correct final answer from a wrong path will fail differently next time; only trajectory scoring catches it.
The system side of these (how the loop and tools are built) is ../building-agents/SKILL.md; a common system-under-test is ../chatbot/SKILL.md.
Show full SKILL.md (520 more words)Show less
The regression gate
Gate policy: block on regression vs a committed baseline, not on an absolute threshold. An absolute threshold flaps CI on judge noise and gives no signal on drift; "did this PR make a tracked metric worse than main?" is the question that matters.
Compute a bootstrap confidence interval on each metric so judge noise alone does not fail the build — only a drop beyond the CI counts.
The runner writes eval-report.json (metrics, per-failure-mode slices, baseline, pass/fail) and exits non-zero on a real regression so the merge is blocked.
python
import json, sys
def gate(current: dict, baseline: dict, margin: float = 0.0) -> int:
regressed = []
for metric, score in current.items():
if metric in baseline and score < baseline[metric] - margin:
regressed.append((metric, baseline[metric], score))
report = {"metrics": current, "baseline": baseline, "regressed": regressed,
"passed": not regressed}
with open("eval-report.json", "w") as f:
json.dump(report, f, indent=2)
if regressed:
for m, b, c in regressed:
print(f"REGRESSION {m}: {b:.3f} -> {c:.3f}", file=sys.stderr)
return 1
return 0
sys.exit(gate(run_eval(), json.load(open("eval-baseline.json"))))
The complete provider-neutral runner (JSONL loader, scorer registry, bootstrap-CI metrics), the GitHub Actions workflow, and side-by-side DeepEval-pytest + Inspect-AI Task/Solver/Scorer versions of the same eval live in references/runner-and-gate.md.
Framework cheat-sheet
Pick by where the eval runs and what it must do. Versions as of 2026-06 — re-verify, they rot.
You need human annotation queues and historical regression tracking.
The two-tool pattern is normal, not over-engineering: a light CI gate (DeepEval / RAGAS / promptfoo) plus a platform (Braintrust / LangSmith / Arize) for annotation and history. They share data; different jobs.
Anti-patterns
Anti-pattern
Why it bites
Do instead
Vibes-gating ("feels better, merge it")
No artifact to defend or reproduce
Gate on a number from a committed dataset
Synthetic-only dataset
Model-written cases miss the model's blind spots
Hand-label real traffic by failure mode
Uncalibrated judge
Confident wrong scores; worse than none
Report agreement vs human gold first
Judge weaker than system
Cannot rank a stronger system; scores noise
Judge model ≥ system under test
Absolute-threshold gate
Flaps CI on judge noise, blind to drift
Block on regression vs baseline + bootstrap CI
Shipping on a leaderboard number
Harness effect = 10–20pt swing
Build your own golden set
Scoring only the final answer
A right answer from a wrong path regresses later
Score the trajectory too
Never relabeling drifted gold
Stale "truth" silently rots the gate
Review and relabel the golden set on a schedule
Project grounding
If the workspace has a 02-DOCS/ harness, record the eval policy in 02-DOCS/wiki/stack/evals.md: dataset location, scorer mix, gate baseline file, judge model, and the failure modes covered. Follow the harness wiki-article-template.md (type: stack) and index it in 02-DOCS/wiki/index.md. This is recorded, not gated — skip silently if there is no harness.
verify.sh
scripts/verify.sh is read-only and tool-detecting. It validates that every *.jsonl golden set in the project parses and that each line carries the required id, input, expected keys; checks the shape of any eval-report.json; and runs ruff / mypy on example Python and markdownlint on docs when those tools are installed. Every missing tool prints a yellow WARN and is skipped — never a failure. An empty or clean target exits 0.
Agent Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt…
A skill your agent uses when a support or sales bot on a live website must behave: persona/system prompt, grounding so it cannot invent prices or policy, jailbreak and injection defense, the human…
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
A skill your agent uses when building, refactoring, or debugging Angular (v20/21+): standalone components, signals, zoneless change detection, @if/@for/@defer control flow, inject() DI…
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…. Agent Eval is an agent skill from ericrisco/rsc-harness. Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework.
When should I use Agent Eval?
Agent Eval fits situations like: measuring whether an LLM; agent system actually got better and gating merges on it: golden sets; fixing an inflated LLM-as-judge; scoring RAG (faithfulness.
How do I install Agent Eval in Claude Code?
Run `npx skills add ericrisco/rsc-harness --skill agent-eval -a claude-code`. Or copy the skill folder (skills/agent-eval in ericrisco/rsc-harness) into .claude/skills/agent-eval in your project. Claude Code loads it when a task matches its description.
How do I install Agent Eval in Codex?
Run `npx skills add ericrisco/rsc-harness --skill agent-eval -a codex`. Or copy the skill folder (skills/agent-eval in ericrisco/rsc-harness) into .agents/skills/agent-eval in your project. Codex loads it when a task matches its description.
Can I use Agent Eval in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill agent-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-eval, .gemini/skills/agent-eval, .github/skills/agent-eval and .opencode/skills/agent-eval in your project.
What does Agent Eval need to run?
Going by SKILL.md and its folder, Agent Eval needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.
Does Agent Eval access the network?
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Is Agent Eval safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
What licence does Agent Eval use?
Agent Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Agent Eval use?
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.
What are the alternatives to Agent Eval?
Skills that share tags, products or a category with Agent Eval: Agent Harness Design (AnastasiyaW/codex-claude-code-config, 154 stars), Jd Gap Analysis (starkyru/learn-ai, 105 stars), Building Agent Systems (telagod/code-abyss, 244 stars) and Chatbot (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Agent Eval?
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.