Genai Prompt Eval
timothywarner-org/claude-code
Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
$ npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai-evals-course/evals-skills validate-evaluator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/validate-evaluator .claude/skills/validate-evaluator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "validate-evaluator" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/validate-evaluator into .claude/skills/validate-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-evaluator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai-evals-course/evals-skills/tree/main/skills/validate-evaluatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai-evals-course/evals-skills validate-evaluator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/validate-evaluator .agents/skills/validate-evaluator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "validate-evaluator" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/validate-evaluator into .agents/skills/validate-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-evaluator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai-evals-course/evals-skills validate-evaluator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/validate-evaluator .cursor/skills/validate-evaluator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "validate-evaluator" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/validate-evaluator into .cursor/skills/validate-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-evaluator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai-evals-course/evals-skills.git --path skills/validate-evaluator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai-evals-course/evals-skills validate-evaluator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/validate-evaluator .gemini/skills/validate-evaluator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "validate-evaluator" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/validate-evaluator into .gemini/skills/validate-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-evaluator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai-evals-course/evals-skills validate-evaluatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/validate-evaluator .github/skills/validate-evaluator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "validate-evaluator" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/validate-evaluator into .github/skills/validate-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-evaluator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai-evals-course/evals-skills validate-evaluator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/validate-evaluator .opencode/skills/validate-evaluator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "validate-evaluator" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/validate-evaluator into .opencode/skills/validate-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "validate-evaluator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
validate-evaluatorChecks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
Starting from human-labeled data, the skill splits it into three disjoint sets: training at 10 to 20 percent for few-shot examples, dev at 40 to 45 percent for iterating on the judge, and test at 40 to 45 percent, used once at the end. Splits are balanced between Pass and Fail even if real traffic is skewed, so there are enough Fail cases to measure TNR.
The judge runs on the dev set and is scored with TPR, how often it agrees with a human Pass, and TNR, how often it agrees with a human Fail. The loop repeats until both exceed 90 percent on dev, then the held-out test set is run once for final numbers and a bias correction formula is applied to production data. Examples use scikit-learn's train_test_split and confusion_matrix.
It assumes a judge prompt already built with write-judge-prompt and about 100 binary-labeled traces per failure mode, roughly 50 Pass and 50 Fail, with labels from a domain expert rather than outsourced annotators. Code-based evaluators are out of scope because they are deterministic and are tested with unit tests.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
LLM Judge Validation loads about 2.2k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 886 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 886 words, ~2,225 tokens.
.claude/skills/validate-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Calibrate an LLM judge against human judgment.
Split human-labeled data into three disjoint sets:
| Split | Size | Purpose | Rules |
|---|---|---|---|
| Training | 10-20% (~10-20 examples) | Source of few-shot examples for the judge prompt | Only clear-cut Pass and Fail cases. Used directly in the prompt. |
| Dev | 40-45% (~40-45 examples) | Iterative evaluator refinement | Never include in the prompt. Evaluate against repeatedly. |
| Test | 40-45% (~40-45 examples) | Final unbiased accuracy measurement | Do NOT look at during development. Used once at the end. |
Target: 30-50 examples of each class (Pass and Fail) across dev and test combined. Use balanced splits even if real-world prevalence is skewed — you need enough Fail examples to measure TNR reliably.
from sklearn.model_selection import train_test_split
# First split: separate test set
train_dev, test = train_test_split(
labeled_data, test_size=0.4, stratify=labeled_data['label'], random_state=42
)
# Second split: separate training examples from dev set
train, dev = train_test_split(
train_dev, test_size=0.75, stratify=train_dev['label'], random_state=42
)
# Result: ~15% train, ~45% dev, ~40% testRun the judge on every example in the dev set. Compare predictions to human labels.
TPR (True Positive Rate): When a human says Pass, how often does the judge also say Pass?
TPR = (judge says Pass AND human says Pass) / (human says Pass)TNR (True Negative Rate): When a human says Fail, how often does the judge also say Fail?
TNR = (judge says Fail AND human says Fail) / (human says Fail)from sklearn.metrics import confusion_matrix
tn, fp, fn, tp = confusion_matrix(human_labels, evaluator_labels,
labels=['Fail', 'Pass']).ravel()
tpr = tp / (tp + fn)
tnr = tn / (tn + fp)Use TPR/TNR, not Precision/Recall or raw accuracy. These two metrics directly map to the bias correction formula. Use Cohen's Kappa only for measuring agreement between two human annotators, not for judge-vs-ground-truth.
Examine every case where the judge disagrees with human labels:
| Disagreement Type | Judge | Human | Fix |
|---|---|---|---|
| False Pass | Pass | Fail | Judge is too lenient. Strengthen Fail definitions or add edge-case examples. |
| False Fail | Fail | Pass | Judge is too strict. Clarify Pass definitions or adjust examples. |
For each disagreement, determine whether to:
Refine the judge prompt and re-run on the dev set. Repeat until TPR and TNR stabilize.
Stopping criteria:
If alignment stalls:
| Problem | Solution |
|---|---|
| TPR and TNR both low | Use a more capable LLM for the judge |
| One metric low, one acceptable | Inspect disagreements for the low metric specifically |
| Both plateau below target | Decompose the criterion into smaller, more atomic checks |
| Consistently wrong on certain input types | Add targeted few-shot examples from training set |
| Labels themselves seem inconsistent | Re-examine human labels; the rubric may need refinement |
Run the judge exactly once on the held-out test set. Record final TPR and TNR.
Do not iterate after seeing test set results. Go back to step 4 with new dev data if needed.
Raw judge scores on unlabeled production data are biased. If you need an accurate aggregate pass rate, correct for known judge errors:
theta_hat = (p_obs + TNR - 1) / (TPR + TNR - 1)Where:
p_obs = fraction of unlabeled traces the judge scored as PassTPR, TNR = from test set measurementtheta_hat = corrected estimate of true success rateClip to [0, 1]. Invalid when TPR + TNR - 1 is near 0 (judge is no better than random).
Example:
Compute a bootstrap confidence interval. A point estimate alone is not enough.
import numpy as np
def bootstrap_ci(human_labels, eval_labels, p_obs, n_bootstrap=2000):
"""Bootstrap 95% CI for corrected success rate."""
n = len(human_labels)
estimates = []
for _ in range(n_bootstrap):
idx = np.random.choice(n, size=n, replace=True)
h = np.array(human_labels)[idx]
e = np.array(eval_labels)[idx]
tp = ((h == 'Pass') & (e == 'Pass')).sum()
fn = ((h == 'Pass') & (e == 'Fail')).sum()
tn = ((h == 'Fail') & (e == 'Fail')).sum()
fp = ((h == 'Fail') & (e == 'Pass')).sum()
tpr_b = tp / (tp + fn) if (tp + fn) > 0 else 0
tnr_b = tn / (tn + fp) if (tn + fp) > 0 else 0
denom = tpr_b + tnr_b - 1
if abs(denom) < 1e-6:
continue
theta = (p_obs + tnr_b - 1) / denom
estimates.append(np.clip(theta, 0, 1))
return np.percentile(estimates, 2.5), np.percentile(estimates, 97.5)
lower, upper = bootstrap_ci(test_human, test_eval, p_obs=0.80)
print(f"95% CI: [{lower:.2f}, {upper:.2f}]")Or use judgy (pip install judgy):
from judgy import estimate_success_rate
# judgy expects 0/1 integer labels (1 = Pass, 0 = Fail)
test_labels = [1 if l == 'Pass' else 0 for l in test_human_labels]
test_preds = [1 if l == 'Pass' else 0 for l in test_eval_labels]
unlabeled_preds = [1 if l == 'Pass' else 0 for l in prod_eval_labels]
theta_hat, lower, upper = estimate_success_rate(
test_labels, test_preds, unlabeled_preds
)
print(f"Corrected rate: {theta_hat:.2f}")
print(f"95% CI: [{lower:.2f}, {upper:.2f}]")<model>-<YYYY-MM-DD>, not a floating alias). Providers update models without notice, causing silent drift.(TPR + TNR - 1), so a low TPR shrinks the denominator and amplifies estimation errors into wide CIs.© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/validate-evaluator of ai-evals-course/evals-skills.
Open the folder on GitHubat commit 80d5f7b
LLM Judge Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLM Judge Validation this skillai-evals-course/evals-skills | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Genai Prompt Evaltimothywarner-org/claude-code | 224 | — | ~696 | Automated safety check: Notes | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Statistical Data Analysislingzhi227/agent-research-skills | 384 | — | ~886 | Automated safety check: Pass | None | |
| scikit-survival Time-to-Event Modelingdavila7/claude-code-templates | 32k | 12 repos | ~3.7k | Automated safety check: Pass | MIT | |
| Clawpathy AutoresearchClawBio/ClawBio | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT |
timothywarner-org/claude-code
Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
lingzhi227/agent-research-skills
Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.
davila7/claude-code-templates
Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.
ClawBio/ClawBio
Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.
LeoYeAI/openclaw-master-skills
A skill your agent uses when users need to search for datasets, download data files, or explore data repositories.
ai-evals-course/evals-skills
Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.
ai-evals-course/evals-skills
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
ai-evals-course/evals-skills
Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.
ai-evals-course/evals-skills
Write code evaluators for known failure modes with objective rules.
Works with
Categories
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. Starting from human-labeled data, the skill splits it into three disjoint sets: training at 10 to 20 percent for few-shot examples, dev at 40 to 45 percent for iterating on the judge, and test at 40 to 45 percent, used once at the end. Splits are balanced between Pass and Fail even if real traffic is skewed, so there are enough Fail cases to measure TNR.
LLM Judge Validation fits situations like: verifying an LLM judge agrees with expert labels before relying on it; measuring a judge's TPR and TNR on held-out data; correcting production pass rates for a judge's known bias.
Run `npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a claude-code`. Or copy the skill folder (skills/validate-evaluator in ai-evals-course/evals-skills) into .claude/skills/validate-evaluator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a codex`. Or copy the skill folder (skills/validate-evaluator in ai-evals-course/evals-skills) into .agents/skills/validate-evaluator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-evaluator, .gemini/skills/validate-evaluator, .github/skills/validate-evaluator and .opencode/skills/validate-evaluator in your project.
Going by SKILL.md and its folder, LLM Judge Validation needs the command-line tools its instructions call (pip). Our summary lists: A judge prompt built beforehand; Human-labeled traces with Pass or Fail labels from a domain expert; Python with scikit-learn.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
LLM Judge Validation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with LLM Judge Validation: Genai Prompt Eval (timothywarner-org/claude-code, 224 stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Statistical Data Analysis (lingzhi227/agent-research-skills, 384 stars) and scikit-survival Time-to-Event Modeling (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,468 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.
Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.