Agent skill

LLM Judge Validation

by ai-evals-course in ai-evals-course/evals-skills

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

Apache-2.0Auto-check passedAI & LLM Engineering

Install LLM Judge Validation

skills CLI
$ npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-evals-course/evals-skills validate-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/validate-evaluator .claude/skills/validate-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validate-evaluator
GitHub stars
1.5k
Token cost
~2.2k tokens
SKILL.md length
886 words
Files
2
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

  • Works in 8 steps: Create Data Splits → Run Evaluator on Dev Set → Measure TPR and TNR → …
  • Verifying an LLM judge agrees with expert labels before relying on it
  • SKILL.md covers Overview, Prerequisites, Core Instructions and Practical Guidance, plus 1 more section
  • Calls pip

What it does

Starting from human-labeled data, the skill splits it into three disjoint sets: training at 10 to 20 percent for few-shot examples, dev at 40 to 45 percent for iterating on the judge, and test at 40 to 45 percent, used once at the end. Splits are balanced between Pass and Fail even if real traffic is skewed, so there are enough Fail cases to measure TNR.

The judge runs on the dev set and is scored with TPR, how often it agrees with a human Pass, and TNR, how often it agrees with a human Fail. The loop repeats until both exceed 90 percent on dev, then the held-out test set is run once for final numbers and a bias correction formula is applied to production data. Examples use scikit-learn's train_test_split and confusion_matrix.

It assumes a judge prompt already built with write-judge-prompt and about 100 binary-labeled traces per failure mode, roughly 50 Pass and 50 Fail, with labels from a domain expert rather than outsourced annotators. Code-based evaluators are out of scope because they are deterministic and are tested with unit tests.

When your agent uses it

  • Verifying an LLM judge agrees with expert labels before relying on it
  • Measuring a judge's TPR and TNR on held-out data
  • Correcting production pass rates for a judge's known bias

Example prompts

  • “Split my 100 labeled traces into train, dev and test sets for judge validation.”
  • “Compute TPR and TNR for the judge on the dev set and tell me whether both clear the bar.”
  • “Apply the bias correction to the production pass rate using the test-set TPR and TNR.”

Requirements

  • A judge prompt built beforehand
  • Human-labeled traces with Pass or Fail labels from a domain expert
  • Python with scikit-learn

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Create Data Splits
  2. Run Evaluator on Dev Set
  3. Measure TPR and TNR
  4. Inspect Disagreements
  5. Iterate
  6. Final Measurement on Test Set
  7. (Optional): Estimate True Success Rate (Rogan-Gladen Correction)
  8. Confidence Interval

What it can do on your machine

Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Judge Validation loads about 2.2k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 886 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 886 words, ~2,225 tokens.

Download SKILL.mdSave it as .claude/skills/validate-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
validate-evaluator
description
Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you need to verify alignment before trusting its outputs. Do NOT use for code-based evaluators (those are deterministic; test with unit tests per `write-code-eval`).

Validate Evaluator

Calibrate an LLM judge against human judgment.

Overview

  1. Split human-labeled data into train (10-20%), dev (40-45%), test (40-45%)
  2. Run judge on dev set and measure TPR/TNR
  3. Iterate on the judge until TPR and TNR > 90% on dev set
  4. Run once on held-out test set for final TPR/TNR
  5. Apply bias correction formula to production data

Prerequisites

  • A built LLM judge prompt (from write-judge-prompt)
  • Human-labeled data: ~100 traces with binary Pass/Fail labels per failure mode
    • Aim for ~50 Pass and ~50 Fail (balanced, even if real distribution is skewed)
    • Labels must come from a domain expert, not outsourced annotators
  • Candidate few-shot examples from your labeled data

Core Instructions

Step 1: Create Data Splits

Split human-labeled data into three disjoint sets:

SplitSizePurposeRules
Training10-20% (~10-20 examples)Source of few-shot examples for the judge promptOnly clear-cut Pass and Fail cases. Used directly in the prompt.
Dev40-45% (~40-45 examples)Iterative evaluator refinementNever include in the prompt. Evaluate against repeatedly.
Test40-45% (~40-45 examples)Final unbiased accuracy measurementDo NOT look at during development. Used once at the end.

Target: 30-50 examples of each class (Pass and Fail) across dev and test combined. Use balanced splits even if real-world prevalence is skewed — you need enough Fail examples to measure TNR reliably.

python
from sklearn.model_selection import train_test_split

# First split: separate test set
train_dev, test = train_test_split(
    labeled_data, test_size=0.4, stratify=labeled_data['label'], random_state=42
)
# Second split: separate training examples from dev set
train, dev = train_test_split(
    train_dev, test_size=0.75, stratify=train_dev['label'], random_state=42
)
# Result: ~15% train, ~45% dev, ~40% test
Step 2: Run Evaluator on Dev Set

Run the judge on every example in the dev set. Compare predictions to human labels.

Step 3: Measure TPR and TNR

TPR (True Positive Rate): When a human says Pass, how often does the judge also say Pass?

TPR = (judge says Pass AND human says Pass) / (human says Pass)

TNR (True Negative Rate): When a human says Fail, how often does the judge also say Fail?

TNR = (judge says Fail AND human says Fail) / (human says Fail)
python
from sklearn.metrics import confusion_matrix

tn, fp, fn, tp = confusion_matrix(human_labels, evaluator_labels,
                                   labels=['Fail', 'Pass']).ravel()
tpr = tp / (tp + fn)
tnr = tn / (tn + fp)

Use TPR/TNR, not Precision/Recall or raw accuracy. These two metrics directly map to the bias correction formula. Use Cohen's Kappa only for measuring agreement between two human annotators, not for judge-vs-ground-truth.

Step 4: Inspect Disagreements

Examine every case where the judge disagrees with human labels:

Disagreement TypeJudgeHumanFix
False PassPassFailJudge is too lenient. Strengthen Fail definitions or add edge-case examples.
False FailFailPassJudge is too strict. Clarify Pass definitions or adjust examples.

For each disagreement, determine whether to:

  • Clarify wording in the judge prompt
  • Swap or add few-shot examples from the training set
  • Add explicit rules for the edge case
  • Split the criterion into more specific sub-checks
Step 5: Iterate

Refine the judge prompt and re-run on the dev set. Repeat until TPR and TNR stabilize.

Stopping criteria:

  • Target: TPR > 90% AND TNR > 90%
  • Minimum acceptable: TPR > 80% AND TNR > 80%

If alignment stalls:

ProblemSolution
TPR and TNR both lowUse a more capable LLM for the judge
One metric low, one acceptableInspect disagreements for the low metric specifically
Both plateau below targetDecompose the criterion into smaller, more atomic checks
Consistently wrong on certain input typesAdd targeted few-shot examples from training set
Labels themselves seem inconsistentRe-examine human labels; the rubric may need refinement
Step 6: Final Measurement on Test Set

Run the judge exactly once on the held-out test set. Record final TPR and TNR.

Do not iterate after seeing test set results. Go back to step 4 with new dev data if needed.

Show full SKILL.md (350 more words)Show less
Step 7 (Optional): Estimate True Success Rate (Rogan-Gladen Correction)

Raw judge scores on unlabeled production data are biased. If you need an accurate aggregate pass rate, correct for known judge errors:

theta_hat = (p_obs + TNR - 1) / (TPR + TNR - 1)

Where:

  • p_obs = fraction of unlabeled traces the judge scored as Pass
  • TPR, TNR = from test set measurement
  • theta_hat = corrected estimate of true success rate

Clip to [0, 1]. Invalid when TPR + TNR - 1 is near 0 (judge is no better than random).

Example:

  • Judge TPR = 0.92, TNR = 0.88
  • 500 production traces: 400 scored Pass -> p_obs = 0.80
  • theta_hat = (0.80 + 0.88 - 1) / (0.92 + 0.88 - 1) = 0.68 / 0.80 = 0.85
  • True success rate is ~85%, not the raw 80%
Step 8: Confidence Interval

Compute a bootstrap confidence interval. A point estimate alone is not enough.

python
import numpy as np

def bootstrap_ci(human_labels, eval_labels, p_obs, n_bootstrap=2000):
    """Bootstrap 95% CI for corrected success rate."""
    n = len(human_labels)
    estimates = []
    for _ in range(n_bootstrap):
        idx = np.random.choice(n, size=n, replace=True)
        h = np.array(human_labels)[idx]
        e = np.array(eval_labels)[idx]

        tp = ((h == 'Pass') & (e == 'Pass')).sum()
        fn = ((h == 'Pass') & (e == 'Fail')).sum()
        tn = ((h == 'Fail') & (e == 'Fail')).sum()
        fp = ((h == 'Fail') & (e == 'Pass')).sum()

        tpr_b = tp / (tp + fn) if (tp + fn) > 0 else 0
        tnr_b = tn / (tn + fp) if (tn + fp) > 0 else 0
        denom = tpr_b + tnr_b - 1

        if abs(denom) < 1e-6:
            continue
        theta = (p_obs + tnr_b - 1) / denom
        estimates.append(np.clip(theta, 0, 1))

    return np.percentile(estimates, 2.5), np.percentile(estimates, 97.5)

lower, upper = bootstrap_ci(test_human, test_eval, p_obs=0.80)
print(f"95% CI: [{lower:.2f}, {upper:.2f}]")

Or use judgy (pip install judgy):

python
from judgy import estimate_success_rate

# judgy expects 0/1 integer labels (1 = Pass, 0 = Fail)
test_labels = [1 if l == 'Pass' else 0 for l in test_human_labels]
test_preds = [1 if l == 'Pass' else 0 for l in test_eval_labels]
unlabeled_preds = [1 if l == 'Pass' else 0 for l in prod_eval_labels]

theta_hat, lower, upper = estimate_success_rate(
    test_labels, test_preds, unlabeled_preds
)
print(f"Corrected rate: {theta_hat:.2f}")
print(f"95% CI: [{lower:.2f}, {upper:.2f}]")

Practical Guidance

  • Pin exact model versions for LLM judges (a dated snapshot id like <model>-<YYYY-MM-DD>, not a floating alias). Providers update models without notice, causing silent drift.
  • Re-validate after changing the judge prompt, switching models, or when production confidence intervals widen unexpectedly.
  • Use ~100 labeled examples (50 Pass, 50 Fail). Below 60, confidence intervals become wide.
  • One trusted domain expert is the most efficient labeling path. If not feasible, have two annotators label 20-50 traces independently and resolve disagreements before proceeding.
  • Improving TPR narrows the confidence interval more than improving TNR. The correction divides by (TPR + TNR - 1), so a low TPR shrinks the denominator and amplifies estimation errors into wide CIs.

Anti-Patterns

  • Assuming judges "just work" without validation. A judge may consistently miss failures or flag passing traces.
  • Using raw accuracy or percent agreement. Use TPR and TNR. With class imbalance, raw accuracy is misleading.
  • Dev/test examples as few-shot examples. This is data leakage.
  • Reporting dev set performance as final accuracy. Dev numbers are optimistic. The test set gives the unbiased estimate.
  • Raw judge scores without bias correction. If you report an aggregate pass rate, apply the Rogan-Gladen formula (Step 7).
  • Point estimates without confidence intervals. A corrected rate of 85% could easily be 78-92% with small test sets. Report the range so stakeholders know how much to trust the number.

© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/validate-evaluator of ai-evals-course/evals-skills.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 80d5f7b

Compare with similar skills

LLM Judge Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Judge Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Judge Validation this skillai-evals-course/evals-skills1.5k—~2.2kAutomated safety check: PassApache-2.0
Genai Prompt Evaltimothywarner-org/claude-code224—~696Automated safety check: NotesMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Statistical Data Analysislingzhi227/agent-research-skills384—~886Automated safety check: PassNone
scikit-survival Time-to-Event Modelingdavila7/claude-code-templates32k12 repos~3.7kAutomated safety check: PassMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Genai Prompt Eval

    timothywarner-org/claude-code

    Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

    224 GitHub stars~696 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    384 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • scikit-survival Time-to-Event Modeling

    davila7/claude-code-templates

    Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.

    32k GitHub starsUsed in 12 repos~3.7k tokens
    Data & AnalyticsAuto-check passed
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Dataset Finder

    LeoYeAI/openclaw-master-skills

    A skill your agent uses when users need to search for datasets, download data files, or explore data repositories.

    2.2k GitHub stars~5.4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from ai-evals-course/evals-skills

All 9 skills in this repo
  • LLM Trace Review Interface

    ai-evals-course/evals-skills

    Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

    1.5k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 14 days ago
    Auto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 14 days ago
    Auto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • LLM-as-Judge Prompt Writer

    ai-evals-course/evals-skills

    Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

    1.5k GitHub stars~1.9k tokensUpdated 14 days ago
    Auto-check passed
  • Write Code Eval

    ai-evals-course/evals-skills

    Write code evaluators for known failure modes with objective rules.

    1.5k GitHub stars~385 tokensUpdated 14 days ago
    Auto-check passed

Questions about LLM Judge Validation

What does LLM Judge Validation do?

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. Starting from human-labeled data, the skill splits it into three disjoint sets: training at 10 to 20 percent for few-shot examples, dev at 40 to 45 percent for iterating on the judge, and test at 40 to 45 percent, used once at the end. Splits are balanced between Pass and Fail even if real traffic is skewed, so there are enough Fail cases to measure TNR.

When should I use LLM Judge Validation?

LLM Judge Validation fits situations like: verifying an LLM judge agrees with expert labels before relying on it; measuring a judge's TPR and TNR on held-out data; correcting production pass rates for a judge's known bias.

How do I install LLM Judge Validation in Claude Code?

Run `npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a claude-code`. Or copy the skill folder (skills/validate-evaluator in ai-evals-course/evals-skills) into .claude/skills/validate-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install LLM Judge Validation in Codex?

Run `npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a codex`. Or copy the skill folder (skills/validate-evaluator in ai-evals-course/evals-skills) into .agents/skills/validate-evaluator in your project. Codex loads it when a task matches its description.

Can I use LLM Judge Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill validate-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-evaluator, .gemini/skills/validate-evaluator, .github/skills/validate-evaluator and .opencode/skills/validate-evaluator in your project.

What does LLM Judge Validation need to run?

Going by SKILL.md and its folder, LLM Judge Validation needs the command-line tools its instructions call (pip). Our summary lists: A judge prompt built beforehand; Human-labeled traces with Pass or Fail labels from a domain expert; Python with scikit-learn.

Does LLM Judge Validation access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is LLM Judge Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Judge Validation use?

LLM Judge Validation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Judge Validation use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Judge Validation?

Skills that share tags, products or a category with LLM Judge Validation: Genai Prompt Eval (timothywarner-org/claude-code, 224 stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Statistical Data Analysis (lingzhi227/agent-research-skills, 384 stars) and scikit-survival Time-to-Event Modeling (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Judge Validation?

ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,468 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.

Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.