Agent skill

Output Eval Validate Judge

by growthxai in growthxai/output

Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Output Eval Validate Judge

skills CLI
$ npx skills add growthxai/output --skill output-eval-validate-judge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install growthxai/output output-eval-validate-judge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge .claude/skills/output-eval-validate-judge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
output-eval-validate-judge
GitHub stars
442
Token cost
~2.5k tokens
SKILL.md length
1,129 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.

  • Works in 6 steps: Create Data Splits → Run the Judge on Dev Set → Compute TPR and TNR → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Overview, Prerequisites, Step 1: Create Data Splits and Step 2: Run the Judge on Dev Set, plus 6 more sections
  • Calls npx

What it does

Output Eval Validate Judge is an agent skill from growthxai/output. Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Use after writing a judge prompt to verify it agrees with human judgment.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/output-eval-validate-judge”

Requirements

  • Node.js
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Create Data Splits
  2. Run the Judge on Dev Set
  3. Compute TPR and TNR
  4. Inspect Disagreements
  5. Iterate
  6. Final Measurement on Test Set

What it can do on your machine

Read from SKILL.md and the folder at commit 99ee298. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Output Eval Validate Judge loads about 2.5k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 1,129 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from growthxai/output at commit 99ee298, republished under its Apache-2.0 licence (© growthxai). 1,129 words, ~2,451 tokens.

Download SKILL.mdSave it as .claude/skills/output-eval-validate-judge/SKILL.md (or your agent's skills folder).
name
output-eval-validate-judge
description
Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Use after writing a judge prompt to verify it agrees with human judgment.
allowed-tools
Bash, Read, Write, Edit

Validating LLM Judges

Overview

An LLM judge is only useful if it agrees with human judgment. This skill walks you through calibrating a judge against human-labeled data using True Positive Rate (TPR) and True Negative Rate (TNR) metrics. Do this before trusting any judgeVerdict(), judgeScore(), or judgeLabel() evaluator in your eval suite.

Prerequisites

  1. A judge .prompt file — Written following output-eval-judge-prompt
  2. ~100 human-labeled traces — With binary pass/fail labels for the failure mode this judge targets. Aim for ~50 pass and ~50 fail. Minimum: 20 pass and 20 fail.
  3. Labels stored in dataset YAML — Each dataset has ground_truth.evals.<evaluator_name>.verdict: pass or fail

This process applies only to LLM-based judges. For code-based Verdict.* evaluators, write unit tests instead.

Step 1: Create Data Splits

Split your labeled datasets into three groups:

Split% of DataPurposeExample (100 datasets)
Train10-20%Source of few-shot examples in the judge prompt15 datasets
Dev40-45%Iterate on judge prompt, measure TPR/TNR42 datasets
Test40-45%Final held-out measurement, run once43 datasets
Organizing splits

Use a naming convention or subdirectories to separate splits:

Option A: Name prefixes

tests/datasets/
├── train_formal_pass_01.yml
├── train_casual_fail_01.yml
├── dev_technical_pass_01.yml
├── dev_ambiguous_fail_01.yml
├── test_simple_pass_01.yml
├── test_contradictory_fail_01.yml
└── ...

Option B: Subdirectories

tests/datasets/
├── train/
│   ├── formal_pass_01.yml
│   └── casual_fail_01.yml
├── dev/
│   ├── technical_pass_01.yml
│   └── ambiguous_fail_01.yml
└── test/
    ├── simple_pass_01.yml
    └── contradictory_fail_01.yml
Splitting rules
  • Balance pass/fail in each split — Don't put all failures in dev and all passes in test
  • Randomize — Don't sort by difficulty or topic
  • Training examples in the prompt — Use only train-split examples as few-shot in the judge .prompt file. Never use dev or test examples — that's data leakage
  • Lock the test split — Once created, do not look at test data until final measurement

Step 2: Run the Judge on Dev Set

Execute the eval workflow against only the dev-split datasets:

bash
# Run with cached output on dev datasets
npx output workflow test <workflowName> --cached \
  --dataset dev_technical_pass_01,dev_ambiguous_fail_01,dev_formal_pass_02,...

Or if using subdirectories, list the dev dataset names:

bash
npx output workflow test <workflowName> --cached \
  --dataset $(ls tests/datasets/dev/ | sed 's/.yml//' | tr '\n' ',')

Save the output. You need the judge's verdict for each dataset to compare against ground truth.

Extracting results

Use --json to get machine-readable results:

bash
npx output workflow test <workflowName> --cached --dataset <dev_datasets> --json

The output includes per-dataset, per-evaluator verdicts that you can compare against ground_truth.evals.<evaluator_name>.verdict.

Step 3: Compute TPR and TNR

For the evaluator you're validating, build a confusion matrix from the dev results.

Definitions

Using "fail" as the positive class (what you're trying to detect):

Judge says FailJudge says Pass
Human says FailTrue Positive (TP)False Negative (FN)
Human says PassFalse Positive (FP)True Negative (TN)

TPR (True Positive Rate) = TP / (TP + FN)

  • "Of all the real failures, what fraction did the judge catch?"
  • Low TPR means the judge misses real failures (dangerous)

TNR (True Negative Rate) = TN / (TN + FP)

  • "Of all the real passes, what fraction did the judge correctly approve?"
  • Low TNR means the judge flags passing traces as failures (noisy)
Example computation

Dev set results for check_tone evaluator (42 datasets):

Judge: FailJudge: Pass
Human: Fail18 (TP)3 (FN)
Human: Pass2 (FP)19 (TN)
  • TPR = 18 / (18 + 3) = 85.7%
  • TNR = 19 / (19 + 2) = 90.5%
Why not raw accuracy?

Raw accuracy = (TP + TN) / total = (18 + 19) / 42 = 88.1%

This looks fine, but masks problems. If your dataset were 90% pass (class imbalance), a judge that always says "pass" would get 90% accuracy while catching zero failures (TPR = 0%). TPR and TNR measure what actually matters: catching failures and not crying wolf.

Step 4: Inspect Disagreements

For every case where the judge disagrees with the human label, determine the root cause.

False Negatives (judge missed a real failure)

The judge said "pass" but the human said "fail." For each:

  1. Read the trace and the judge's critique
  2. Determine why the judge missed it:
    • Criterion too narrow — The prompt defines failure too narrowly. Broaden the fail definition.
    • Missing few-shot example — The failure pattern isn't represented in examples. Add a similar borderline example from the train split.
    • Insufficient context — The judge doesn't have the information needed to detect this failure. Add the missing variable to the prompt.
False Positives (judge flagged a passing trace)

The judge said "fail" but the human said "pass." For each:

  1. Read the trace and the judge's critique
  2. Determine why the judge flagged it:
    • Criterion too broad — The prompt defines failure too broadly. Tighten the fail definition.
    • Misleading few-shot example — A borderline example is being overgeneralized. Clarify or replace it.
    • Overly strict — The judge applies the criterion more strictly than intended. Add explicit exceptions to the prompt.
Show full SKILL.md (433 more words)Show less
Logging disagreements

Track each disagreement to guide prompt iteration:

DatasetHumanJudgeRoot CauseFix
dev_technical_pass_03passfailJudge flagged "it's" as casual but context was a direct quoteAdd exception: "Contractions within direct quotes are acceptable"
dev_ambiguous_fail_02failpassJudge missed subtle tone shift in paragraph 3Add borderline few-shot example showing mid-text tone drift

Step 5: Iterate

Apply the fixes from Step 4 to the judge .prompt file. Then re-run on the dev set:

bash
npx output workflow test <workflowName> --cached --dataset <dev_datasets>

Recompute TPR and TNR. Repeat until both metrics meet the target.

Targets
MetricTargetMinimum Acceptable
TPR> 90%> 80%
TNR> 90%> 80%

If you can't reach 80%/80% after 3-4 iterations:

  1. Upgrade the model — Switch from Haiku to Sonnet in the .prompt frontmatter
  2. Split the criterion — The failure mode may contain two distinct sub-failures that need separate judges
  3. Revisit the labels — Some human labels may be inconsistent. Re-label disagreements with a second reviewer
Iteration checklist

Each iteration:

  • Identified root cause for each disagreement
  • Applied targeted fix to .prompt file (not random changes)
  • Re-ran on dev set
  • Recomputed TPR and TNR
  • Logged the iteration number, changes made, and resulting metrics

Step 6: Final Measurement on Test Set

Once dev metrics meet the target, run the judge on the held-out test set exactly once:

bash
npx output workflow test <workflowName> --cached --dataset <test_datasets> --json

Compute TPR and TNR on the test results. Record these as the final metrics.

Interpreting final results
  • Test metrics close to dev metrics — The judge generalizes well. Ship it.
  • Test metrics significantly lower — The judge may be overfit to dev set patterns. Do not iterate on the test set. Instead, gather more labeled data, re-split, and restart from Step 2.
Recording results

Document the final validation results alongside the judge prompt:

markdown
# Validation: check_tone (judge_tone@v1.prompt)
# Date: 2026-03-25
# Model: claude-haiku-4-5-20251001

## Dev Set (42 datasets)
- TPR: 90.5% (19/21)
- TNR: 95.2% (20/21)

## Test Set (43 datasets)
- TPR: 88.0% (22/25)
- TNR: 94.4% (17/18)

## Conclusion: APPROVED — both metrics above 80% minimum

Store this in a VALIDATION.md file next to the judge prompt or in the evaluator's documentation.

Anti-Patterns

  • Assuming judges work without validation — An unvalidated judge may consistently miss failures or flag passing traces
  • Using dev/test examples as few-shot — Data leakage inflates metrics and hides real performance
  • Optimizing for raw accuracy — Use TPR and TNR instead; accuracy hides class imbalance problems
  • Iterating on the test set — Test is held-out. If test metrics are bad, gather more data and re-split
  • Skipping disagreement analysis — Random prompt tweaks without understanding root causes don't converge
  • Too few labeled examples — Below 40 total (20 pass + 20 fail), metrics are unreliable due to small sample size
  • output-eval-judge-prompt — Design the judge prompt being validated
  • output-eval-error-analysis — Source of human-labeled data for validation
  • output-eval-dataset-design — Generate additional labeled datasets if you need more data
  • output-dev-eval-testing — output workflow test CLI, --cached and --dataset flags
  • output-eval-audit — Audit whether existing judges have been validated

© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge of growthxai/output.

Open the folder on GitHubat commit 99ee298

Compare with similar skills

Output Eval Validate Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Output Eval Validate Judge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Output Eval Validate Judge this skillgrowthxai/output442—~2.5kAutomated safety check: NotesApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from growthxai/output

All 50 skills in this repo
  • Output Build Workflow

    growthxai/output

    Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.

    442 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Output Credentials Edit

    growthxai/output

    View, edit, and set encrypted credentials in an Output.ai project.

    442 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes
  • Wire encrypted credentials to environment variables using the credential: convention.

    442 GitHub stars~930 tokensUpdated yesterday
    Auto-check: notes
  • Output Credentials Init

    growthxai/output

    Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.

    442 GitHub stars~803 tokensUpdated yesterday
    Auto-check: notes
  • Output Debug Workflow

    growthxai/output

    Debug Output SDK workflow issues. An agent skill from growthxai/output.

    442 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Output Dev Agent Class

    growthxai/output

    Use the Agent class for multi-step tool loops, conversation history, streaming progress, and reusable LLM agents.

    442 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed

Questions about Output Eval Validate Judge

What does Output Eval Validate Judge do?

Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Output Eval Validate Judge is an agent skill from growthxai/output. Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.

When should I use Output Eval Validate Judge?

Output Eval Validate Judge fits situations like: tasks that involve LLM evaluation.

How do I install Output Eval Validate Judge in Claude Code?

Run `npx skills add growthxai/output --skill output-eval-validate-judge -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge in growthxai/output) into .claude/skills/output-eval-validate-judge in your project. Claude Code loads it when a task matches its description.

How do I install Output Eval Validate Judge in Codex?

Run `npx skills add growthxai/output --skill output-eval-validate-judge -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge in growthxai/output) into .agents/skills/output-eval-validate-judge in your project. Codex loads it when a task matches its description.

Can I use Output Eval Validate Judge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-eval-validate-judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-eval-validate-judge, .gemini/skills/output-eval-validate-judge, .github/skills/output-eval-validate-judge and .opencode/skills/output-eval-validate-judge in your project.

What does Output Eval Validate Judge need to run?

Going by SKILL.md and its folder, Output Eval Validate Judge needs the command-line tools its instructions call (npx). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Output Eval Validate Judge access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Output Eval Validate Judge safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Output Eval Validate Judge use?

Output Eval Validate Judge is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Output Eval Validate Judge use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Output Eval Validate Judge?

Skills that share tags, products or a category with Output Eval Validate Judge: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Output Eval Validate Judge?

growthxai (a GitHub organization) maintains it in growthxai/output, which has 442 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.