LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.
$ npx skills add growthxai/output --skill output-eval-validate-judge -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install growthxai/output output-eval-validate-judge --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge .claude/skills/output-eval-validate-judge && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "output-eval-validate-judge" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge into .claude/skills/output-eval-validate-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-validate-judge", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judgeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add growthxai/output --skill output-eval-validate-judge -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install growthxai/output output-eval-validate-judge --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .agents/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge .agents/skills/output-eval-validate-judge && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "output-eval-validate-judge" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge into .agents/skills/output-eval-validate-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-validate-judge", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-eval-validate-judge -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install growthxai/output output-eval-validate-judge --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge .cursor/skills/output-eval-validate-judge && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "output-eval-validate-judge" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge into .cursor/skills/output-eval-validate-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-validate-judge", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/growthxai/output.git --path coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add growthxai/output --skill output-eval-validate-judge -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install growthxai/output output-eval-validate-judge --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge .gemini/skills/output-eval-validate-judge && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "output-eval-validate-judge" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge into .gemini/skills/output-eval-validate-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-validate-judge", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install growthxai/output output-eval-validate-judgeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add growthxai/output --skill output-eval-validate-judge -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .github/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge .github/skills/output-eval-validate-judge && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "output-eval-validate-judge" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge into .github/skills/output-eval-validate-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-validate-judge", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-eval-validate-judge -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install growthxai/output output-eval-validate-judge --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge .opencode/skills/output-eval-validate-judge && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "output-eval-validate-judge" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge into .opencode/skills/output-eval-validate-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-validate-judge", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
output-eval-validate-judgeValidate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.
Output Eval Validate Judge is an agent skill from growthxai/output. Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Use after writing a judge prompt to verify it agrees with human judgment.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 99ee298. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteEditFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Output Eval Validate Judge loads about 2.5k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 1,129 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Write, EditAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from growthxai/output at commit 99ee298, republished under its Apache-2.0 licence (© growthxai). 1,129 words, ~2,451 tokens.
.claude/skills/output-eval-validate-judge/SKILL.md (or your agent's skills folder).An LLM judge is only useful if it agrees with human judgment. This skill walks you through calibrating a judge against human-labeled data using True Positive Rate (TPR) and True Negative Rate (TNR) metrics. Do this before trusting any judgeVerdict(), judgeScore(), or judgeLabel() evaluator in your eval suite.
.prompt file — Written following output-eval-judge-promptground_truth.evals.<evaluator_name>.verdict: pass or failThis process applies only to LLM-based judges. For code-based Verdict.* evaluators, write unit tests instead.
Split your labeled datasets into three groups:
| Split | % of Data | Purpose | Example (100 datasets) |
|---|---|---|---|
| Train | 10-20% | Source of few-shot examples in the judge prompt | 15 datasets |
| Dev | 40-45% | Iterate on judge prompt, measure TPR/TNR | 42 datasets |
| Test | 40-45% | Final held-out measurement, run once | 43 datasets |
Use a naming convention or subdirectories to separate splits:
Option A: Name prefixes
tests/datasets/
├── train_formal_pass_01.yml
├── train_casual_fail_01.yml
├── dev_technical_pass_01.yml
├── dev_ambiguous_fail_01.yml
├── test_simple_pass_01.yml
├── test_contradictory_fail_01.yml
└── ...Option B: Subdirectories
tests/datasets/
├── train/
│ ├── formal_pass_01.yml
│ └── casual_fail_01.yml
├── dev/
│ ├── technical_pass_01.yml
│ └── ambiguous_fail_01.yml
└── test/
├── simple_pass_01.yml
└── contradictory_fail_01.yml.prompt file. Never use dev or test examples — that's data leakageExecute the eval workflow against only the dev-split datasets:
# Run with cached output on dev datasets
npx output workflow test <workflowName> --cached \
--dataset dev_technical_pass_01,dev_ambiguous_fail_01,dev_formal_pass_02,...Or if using subdirectories, list the dev dataset names:
npx output workflow test <workflowName> --cached \
--dataset $(ls tests/datasets/dev/ | sed 's/.yml//' | tr '\n' ',')Save the output. You need the judge's verdict for each dataset to compare against ground truth.
Use --json to get machine-readable results:
npx output workflow test <workflowName> --cached --dataset <dev_datasets> --jsonThe output includes per-dataset, per-evaluator verdicts that you can compare against ground_truth.evals.<evaluator_name>.verdict.
For the evaluator you're validating, build a confusion matrix from the dev results.
Using "fail" as the positive class (what you're trying to detect):
| Judge says Fail | Judge says Pass | |
|---|---|---|
| Human says Fail | True Positive (TP) | False Negative (FN) |
| Human says Pass | False Positive (FP) | True Negative (TN) |
TPR (True Positive Rate) = TP / (TP + FN)
TNR (True Negative Rate) = TN / (TN + FP)
Dev set results for check_tone evaluator (42 datasets):
| Judge: Fail | Judge: Pass | |
|---|---|---|
| Human: Fail | 18 (TP) | 3 (FN) |
| Human: Pass | 2 (FP) | 19 (TN) |
Raw accuracy = (TP + TN) / total = (18 + 19) / 42 = 88.1%
This looks fine, but masks problems. If your dataset were 90% pass (class imbalance), a judge that always says "pass" would get 90% accuracy while catching zero failures (TPR = 0%). TPR and TNR measure what actually matters: catching failures and not crying wolf.
For every case where the judge disagrees with the human label, determine the root cause.
The judge said "pass" but the human said "fail." For each:
The judge said "fail" but the human said "pass." For each:
Track each disagreement to guide prompt iteration:
| Dataset | Human | Judge | Root Cause | Fix |
|---|---|---|---|---|
| dev_technical_pass_03 | pass | fail | Judge flagged "it's" as casual but context was a direct quote | Add exception: "Contractions within direct quotes are acceptable" |
| dev_ambiguous_fail_02 | fail | pass | Judge missed subtle tone shift in paragraph 3 | Add borderline few-shot example showing mid-text tone drift |
Apply the fixes from Step 4 to the judge .prompt file. Then re-run on the dev set:
npx output workflow test <workflowName> --cached --dataset <dev_datasets>Recompute TPR and TNR. Repeat until both metrics meet the target.
| Metric | Target | Minimum Acceptable |
|---|---|---|
| TPR | > 90% | > 80% |
| TNR | > 90% | > 80% |
If you can't reach 80%/80% after 3-4 iterations:
.prompt frontmatterEach iteration:
.prompt file (not random changes)Once dev metrics meet the target, run the judge on the held-out test set exactly once:
npx output workflow test <workflowName> --cached --dataset <test_datasets> --jsonCompute TPR and TNR on the test results. Record these as the final metrics.
Document the final validation results alongside the judge prompt:
# Validation: check_tone (judge_tone@v1.prompt)
# Date: 2026-03-25
# Model: claude-haiku-4-5-20251001
## Dev Set (42 datasets)
- TPR: 90.5% (19/21)
- TNR: 95.2% (20/21)
## Test Set (43 datasets)
- TPR: 88.0% (22/25)
- TNR: 94.4% (17/18)
## Conclusion: APPROVED — both metrics above 80% minimumStore this in a VALIDATION.md file next to the judge prompt or in the evaluator's documentation.
output-eval-judge-prompt — Design the judge prompt being validatedoutput-eval-error-analysis — Source of human-labeled data for validationoutput-eval-dataset-design — Generate additional labeled datasets if you need more dataoutput-dev-eval-testing — output workflow test CLI, --cached and --dataset flagsoutput-eval-audit — Audit whether existing judges have been validated© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge of growthxai/output.
Open the folder on GitHubat commit 99ee298
Output Eval Validate Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Output Eval Validate Judge this skillgrowthxai/output | 442 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
growthxai/output
Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.
growthxai/output
View, edit, and set encrypted credentials in an Output.ai project.
growthxai/output
Wire encrypted credentials to environment variables using the credential: convention.
growthxai/output
Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.
growthxai/output
Debug Output SDK workflow issues. An agent skill from growthxai/output.
growthxai/output
Use the Agent class for multi-step tool loops, conversation history, streaming progress, and reusable LLM agents.
Categories
Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Output Eval Validate Judge is an agent skill from growthxai/output. Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.
Output Eval Validate Judge fits situations like: tasks that involve LLM evaluation.
Run `npx skills add growthxai/output --skill output-eval-validate-judge -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge in growthxai/output) into .claude/skills/output-eval-validate-judge in your project. Claude Code loads it when a task matches its description.
Run `npx skills add growthxai/output --skill output-eval-validate-judge -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-validate-judge in growthxai/output) into .agents/skills/output-eval-validate-judge in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-eval-validate-judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-eval-validate-judge, .gemini/skills/output-eval-validate-judge, .github/skills/output-eval-validate-judge and .opencode/skills/output-eval-validate-judge in your project.
Going by SKILL.md and its folder, Output Eval Validate Judge needs the command-line tools its instructions call (npx). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Output Eval Validate Judge is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Output Eval Validate Judge: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
growthxai (a GitHub organization) maintains it in growthxai/output, which has 442 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.
Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.