LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output.
$ npx skills add growthxai/output --skill output-eval-audit -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install growthxai/output output-eval-audit --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-audit .claude/skills/output-eval-audit && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "output-eval-audit" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-audit into .claude/skills/output-eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-audit", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-auditType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add growthxai/output --skill output-eval-audit -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install growthxai/output output-eval-audit --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .agents/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-audit .agents/skills/output-eval-audit && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "output-eval-audit" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-audit into .agents/skills/output-eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-audit", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-eval-audit -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install growthxai/output output-eval-audit --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-audit .cursor/skills/output-eval-audit && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "output-eval-audit" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-audit into .cursor/skills/output-eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-audit", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/growthxai/output.git --path coding_assistants/claude/plugins/outputai/skills/output-eval-audit--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add growthxai/output --skill output-eval-audit -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install growthxai/output output-eval-audit --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-audit .gemini/skills/output-eval-audit && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "output-eval-audit" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-audit into .gemini/skills/output-eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-audit", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install growthxai/output output-eval-auditInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add growthxai/output --skill output-eval-audit -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .github/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-audit .github/skills/output-eval-audit && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "output-eval-audit" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-audit into .github/skills/output-eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-audit", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-eval-audit -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install growthxai/output output-eval-audit --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-audit .opencode/skills/output-eval-audit && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "output-eval-audit" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-audit into .opencode/skills/output-eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-audit", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
output-eval-auditAudit an existing eval suite for trustworthiness. An agent skill from growthxai/output.
Output Eval Audit is an agent skill from growthxai/output. Audit an existing eval suite for trustworthiness. Use when inheriting evals, suspecting evals miss real failures, or after significant pipeline changes.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 99ee298. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Output Eval Audit loads about 2.5k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 892 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, ReadAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from growthxai/output at commit 99ee298, republished under its Apache-2.0 licence (© growthxai). 892 words, ~2,463 tokens.
.claude/skills/output-eval-audit/SKILL.md (or your agent's skills folder).Audit your eval suite to determine whether it actually catches real failures. This skill provides a structured diagnostic that identifies gaps in error analysis, evaluator design, judge validation, and dataset coverage, with concrete remediation steps for each finding.
Read the eval infrastructure files for the workflow being audited:
src/workflows/<workflow_name>/
├── tests/
│ ├── datasets/ # YAML dataset files
│ │ ├── *.yml
│ │ └── ...
│ └── evals/
│ ├── evaluators.ts # Evaluator definitions
│ ├── workflow.ts # Eval workflow definition
│ └── *.prompt # Judge prompt filesInventory what exists:
| Artifact | File(s) | Count |
|---|---|---|
| Evaluators | tests/evals/evaluators.ts | ? |
| Eval workflow | tests/evals/workflow.ts | ? entries in evals array |
| Judge prompts | tests/evals/*.prompt | ? |
| Datasets | tests/datasets/*.yml | ? |
| Datasets with ground_truth | ? of above | ? |
| Datasets with last_output | ? of above | ? |
If any of these are missing entirely, note it and skip to "Starting From Zero" at the bottom.
Evaluate each of the four areas below. For each, assign a status:
Question: Were the evaluators derived from observed failure modes in real workflow traces?
Check:
Pass criteria:
Common failures:
evaluate_quality, check_overall, rate_output — generic, not grounded in observed failuresRemediation: output-eval-error-analysis — Review 50+ traces and categorize actual failure modes before modifying evaluators
Question: Are the evaluators well-designed for reliable automated evaluation?
Check each evaluator in tests/evals/evaluators.ts:
| Check | What to look for |
|---|---|
| One failure mode per judge | Each judgeVerdict() evaluator targets exactly one criterion |
| Binary verdicts | Judge prompts use pass/fail, not Likert scales (1-5) or multi-axis ratings |
| Code-based where possible | Objective checks use Verdict.* helpers, not LLM judges |
| Few-shot examples in judges | Judge .prompt files include pass, fail, and borderline examples |
| Critique before verdict | Judge prompts request critique/reasoning before the verdict in structured output |
| Appropriate criticality | required for blocking failures, informational for nice-to-have checks |
| Correct interpret type | interpret config matches what the evaluator returns |
Pass criteria:
Common failures:
Verdict.*interpret type doesn't match evaluator return type (e.g., judgeVerdict() with interpret: { type: 'boolean' })Remediation: output-eval-judge-prompt — Redesign judge prompts following the four-component structure
Question: Have LLM judges been validated against human labels?
Check for each LLM-based evaluator (those using judgeVerdict(), judgeScore(), judgeLabel()):
| Check | What to look for |
|---|---|
| Human labels exist | Datasets have ground_truth.evals.<evaluator_name>.verdict populated |
| TPR/TNR measured | Validation results documented (file, comment, or commit) |
| Train/dev/test split | Few-shot examples in the judge prompt come from a designated train split, not from the same data used for measurement |
| Metrics meet threshold | TPR > 80% and TNR > 80% (target: > 90%) |
Pass criteria:
Common failures:
Remediation: output-eval-validate-judge — Calibrate each judge against human labels using TPR/TNR
Question: Do the datasets adequately cover the failure space?
Check:
| Check | What to look for |
|---|---|
| Dataset count | Minimum 10 for simple workflows, 20+ for complex ones |
| Diversity | Datasets vary across multiple input dimensions, not just happy paths |
| Failure representation | At least 30% of datasets have human_verdict: fail in ground_truth |
| Ground truth populated | Most datasets have ground_truth with per-evaluator labels |
| Real + synthetic mix | Includes production traces alongside synthetic test cases |
| No near-duplicates | Each dataset tests a meaningfully different scenario |
Pass criteria:
Common failures:
Remediation: output-eval-dataset-design — Design diverse datasets using dimension-based variation
Summarize findings in a structured format:
# Eval Audit: <workflow_name>
# Date: YYYY-MM-DD
# Auditor: <name>
## Summary
| Area | Status | Key Finding |
|------|--------|-------------|
| Error Analysis Grounding | Warn | Evaluators seem reasonable but no documented trace review |
| Evaluator Design | Fail | Single judge evaluates 3 criteria simultaneously |
| Judge Validation | Fail | No validation performed on any LLM judge |
| Dataset Coverage | Warn | 12 datasets but only 2 are failure cases |
## Findings
### 1. Error Analysis Grounding — WARN
Evaluators target reasonable criteria (tone, topic, length) but there is no evidence
that these were derived from observed failures. The eval suite may be missing the
workflow's actual top failure modes.
**Next step:** Run error analysis on 50+ production traces (`output-eval-error-analysis`)
### 2. Evaluator Design — FAIL
`evaluate_overall_quality` in evaluators.ts uses a single judgeVerdict() call that
assesses tone, accuracy, and completeness simultaneously. This makes failures
unactionable — when it fails, you don't know which criterion failed.
**Next step:** Split into three focused judges (`output-eval-judge-prompt`)
### 3. Judge Validation — FAIL
No TPR/TNR metrics exist for any LLM judge. The judge_quality@v1.prompt has no
few-shot examples.
**Next step:** Label 100 datasets, validate each judge (`output-eval-validate-judge`)
### 4. Dataset Coverage — WARN
12 datasets exist with cached output. Only 2 have ground_truth.human_verdict: fail.
All inputs are simple topics with no edge cases.
**Next step:** Design 20+ diverse datasets (`output-eval-dataset-design`)
## Priority Order
1. Error analysis (foundational — may change which evaluators are needed)
2. Split holistic judge into focused judges
3. Expand datasets to 30+ with balanced pass/fail
4. Validate all LLM judgesIf the workflow has no eval infrastructure at all:
output-eval-error-analysis. Review 50+ workflow traces.output-eval-dataset-design. Create 20+ diverse datasets.output-dev-eval-testing. Write verify() evaluators and evalWorkflow().output-eval-judge-prompt. For subjective criteria only.output-eval-validate-judge. Before trusting any LLM judge.Do not skip error analysis. Building evaluators without understanding how the workflow fails wastes effort on the wrong things.
output-eval-error-analysis — Systematic trace review and failure categorizationoutput-eval-judge-prompt — Design effective LLM judge promptsoutput-eval-dataset-design — Generate diverse test datasetsoutput-eval-validate-judge — Calibrate LLM judges against human labelsoutput-dev-eval-testing — Implementation reference for offline eval testingoutput-dev-evaluator-function — Implementation reference for runtime evaluators© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-eval-audit of growthxai/output.
Open the folder on GitHubat commit 99ee298
Output Eval Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Output Eval Audit this skillgrowthxai/output | 442 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
growthxai/output
Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.
growthxai/output
View, edit, and set encrypted credentials in an Output.ai project.
growthxai/output
Wire encrypted credentials to environment variables using the credential: convention.
growthxai/output
Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.
growthxai/output
Debug Output SDK workflow issues. An agent skill from growthxai/output.
growthxai/output
Use the Agent class for multi-step tool loops, conversation history, streaming progress, and reusable LLM agents.
Categories
Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output. Output Eval Audit is an agent skill from growthxai/output. Audit an existing eval suite for trustworthiness.
Output Eval Audit fits situations like: inheriting evals; suspecting evals miss real failures; after significant pipeline changes.
Run `npx skills add growthxai/output --skill output-eval-audit -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-audit in growthxai/output) into .claude/skills/output-eval-audit in your project. Claude Code loads it when a task matches its description.
Run `npx skills add growthxai/output --skill output-eval-audit -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-audit in growthxai/output) into .agents/skills/output-eval-audit in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-eval-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-eval-audit, .gemini/skills/output-eval-audit, .github/skills/output-eval-audit and .opencode/skills/output-eval-audit in your project.
SKILL.md names no scripts, command-line tools or credentials: Output Eval Audit is instructions for the agent only. Its frontmatter pre-approves these tools: Bash, Read.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Output Eval Audit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Output Eval Audit: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
growthxai (a GitHub organization) maintains it in growthxai/output, which has 442 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.
Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.