Agent skill

Output Eval Audit

by growthxai in growthxai/output

Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Output Eval Audit

skills CLI
$ npx skills add growthxai/output --skill output-eval-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install growthxai/output output-eval-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-audit .claude/skills/output-eval-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
output-eval-audit
GitHub stars
442
Token cost
~2.5k tokens
SKILL.md length
892 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output.

  • Works in 3 steps: Gather Artifacts → Run the Diagnostic → Compile the Report
  • Inheriting evals
  • SKILL.md covers Overview, When to Use, Step 1: Gather Artifacts and Step 2: Run the Diagnostic, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Output Eval Audit is an agent skill from growthxai/output. Audit an existing eval suite for trustworthiness. Use when inheriting evals, suspecting evals miss real failures, or after significant pipeline changes.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.

When your agent uses it

  • Inheriting evals
  • Suspecting evals miss real failures
  • After significant pipeline changes

Example prompts

  • “/output-eval-audit”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Gather Artifacts
  2. Run the Diagnostic
  3. Compile the Report

What it can do on your machine

Read from SKILL.md and the folder at commit 99ee298. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Output Eval Audit loads about 2.5k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 892 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from growthxai/output at commit 99ee298, republished under its Apache-2.0 licence (© growthxai). 892 words, ~2,463 tokens.

Download SKILL.mdSave it as .claude/skills/output-eval-audit/SKILL.md (or your agent's skills folder).
name
output-eval-audit
description
Audit an existing eval suite for trustworthiness. Use when inheriting evals, suspecting evals miss real failures, or after significant pipeline changes.
allowed-tools
Bash, Read

Auditing an Eval Suite

Overview

Audit your eval suite to determine whether it actually catches real failures. This skill provides a structured diagnostic that identifies gaps in error analysis, evaluator design, judge validation, and dataset coverage, with concrete remediation steps for each finding.

When to Use

  • Inheriting an eval suite from another team or developer
  • Suspecting that evals pass but production quality is poor
  • After switching models, rewriting prompts, or changing pipeline logic
  • Periodic health check (quarterly or after major releases)

Step 1: Gather Artifacts

Read the eval infrastructure files for the workflow being audited:

src/workflows/<workflow_name>/
├── tests/
│   ├── datasets/           # YAML dataset files
│   │   ├── *.yml
│   │   └── ...
│   └── evals/
│       ├── evaluators.ts   # Evaluator definitions
│       ├── workflow.ts      # Eval workflow definition
│       └── *.prompt         # Judge prompt files

Inventory what exists:

ArtifactFile(s)Count
Evaluatorstests/evals/evaluators.ts?
Eval workflowtests/evals/workflow.ts? entries in evals array
Judge promptstests/evals/*.prompt?
Datasetstests/datasets/*.yml?
Datasets with ground_truth? of above?
Datasets with last_output? of above?

If any of these are missing entirely, note it and skip to "Starting From Zero" at the bottom.

Step 2: Run the Diagnostic

Evaluate each of the four areas below. For each, assign a status:

  • Pass — Meets the standard
  • Warn — Partially meets the standard, improvements needed
  • Fail — Does not meet the standard, significant risk

Area 1: Error Analysis Grounding

Question: Were the evaluators derived from observed failure modes in real workflow traces?

Check:

  • Do failure categories exist (documented in a file, comments, or commit history)?
  • Does each evaluator map to a specific failure category?
  • Or are evaluators measuring generic qualities ("quality score", "overall rating")?

Pass criteria:

  • Each evaluator targets a named failure mode (e.g., "check_tone" targets tone mismatch, not "evaluate general quality")
  • Failure categories were derived from reviewing real traces (not brainstormed)

Common failures:

  • Evaluators named evaluate_quality, check_overall, rate_output — generic, not grounded in observed failures
  • Evaluators were written based on what seemed important, not what actually fails
  • No evidence of trace review before evaluator creation

Remediation: output-eval-error-analysis — Review 50+ traces and categorize actual failure modes before modifying evaluators


Area 2: Evaluator Design

Question: Are the evaluators well-designed for reliable automated evaluation?

Check each evaluator in tests/evals/evaluators.ts:

CheckWhat to look for
One failure mode per judgeEach judgeVerdict() evaluator targets exactly one criterion
Binary verdictsJudge prompts use pass/fail, not Likert scales (1-5) or multi-axis ratings
Code-based where possibleObjective checks use Verdict.* helpers, not LLM judges
Few-shot examples in judgesJudge .prompt files include pass, fail, and borderline examples
Critique before verdictJudge prompts request critique/reasoning before the verdict in structured output
Appropriate criticalityrequired for blocking failures, informational for nice-to-have checks
Correct interpret typeinterpret config matches what the evaluator returns

Pass criteria:

  • All checks above are met for every evaluator

Common failures:

  • A single judge prompt evaluates 3+ criteria simultaneously ("Rate tone, accuracy, and completeness")
  • Judge prompts have no few-shot examples
  • Deterministic checks (length, string contains, regex) use LLM judges instead of Verdict.*
  • interpret type doesn't match evaluator return type (e.g., judgeVerdict() with interpret: { type: 'boolean' })

Remediation: output-eval-judge-prompt — Redesign judge prompts following the four-component structure


Show full SKILL.md (419 more words)Show less
Area 3: Judge Validation

Question: Have LLM judges been validated against human labels?

Check for each LLM-based evaluator (those using judgeVerdict(), judgeScore(), judgeLabel()):

CheckWhat to look for
Human labels existDatasets have ground_truth.evals.<evaluator_name>.verdict populated
TPR/TNR measuredValidation results documented (file, comment, or commit)
Train/dev/test splitFew-shot examples in the judge prompt come from a designated train split, not from the same data used for measurement
Metrics meet thresholdTPR > 80% and TNR > 80% (target: > 90%)

Pass criteria:

  • Every LLM judge has documented TPR/TNR metrics above 80%
  • Train/dev/test split was used (no data leakage)

Common failures:

  • No validation at all — judges were written and immediately deployed
  • Few-shot examples in the judge prompt are the same examples used to measure metrics (data leakage)
  • "It seems to work" without quantitative measurement
  • Only raw accuracy reported (masks class imbalance)

Remediation: output-eval-validate-judge — Calibrate each judge against human labels using TPR/TNR


Area 4: Dataset Coverage

Question: Do the datasets adequately cover the failure space?

Check:

CheckWhat to look for
Dataset countMinimum 10 for simple workflows, 20+ for complex ones
DiversityDatasets vary across multiple input dimensions, not just happy paths
Failure representationAt least 30% of datasets have human_verdict: fail in ground_truth
Ground truth populatedMost datasets have ground_truth with per-evaluator labels
Real + synthetic mixIncludes production traces alongside synthetic test cases
No near-duplicatesEach dataset tests a meaningfully different scenario

Pass criteria:

  • 20+ diverse datasets with ground truth
  • Both pass and fail cases represented (not 95% passes)
  • Datasets cover different input dimensions

Common failures:

  • Only 3-5 datasets, all happy-path variations
  • 100% of datasets pass (no failure cases to validate judges against)
  • Datasets are synthetic-only with no real production traces
  • Ground truth fields are empty or missing

Remediation: output-eval-dataset-design — Design diverse datasets using dimension-based variation


Step 3: Compile the Report

Summarize findings in a structured format:

markdown
# Eval Audit: <workflow_name>
# Date: YYYY-MM-DD
# Auditor: <name>

## Summary

| Area | Status | Key Finding |
|------|--------|-------------|
| Error Analysis Grounding | Warn | Evaluators seem reasonable but no documented trace review |
| Evaluator Design | Fail | Single judge evaluates 3 criteria simultaneously |
| Judge Validation | Fail | No validation performed on any LLM judge |
| Dataset Coverage | Warn | 12 datasets but only 2 are failure cases |

## Findings

### 1. Error Analysis Grounding — WARN
Evaluators target reasonable criteria (tone, topic, length) but there is no evidence
that these were derived from observed failures. The eval suite may be missing the
workflow's actual top failure modes.

**Next step:** Run error analysis on 50+ production traces (`output-eval-error-analysis`)

### 2. Evaluator Design — FAIL
`evaluate_overall_quality` in evaluators.ts uses a single judgeVerdict() call that
assesses tone, accuracy, and completeness simultaneously. This makes failures
unactionable — when it fails, you don't know which criterion failed.

**Next step:** Split into three focused judges (`output-eval-judge-prompt`)

### 3. Judge Validation — FAIL
No TPR/TNR metrics exist for any LLM judge. The judge_quality@v1.prompt has no
few-shot examples.

**Next step:** Label 100 datasets, validate each judge (`output-eval-validate-judge`)

### 4. Dataset Coverage — WARN
12 datasets exist with cached output. Only 2 have ground_truth.human_verdict: fail.
All inputs are simple topics with no edge cases.

**Next step:** Design 20+ diverse datasets (`output-eval-dataset-design`)

## Priority Order
1. Error analysis (foundational — may change which evaluators are needed)
2. Split holistic judge into focused judges
3. Expand datasets to 30+ with balanced pass/fail
4. Validate all LLM judges

Starting From Zero

If the workflow has no eval infrastructure at all:

  1. Start with error analysis — output-eval-error-analysis. Review 50+ workflow traces.
  2. Build datasets — output-eval-dataset-design. Create 20+ diverse datasets.
  3. Implement evaluators — output-dev-eval-testing. Write verify() evaluators and evalWorkflow().
  4. Design judge prompts — output-eval-judge-prompt. For subjective criteria only.
  5. Validate judges — output-eval-validate-judge. Before trusting any LLM judge.

Do not skip error analysis. Building evaluators without understanding how the workflow fails wastes effort on the wrong things.

  • output-eval-error-analysis — Systematic trace review and failure categorization
  • output-eval-judge-prompt — Design effective LLM judge prompts
  • output-eval-dataset-design — Generate diverse test datasets
  • output-eval-validate-judge — Calibrate LLM judges against human labels
  • output-dev-eval-testing — Implementation reference for offline eval testing
  • output-dev-evaluator-function — Implementation reference for runtime evaluators

© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-eval-audit of growthxai/output.

Open the folder on GitHubat commit 99ee298

Compare with similar skills

Output Eval Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Output Eval Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Output Eval Audit this skillgrowthxai/output442—~2.5kAutomated safety check: NotesApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from growthxai/output

All 50 skills in this repo
  • Output Build Workflow

    growthxai/output

    Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.

    442 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Output Credentials Edit

    growthxai/output

    View, edit, and set encrypted credentials in an Output.ai project.

    442 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes
  • Wire encrypted credentials to environment variables using the credential: convention.

    442 GitHub stars~930 tokensUpdated yesterday
    Auto-check: notes
  • Output Credentials Init

    growthxai/output

    Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.

    442 GitHub stars~803 tokensUpdated yesterday
    Auto-check: notes
  • Output Debug Workflow

    growthxai/output

    Debug Output SDK workflow issues. An agent skill from growthxai/output.

    442 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Output Dev Agent Class

    growthxai/output

    Use the Agent class for multi-step tool loops, conversation history, streaming progress, and reusable LLM agents.

    442 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed

Questions about Output Eval Audit

What does Output Eval Audit do?

Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output. Output Eval Audit is an agent skill from growthxai/output. Audit an existing eval suite for trustworthiness.

When should I use Output Eval Audit?

Output Eval Audit fits situations like: inheriting evals; suspecting evals miss real failures; after significant pipeline changes.

How do I install Output Eval Audit in Claude Code?

Run `npx skills add growthxai/output --skill output-eval-audit -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-audit in growthxai/output) into .claude/skills/output-eval-audit in your project. Claude Code loads it when a task matches its description.

How do I install Output Eval Audit in Codex?

Run `npx skills add growthxai/output --skill output-eval-audit -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-audit in growthxai/output) into .agents/skills/output-eval-audit in your project. Codex loads it when a task matches its description.

Can I use Output Eval Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-eval-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-eval-audit, .gemini/skills/output-eval-audit, .github/skills/output-eval-audit and .opencode/skills/output-eval-audit in your project.

What does Output Eval Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Output Eval Audit is instructions for the agent only. Its frontmatter pre-approves these tools: Bash, Read.

Does Output Eval Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Output Eval Audit safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Output Eval Audit use?

Output Eval Audit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Output Eval Audit use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Output Eval Audit?

Skills that share tags, products or a category with Output Eval Audit: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Output Eval Audit?

growthxai (a GitHub organization) maintains it in growthxai/output, which has 442 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.