Agent skill

Output Eval Error Analysis

by growthxai in growthxai/output

Systematically review workflow traces to identify failure modes before building evaluators.

Apache-2.0Auto-check: notes

Install Output Eval Error Analysis

skills CLI
$ npx skills add growthxai/output --skill output-eval-error-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install growthxai/output output-eval-error-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis .claude/skills/output-eval-error-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
output-eval-error-analysis
GitHub stars
440
Token cost
~2.6k tokens
SKILL.md length
976 words
Files
1
Skills in repo
52
Repo updated
First seen
Licence
Apache-2.0

At a glance

Systematically review workflow traces to identify failure modes before building evaluators.

  • Works in 6 steps: Collect Traces → Review Traces Individually → Group Into Failure Categories → …
  • Starting an eval project
  • SKILL.md covers Overview, When to Use, Step 1: Collect Traces and Step 2: Review Traces…, plus 7 more sections
  • Calls npx

What it does

Output Eval Error Analysis is an agent skill from growthxai/output. Systematically review workflow traces to identify failure modes before building evaluators. Use when starting an eval project, after significant pipeline changes, or when production quality drops.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.

When your agent uses it

  • Starting an eval project
  • After significant pipeline changes
  • Production quality drops

Example prompts

  • “/output-eval-error-analysis”

Requirements

  • Node.js
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Collect Traces
  2. Review Traces Individually
  3. Group Into Failure Categories
  4. Label Datasets
  5. Decide What to Fix vs. Evaluate
  6. Map Categories to Evaluators

What it can do on your machine

Read from SKILL.md and the folder at commit 52b51ac. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Output Eval Error Analysis loads about 2.6k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 976 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from growthxai/output at commit 52b51ac, republished under its Apache-2.0 licence (© growthxai). 976 words, ~2,650 tokens.

Download SKILL.mdSave it as .claude/skills/output-eval-error-analysis/SKILL.md (or your agent's skills folder).
name
output-eval-error-analysis
description
Systematically review workflow traces to identify failure modes before building evaluators. Use when starting an eval project, after significant pipeline changes, or when production quality drops.
allowed-tools
Bash, Read, Write, Edit

Error Analysis for Workflow Evaluation

Overview

Review real workflow traces and categorize how your workflow fails before writing any evaluators. Evaluators built without error analysis target generic qualities ("is this good?") instead of the specific ways your workflow actually breaks. This skill walks you through the process.

When to Use

  • Starting a new eval project for an existing workflow
  • Production quality has dropped and you need to understand why
  • After significant prompt, model, or pipeline changes
  • Before building your first evaluator for a workflow

Step 1: Collect Traces

Gather 50-100 representative workflow executions. More traces = more reliable failure categories.

From recent runs

List recent workflow executions and pull their traces:

bash
# List recent runs for a workflow
npx output workflow runs list <workflowName>

# Pull a specific trace as JSON
npx output workflow debug <workflowId> --json
From production (bulk download)

Download production traces directly into dataset YAML files:

bash
# Download up to 20 recent traces as dataset files
npx output workflow dataset generate <workflowName> --download --limit 20

This creates YAML files in tests/datasets/ with the input and last_output fields populated from real executions.

From scenario-driven generation

If production traces are sparse, generate traces from scenario inputs:

bash
# Generate a dataset from a scenario file
npx output workflow dataset generate <workflowName> basic --name basic_trace

# Generate from inline JSON
npx output workflow dataset generate <workflowName> --input '{"topic": "AI safety"}' --name ai_safety_trace

Run enough inputs to get 50+ traces. Prioritize diversity over volume — vary inputs across the dimensions you expect to matter.

Step 2: Review Traces Individually

Review each trace one at a time. For each trace, record:

FieldWhat to write
Trace IDThe workflow execution ID
VerdictPass or Fail (binary — no "partial" at this stage)
Root causeIf Fail: what specifically went wrong and why
NotesAnything surprising or worth remembering
Review template

Create a file to track your reviews. A simple markdown table works:

markdown
# Error Analysis: <workflow_name>
# Date: YYYY-MM-DD
# Traces reviewed: 0 / 50

| # | Trace ID | Verdict | Root Cause | Notes |
|---|----------|---------|------------|-------|
| 1 | abc-123  | Fail    | Hallucinated a URL that doesn't exist | Common with technical topics |
| 2 | def-456  | Pass    | — | Clean output |
| 3 | ghi-789  | Fail    | Ignored the "formal tone" requirement | Input had conflicting signals |
What to look for in each trace

Open the JSON trace and examine:

  1. Final output — Does it meet the user's intent? Is it correct?
  2. Step-by-step data flow — Did each step receive the right input and produce reasonable output?
  3. LLM responses — Did the model follow instructions? Did it hallucinate?
  4. Error states — Did any step fail, retry, or produce unexpected errors?
Critical rule: read first, categorize second

Review at least 30 traces before naming any failure categories. Premature categorization causes you to see patterns that aren't there and miss patterns that are. Just record what you observe.

Step 3: Group Into Failure Categories

After reviewing 30+ traces, patterns will emerge. Group your failures into 5-10 categories based on root cause, not surface symptoms.

Good categories (root cause)
  • "Hallucinated URLs" — model invents links that don't exist
  • "Tone mismatch" — output tone doesn't match the requested persona
  • "Missing required section" — output omits a section the input explicitly requested
  • "Factual error" — output contains verifiably wrong claims
  • "Prompt injection leak" — user input manipulates the system prompt
Bad categories (surface symptoms)
  • "Bad output" — too vague, not actionable
  • "LLM error" — doesn't identify the specific failure
  • "Quality issue" — could mean anything
Splitting and merging
  • If a category has fewer than 3 examples, merge it into a broader category or note it as rare
  • If a category has 15+ examples and contains distinct sub-patterns, split it
  • Categories should be mutually exclusive — each failure belongs to exactly one category
Example categorization

For a blog generation workflow after reviewing 60 traces:

CategoryCountRateExample
Hallucinated URLs813%Invented links to non-existent pages
Tone mismatch610%Casual tone when formal was requested
Off-topic drift58%Blog about "AI" drifted to unrelated ML history
Missing sections47%Skipped "conclusion" when explicitly requested
Too short35%Under 200 words when 500+ requested
Total failures2643%
Passes3457%

Step 4: Label Datasets

Add ground_truth labels to your dataset YAML files so evaluators can validate against them. Each failure category maps to a future evaluator name.

Show full SKILL.md (399 more words)Show less
YAML structure
yaml
name: ai_safety_trace
input:
  topic: "AI safety"
  tone: "formal"
  min_length: 500
last_output:
  output:
    title: "Understanding AI Safety"
    blog_post: "AI safety is super important and stuff..."
  executionTimeMs: 3200
  date: '2026-03-25T00:00:00.000Z'
ground_truth:
  # Global ground truth (available to all evaluators)
  human_verdict: fail
  failure_categories:
    - tone_mismatch
  notes: "Used casual language despite formal tone request"
  # Per-evaluator ground truth
  evals:
    check_tone:
      expected_tone: formal
      verdict: fail
    check_length:
      min_length: 500
      verdict: pass
    check_hallucinated_urls:
      verdict: pass

The ground_truth.evals.<evaluator_name> fields map directly to the evaluator names you'll use in verify(). Each evaluator receives its own ground truth merged with the top-level ground truth via context.ground_truth.

Labeling efficiently

You don't need to label every dataset for every category. Focus on:

  1. Label all datasets with the global human_verdict (pass/fail)
  2. Label datasets for the top 3 failure categories by rate
  3. Add per-evaluator labels as you build each evaluator

Step 5: Decide What to Fix vs. Evaluate

Not every failure category needs an evaluator. Use this decision tree:

Is this failure caused by a fixable prompt/tool gap?
├─ YES → Fix the prompt or add the missing tool first
│        Re-run error analysis after the fix
└─ NO  → Will this failure recur and need ongoing monitoring?
         ├─ YES → Build an evaluator
         │        Can it be checked with deterministic code?
         │        ├─ YES → Use Verdict.* helpers (contains, matches, gte, etc.)
         │        └─ NO  → Use judgeVerdict() with an LLM judge prompt
         └─ NO  → Document it and move on (rare edge case)
Prioritize by failure rate

Build evaluators for the highest-rate failure categories first. A failure at 13% matters more than one at 2%.

Code-based checks first

Many failures that seem subjective have objective proxies:

FailureSeems like...But you can check with...
"Too short"SubjectiveVerdict.gte(output.length, threshold)
"Missing section"Needs LLMVerdict.contains(output, "## Conclusion")
"Hallucinated URLs"Needs LLMExtract URLs with regex, verify with HTTP HEAD
"Wrong format"Needs LLMVerdict.matches(output, expectedPattern)

Reserve LLM judges for genuinely subjective criteria: tone, relevance, faithfulness, coherence.

Step 6: Map Categories to Evaluators

Create a mapping document that connects your failure categories to planned evaluators:

markdown
# Evaluator Plan: blog_generator

| Category | Rate | Evaluator Type | Evaluator Name | Criticality |
|----------|------|----------------|----------------|-------------|
| Hallucinated URLs | 13% | Code (URL extraction + HTTP check) | check_urls | required |
| Tone mismatch | 10% | LLM judge | check_tone | required |
| Off-topic drift | 8% | LLM judge | check_topic | required |
| Missing sections | 7% | Code (string contains) | check_sections | required |
| Too short | 5% | Code (length check) | check_length | informational |

This becomes your implementation roadmap. Use criticality: 'required' for failure categories that should block a passing verdict. Use 'informational' for nice-to-have checks.

Next Steps

  • Build evaluators — Follow output-dev-eval-testing to implement each evaluator with verify() and wire them into evalWorkflow()
  • Design judge prompts — For LLM-based evaluators, follow output-eval-judge-prompt to write effective .prompt files
  • Expand datasets — If your traces don't cover enough failure regions, follow output-eval-dataset-design to generate diverse test cases
  • Re-run after changes — After fixing prompts, switching models, or modifying pipeline logic, repeat this error analysis to find new failure modes

Anti-Patterns

  • Building evaluators without error analysis — You'll evaluate the wrong things
  • Categorizing before reviewing 30+ traces — Premature categories cause confirmation bias
  • Surface-level categories ("bad output", "LLM error") — Split by root cause
  • One giant evaluator — One evaluator per failure mode, not one evaluator for everything
  • Skipping code-based checks — Don't use an LLM judge when Verdict.contains() works
  • Never re-running — Error analysis is not a one-time activity; repeat after significant changes
  • output-dev-eval-testing — Implement evaluators with verify(), Verdict, and evalWorkflow()
  • output-eval-judge-prompt — Design LLM judge prompts for subjective failure modes
  • output-eval-dataset-design — Generate diverse datasets when real traces are sparse
  • output-eval-validate-judge — Validate LLM judges against human labels
  • output-eval-audit — Audit an existing eval suite for trustworthiness
  • output-workflow-trace — Retrieve and analyze workflow execution traces

© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis of growthxai/output.

Open the folder on GitHubat commit 52b51ac

Compare with similar skills

Output Eval Error Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Output Eval Error Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Output Eval Error Analysis this skillgrowthxai/output440—~2.6kAutomated safety check: NotesApache-2.0
Triage Agent Eval Failuresnovuhq/novu40k—~1.5kAutomated safety check: PassCustom licence
Eval Harnessaffaan-m/ECC275k—~2.2kAutomated safety check: PassMIT
Resilience Hub Failure Mode Assessmentaws/agent-toolkit-for-aws2.8k—~1kAutomated safety check: PassApache-2.0
Dynamic Workflow Modeaffaan-m/ECC275k1 repos~1.3kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT

Similar skills

  • Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge).

    40k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    275k GitHub stars~2.2k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs and interprets AWS Resilience Hub v2 failure mode assessments.

    2.8k GitHub stars~1k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Design task-local harnesses, eval gates, and reusable skill extraction for Claude dynamic workflow mode and other adaptive agent harnesses.

    275k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    275k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from growthxai/output

All 52 skills in this repo
  • Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object().

    440 GitHub stars~597 tokensUpdated yesterday
    Auto-check passed
  • Output Build Workflow

    growthxai/output

    Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.

    440 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Output Credentials Edit

    growthxai/output

    View, edit, and set encrypted credentials in an Output.ai project.

    440 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check: notes
  • Wire encrypted credentials to environment variables using the credential: convention.

    440 GitHub stars~930 tokensUpdated yesterday
    Auto-check: notes
  • Output Credentials Init

    growthxai/output

    Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.

    440 GitHub stars~803 tokensUpdated yesterday
    Auto-check: notes
  • Output Debug Workflow

    growthxai/output

    Debug Output SDK workflow issues. An agent skill from growthxai/output.

    440 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Questions about Output Eval Error Analysis

What does Output Eval Error Analysis do?

Systematically review workflow traces to identify failure modes before building evaluators. Output Eval Error Analysis is an agent skill from growthxai/output. Systematically review workflow traces to identify failure modes before building evaluators.

When should I use Output Eval Error Analysis?

Output Eval Error Analysis fits situations like: starting an eval project; after significant pipeline changes; production quality drops.

How do I install Output Eval Error Analysis in Claude Code?

Run `npx skills add growthxai/output --skill output-eval-error-analysis -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis in growthxai/output) into .claude/skills/output-eval-error-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Output Eval Error Analysis in Codex?

Run `npx skills add growthxai/output --skill output-eval-error-analysis -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis in growthxai/output) into .agents/skills/output-eval-error-analysis in your project. Codex loads it when a task matches its description.

Can I use Output Eval Error Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-eval-error-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-eval-error-analysis, .gemini/skills/output-eval-error-analysis, .github/skills/output-eval-error-analysis and .opencode/skills/output-eval-error-analysis in your project.

What does Output Eval Error Analysis need to run?

Going by SKILL.md and its folder, Output Eval Error Analysis needs the command-line tools its instructions call (npx). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Output Eval Error Analysis access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Output Eval Error Analysis safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Output Eval Error Analysis use?

Output Eval Error Analysis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Output Eval Error Analysis use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Output Eval Error Analysis?

Skills that share tags, products or a category with Output Eval Error Analysis: Triage Agent Eval Failures (novuhq/novu, 40k stars), Eval Harness (affaan-m/ECC, 275k stars), Resilience Hub Failure Mode Assessment (aws/agent-toolkit-for-aws, 2.8k stars) and Dynamic Workflow Mode (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Output Eval Error Analysis?

growthxai (a GitHub organization) maintains it in growthxai/output, which has 440 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on October 7, 2026.

Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.