Triage Agent Eval Failures
novuhq/novu
Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge).
Systematically review workflow traces to identify failure modes before building evaluators.
$ npx skills add growthxai/output --skill output-eval-error-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install growthxai/output output-eval-error-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .claude/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis .claude/skills/output-eval-error-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "output-eval-error-analysis" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis into .claude/skills/output-eval-error-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-error-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add growthxai/output --skill output-eval-error-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install growthxai/output output-eval-error-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .agents/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis .agents/skills/output-eval-error-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "output-eval-error-analysis" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis into .agents/skills/output-eval-error-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-error-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-eval-error-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install growthxai/output output-eval-error-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis .cursor/skills/output-eval-error-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "output-eval-error-analysis" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis into .cursor/skills/output-eval-error-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-error-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/growthxai/output.git --path coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add growthxai/output --skill output-eval-error-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install growthxai/output output-eval-error-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis .gemini/skills/output-eval-error-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "output-eval-error-analysis" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis into .gemini/skills/output-eval-error-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-error-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install growthxai/output output-eval-error-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add growthxai/output --skill output-eval-error-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .github/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis .github/skills/output-eval-error-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "output-eval-error-analysis" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis into .github/skills/output-eval-error-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-error-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add growthxai/output --skill output-eval-error-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install growthxai/output output-eval-error-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/growthxai/output.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis .opencode/skills/output-eval-error-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "output-eval-error-analysis" agent skill from https://github.com/growthxai/output/tree/main/coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis into .opencode/skills/output-eval-error-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "output-eval-error-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
output-eval-error-analysisSystematically review workflow traces to identify failure modes before building evaluators.
Output Eval Error Analysis is an agent skill from growthxai/output. Systematically review workflow traces to identify failure modes before building evaluators. Use when starting an eval project, after significant pipeline changes, or when production quality drops.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code describe what you want, Claude builds it, with all the best practices already… The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 52b51ac. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteEditFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Output Eval Error Analysis loads about 2.6k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 976 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Write, EditAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from growthxai/output at commit 52b51ac, republished under its Apache-2.0 licence (© growthxai). 976 words, ~2,650 tokens.
.claude/skills/output-eval-error-analysis/SKILL.md (or your agent's skills folder).Review real workflow traces and categorize how your workflow fails before writing any evaluators. Evaluators built without error analysis target generic qualities ("is this good?") instead of the specific ways your workflow actually breaks. This skill walks you through the process.
Gather 50-100 representative workflow executions. More traces = more reliable failure categories.
List recent workflow executions and pull their traces:
# List recent runs for a workflow
npx output workflow runs list <workflowName>
# Pull a specific trace as JSON
npx output workflow debug <workflowId> --jsonDownload production traces directly into dataset YAML files:
# Download up to 20 recent traces as dataset files
npx output workflow dataset generate <workflowName> --download --limit 20This creates YAML files in tests/datasets/ with the input and last_output fields populated from real executions.
If production traces are sparse, generate traces from scenario inputs:
# Generate a dataset from a scenario file
npx output workflow dataset generate <workflowName> basic --name basic_trace
# Generate from inline JSON
npx output workflow dataset generate <workflowName> --input '{"topic": "AI safety"}' --name ai_safety_traceRun enough inputs to get 50+ traces. Prioritize diversity over volume — vary inputs across the dimensions you expect to matter.
Review each trace one at a time. For each trace, record:
| Field | What to write |
|---|---|
| Trace ID | The workflow execution ID |
| Verdict | Pass or Fail (binary — no "partial" at this stage) |
| Root cause | If Fail: what specifically went wrong and why |
| Notes | Anything surprising or worth remembering |
Create a file to track your reviews. A simple markdown table works:
# Error Analysis: <workflow_name>
# Date: YYYY-MM-DD
# Traces reviewed: 0 / 50
| # | Trace ID | Verdict | Root Cause | Notes |
|---|----------|---------|------------|-------|
| 1 | abc-123 | Fail | Hallucinated a URL that doesn't exist | Common with technical topics |
| 2 | def-456 | Pass | — | Clean output |
| 3 | ghi-789 | Fail | Ignored the "formal tone" requirement | Input had conflicting signals |Open the JSON trace and examine:
Review at least 30 traces before naming any failure categories. Premature categorization causes you to see patterns that aren't there and miss patterns that are. Just record what you observe.
After reviewing 30+ traces, patterns will emerge. Group your failures into 5-10 categories based on root cause, not surface symptoms.
For a blog generation workflow after reviewing 60 traces:
| Category | Count | Rate | Example |
|---|---|---|---|
| Hallucinated URLs | 8 | 13% | Invented links to non-existent pages |
| Tone mismatch | 6 | 10% | Casual tone when formal was requested |
| Off-topic drift | 5 | 8% | Blog about "AI" drifted to unrelated ML history |
| Missing sections | 4 | 7% | Skipped "conclusion" when explicitly requested |
| Too short | 3 | 5% | Under 200 words when 500+ requested |
| Total failures | 26 | 43% | |
| Passes | 34 | 57% |
Add ground_truth labels to your dataset YAML files so evaluators can validate against them. Each failure category maps to a future evaluator name.
name: ai_safety_trace
input:
topic: "AI safety"
tone: "formal"
min_length: 500
last_output:
output:
title: "Understanding AI Safety"
blog_post: "AI safety is super important and stuff..."
executionTimeMs: 3200
date: '2026-03-25T00:00:00.000Z'
ground_truth:
# Global ground truth (available to all evaluators)
human_verdict: fail
failure_categories:
- tone_mismatch
notes: "Used casual language despite formal tone request"
# Per-evaluator ground truth
evals:
check_tone:
expected_tone: formal
verdict: fail
check_length:
min_length: 500
verdict: pass
check_hallucinated_urls:
verdict: passThe ground_truth.evals.<evaluator_name> fields map directly to the evaluator names you'll use in verify(). Each evaluator receives its own ground truth merged with the top-level ground truth via context.ground_truth.
You don't need to label every dataset for every category. Focus on:
human_verdict (pass/fail)Not every failure category needs an evaluator. Use this decision tree:
Is this failure caused by a fixable prompt/tool gap?
├─ YES → Fix the prompt or add the missing tool first
│ Re-run error analysis after the fix
└─ NO → Will this failure recur and need ongoing monitoring?
├─ YES → Build an evaluator
│ Can it be checked with deterministic code?
│ ├─ YES → Use Verdict.* helpers (contains, matches, gte, etc.)
│ └─ NO → Use judgeVerdict() with an LLM judge prompt
└─ NO → Document it and move on (rare edge case)Build evaluators for the highest-rate failure categories first. A failure at 13% matters more than one at 2%.
Many failures that seem subjective have objective proxies:
| Failure | Seems like... | But you can check with... |
|---|---|---|
| "Too short" | Subjective | Verdict.gte(output.length, threshold) |
| "Missing section" | Needs LLM | Verdict.contains(output, "## Conclusion") |
| "Hallucinated URLs" | Needs LLM | Extract URLs with regex, verify with HTTP HEAD |
| "Wrong format" | Needs LLM | Verdict.matches(output, expectedPattern) |
Reserve LLM judges for genuinely subjective criteria: tone, relevance, faithfulness, coherence.
Create a mapping document that connects your failure categories to planned evaluators:
# Evaluator Plan: blog_generator
| Category | Rate | Evaluator Type | Evaluator Name | Criticality |
|----------|------|----------------|----------------|-------------|
| Hallucinated URLs | 13% | Code (URL extraction + HTTP check) | check_urls | required |
| Tone mismatch | 10% | LLM judge | check_tone | required |
| Off-topic drift | 8% | LLM judge | check_topic | required |
| Missing sections | 7% | Code (string contains) | check_sections | required |
| Too short | 5% | Code (length check) | check_length | informational |This becomes your implementation roadmap. Use criticality: 'required' for failure categories that should block a passing verdict. Use 'informational' for nice-to-have checks.
output-dev-eval-testing to implement each evaluator with verify() and wire them into evalWorkflow()output-eval-judge-prompt to write effective .prompt filesoutput-eval-dataset-design to generate diverse test casesVerdict.contains() worksoutput-dev-eval-testing — Implement evaluators with verify(), Verdict, and evalWorkflow()output-eval-judge-prompt — Design LLM judge prompts for subjective failure modesoutput-eval-dataset-design — Generate diverse datasets when real traces are sparseoutput-eval-validate-judge — Validate LLM judges against human labelsoutput-eval-audit — Audit an existing eval suite for trustworthinessoutput-workflow-trace — Retrieve and analyze workflow execution traces© growthxai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis of growthxai/output.
Open the folder on GitHubat commit 52b51ac
Output Eval Error Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Output Eval Error Analysis this skillgrowthxai/output | 440 | — | ~2.6k | Automated safety check: Notes | Apache-2.0 | |
| Triage Agent Eval Failuresnovuhq/novu | 40k | — | ~1.5k | Automated safety check: Pass | Custom licence | |
| Eval Harnessaffaan-m/ECC | 275k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Resilience Hub Failure Mode Assessmentaws/agent-toolkit-for-aws | 2.8k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Dynamic Workflow Modeaffaan-m/ECC | 275k | 1 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Evalalirezarezvani/claude-skills | 28k | 1 repos | ~618 | Automated safety check: Pass | MIT |
novuhq/novu
Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge).
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
aws/agent-toolkit-for-aws
Runs and interprets AWS Resilience Hub v2 failure mode assessments.
affaan-m/ECC
Design task-local harnesses, eval gates, and reusable skill extraction for Claude dynamic workflow mode and other adaptive agent harnesses.
alirezarezvani/claude-skills
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
affaan-m/ECC
Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi
growthxai/output
Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object().
growthxai/output
Implement an Output SDK workflow from a plan document. An agent skill from growthxai/output.
growthxai/output
View, edit, and set encrypted credentials in an Output.ai project.
growthxai/output
Wire encrypted credentials to environment variables using the credential: convention.
growthxai/output
Initialize encrypted credentials for an Output.ai project. An agent skill from growthxai/output.
growthxai/output
Debug Output SDK workflow issues. An agent skill from growthxai/output.
Systematically review workflow traces to identify failure modes before building evaluators. Output Eval Error Analysis is an agent skill from growthxai/output. Systematically review workflow traces to identify failure modes before building evaluators.
Output Eval Error Analysis fits situations like: starting an eval project; after significant pipeline changes; production quality drops.
Run `npx skills add growthxai/output --skill output-eval-error-analysis -a claude-code`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis in growthxai/output) into .claude/skills/output-eval-error-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add growthxai/output --skill output-eval-error-analysis -a codex`. Or copy the skill folder (coding_assistants/claude/plugins/outputai/skills/output-eval-error-analysis in growthxai/output) into .agents/skills/output-eval-error-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthxai/output --skill output-eval-error-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-eval-error-analysis, .gemini/skills/output-eval-error-analysis, .github/skills/output-eval-error-analysis and .opencode/skills/output-eval-error-analysis in your project.
Going by SKILL.md and its folder, Output Eval Error Analysis needs the command-line tools its instructions call (npx). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Output Eval Error Analysis is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Output Eval Error Analysis: Triage Agent Eval Failures (novuhq/novu, 40k stars), Eval Harness (affaan-m/ECC, 275k stars), Resilience Hub Failure Mode Assessment (aws/agent-toolkit-for-aws, 2.8k stars) and Dynamic Workflow Mode (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
growthxai (a GitHub organization) maintains it in growthxai/output, which has 440 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on October 7, 2026.
Source: growthxai/output on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.