Agent skill

Diagnose Failed Run

by adrianco in adrianco/retort

Determine the TRUE cause of a failed retort run before attributing it.

Apache-2.0Auto-check passedTesting & QA

Install Diagnose Failed Run

skills CLI
$ npx skills add adrianco/retort --skill diagnose-failed-run -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install adrianco/retort diagnose-failed-run --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/diagnose-failed-run .claude/skills/diagnose-failed-run && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
diagnose-failed-run
GitHub stars
207
Token cost
~1.7k tokens
SKILL.md length
864 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Determine the TRUE cause of a failed retort run before attributing it.

  • Works in 4 steps: Ground-truth before concluding. Run the… → **Label every claim [DIRECT] (observed:… → Read the DB and the agent logs. The… → …
  • A run is recorded failed (testcoverage=0 / gate fail)
  • SKILL.md covers Overview, Cardinal rules, Inputs in each runs//repN/ and Procedure, plus 5 more sections
  • Calls cargo, go and mvn

What it does

Diagnose Failed Run is an agent skill from adrianco/retort. Determine the TRUE cause of a failed retort run before attributing it. Ground-truth every failure (run its tests, read its agent logs, inspect its workspace) and classify it as an infrastructure false-fail, a genuine model miss, or an environment issue — never trust the gate verdict or a log signature alone. Use whenever a run is recorded failed (testcoverage=0 / gate fail), a pass rate looks low, or you're deciding whether a "failure" is real before reporting or proceeding.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test coverage. The repository describes itself as: Platform Evolution Engine. Distill the best from the combinatorial mess. The licence is Apache-2.0.

When your agent uses it

  • A run is recorded failed (testcoverage=0 / gate fail)
  • A pass rate looks low
  • Youre deciding whether a failure is real before reporting

Example prompts

  • “re deciding whether a”
  • “/diagnose-failed-run”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Ground-truth before concluding. Run the tests yourself, inspect the workspace, read
  2. **Label every claim [DIRECT] (observed: command output, file state, a reproduced
  3. Read the DB and the agent logs. The recorded scores/session show which tool or
  4. Verify any fix end-to-end (through the harness), not just CLI/unit. A fix that

What it can do on your machine

Read from SKILL.md and the folder at commit 1f75769. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo
    • go
    • mvn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Diagnose Failed Run loads about 1.7k tokens when it runs. Until then it costs about 125 tokens; SKILL.md has 864 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from adrianco/retort at commit 1f75769, republished under its Apache-2.0 licence (© adrianco). 864 words, ~1,744 tokens.

Download SKILL.mdSave it as .claude/skills/diagnose-failed-run/SKILL.md (or your agent's skills folder).
name
diagnose-failed-run
description
Determine the TRUE cause of a failed retort run before attributing it. Ground-truth every failure (run its tests, read its agent logs, inspect its workspace) and classify it as an infrastructure false-fail, a genuine model miss, or an environment issue — never trust the gate verdict or a log signature alone. Use whenever a run is recorded failed (test_coverage=0 / gate fail), a pass rate looks low, or you're deciding whether a "failure" is real before reporting or proceeding.
type
anthropic-skill
version
1.0

Diagnose a Failed Retort Run

Overview

A run failing the mechanical gate (test_coverage=0 → all metrics zeroed → status=failed) says only that the tests didn't run/pass under the scorer — not that the model produced bad code. Historically most "failures" in this repo were harness/measurement artifacts, not model defects. This skill is the judgment layer on top of retort diagnose: it establishes the real cause with direct evidence before anyone attributes, reports, or gates on it.

It exists because we repeatedly drew wrong conclusions by inferring from a signature ("flaky", "concurrency", a wrong env var, a wrong tool-permission fix) that end-to-end checks later overturned. The cost of guessing here is hours and bad decisions.

Cardinal rules

  1. Ground-truth before concluding. Run the tests yourself, inspect the workspace, read the agent logs. Never attribute a cause from the gate verdict, a duration, or a single log line alone.
  2. Label every claim [DIRECT] (observed: command output, file state, a reproduced result) vs [HYPOTHESIS] (inferred, not yet tested). Don't ship a hypothesis as a cause. If a cause is unconfirmed, say so and name the experiment that would confirm it.
  3. Read the DB and the agent logs. The recorded scores/session show which tool or step failed; the agent stderr shows why (the permission type, the error). You usually need both — e.g. the session said "read rejected", but only --print-logs stderr revealed the real permission was external_directory.
  4. Verify any fix end-to-end (through the harness), not just CLI/unit. A fix that passes a CLI probe can still fail in the runner (a wrong env var passed both).

Inputs in each runs/<cell>/repN/

FileUse
TASK.mdWhat the agent was asked to build
stack.jsonlanguage / agent / model / tooling for this cell
scores.jsonrecorded metrics (all 0 ⇒ gate-failed)
generated sourcewhat the agent actually produced — check it exists and where (subdirs/packages count)
_agent_stdout.logthe agent's full --format json / --mode json event stream (tool calls + their status, errors, the final text)
_agent_stderr.logthe agent's internal logs (permission evaluations, hangs, provider errors) — for opencode, requires --print-logs

If _agent_*.log are absent, the run predates log capture; reconstruct from the agent's own session db (e.g. opencode's OPENCODE_DB sibling) and add capture before re-running.

Procedure

  1. Mechanical pass first: retort diagnose --experiment-dir <dir> labels each failed run TOOLING (recovers with retort rescore --only-failed) vs GENUINE. Start here; this skill handles the cases it can't auto-resolve.
  2. Ground-truth each failure (do NOT skip even for GENUINE):
    • Does code exist, and where? A "no code" verdict is often just a package subdir (app/, src/) the eye/glob missed.
    • Run the tests yourself in a clean copy with deps installed (pytest / cargo test / go test ./... / mvn test / …). Passing tests ⇒ it's a scorer false-fail, not a model miss.
    • Read _agent_stdout.log for the failing tool call + status; read _agent_stderr.log for the reason.
  3. Classify the cause (see taxonomy). State it [DIRECT] with the evidence.
  4. Isolate intermittent/ambiguous causes with a controlled A/B — change exactly one variable (e.g. shared vs isolated db; baseline vs with-fix batch) and compare. Log what you dropped; don't conclude from n=1.
  5. Reproduce-with-full-capture before declaring an intermittent cause — run it enough times to catch one with _agent_*.log captured, then read the events at the abort.
  6. Verify the fix end-to-end through the harness, then re-score / re-run.
Show full SKILL.md (323 more words)Show less

Cause taxonomy

A. Infrastructure false-fail (code works; the harness mis-scored it → fix the harness, then rescore):

  • Scorer gap — language/runner the scorer doesn't handle (e.g. a Bun bun:test project the TS scorer can't measure).
  • Undeclared / transitive deps — model omits requirements.txt; tests fail at import (incl. transitive deps like httpx for fastapi/starlette TestClient).
  • Permission auto-deny — headless agent denies a tool (opencode's external_directory: ask→deny on the temp workspace) → aborts with no code.
  • Concurrency/env contention — shared state (opencode's single opencode.db) → fast-fail at startup under parallelism.
  • Mis-measurement — coverage parsed as 0 on passing tests, etc.

B. Genuine model miss (real, count it; not fixable in the harness):

  • Won't compile / tests genuinely fail / no tests written / no or partial deliverable / code in the wrong place the task didn't ask for.

C. Environment (transient/operational, not the model or harness logic):

  • Rate-limit / credit cutoff (check the provider key's usage/limit), stray processes, resource exhaustion, OS cleanup.

Tells (and their trap)

The classic tell — fails instantly for ~$0 ⇒ harness; burns model time ⇒ genuine — is a starting hint, not a verdict. It breaks when scoring is the failure point: a run can burn minutes producing working code and still be scored 0 (scorer false-fail). And a fast no-code abort can be a harness permission denial, not a model giving up. Always ground-truth.

Output

Per failure, report: cell · cause (one of the taxonomy) · [DIRECT]/[HYPOTHESIS] with the evidence · recommended action (retort rescore; fix scorer/harness then rescore; count as genuine; re-run with capture). Then the rolled-up verdict: how many failures are infra vs genuine, and whether the infrastructure is trustworthy enough to proceed.

Logging setup (prerequisite)

  • The runner persists _agent_stdout.log / _agent_stderr.log per run (_persist_agent_output).
  • For opencode, the harness passes --print-logs so stderr carries the diagnostic logs (permission evaluations, step loop). Raising opencode's level (--log-level DEBUG, OPENCODE_LOG_LEVEL=verbose) adds ~nothing — the --format json stdout stream is the signal.

See also

  • retort diagnose / retort rescore — automated TOOLING/GENUINE triage + recovery.
  • evaluate-run skill — the per-run spec/quality evaluation (this skill is for failures).

© adrianco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/diagnose-failed-run of adrianco/retort.

Open the folder on GitHubat commit 1f75769

Compare with similar skills

Diagnose Failed Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Diagnose Failed Run compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Diagnose Failed Run this skilladrianco/retort207—~1.7kAutomated safety check: PassApache-2.0
Requirementsrizsotto/Bear6.5k—~2kAutomated safety check: PassGPL-3.0
Crap Analysisardalis/RiverBooks1352 repos~3.4kAutomated safety check: PassNone
Code Coverages3s-project/s3s311—~789Automated safety check: PassApache-2.0
Project Statusbactopia/bactopia522—~787Automated safety check: PassMIT
Check Coverageldayton/Dippy243—~403Automated safety check: PassMIT

Similar skills

  • Requirements

    rizsotto/Bear

    Write, modify, or review a requirement file under docs/requirements -- pick the single owning file, keep the text contract-only, name IDs so they need no explanation, and verify cross-references and…

    6.5k GitHub stars~2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Crap Analysis

    ardalis/RiverBooks

    Analyze code coverage and CRAP (Change Risk Anti-Patterns) scores to identify high-risk code.

    135 GitHub starsUsed in 2 repos~3.4k tokens
    Testing & QAAuto-check passed
  • Code Coverage

    s3s-project/s3s

    Measure and grow the line coverage of the s3s crate. An agent skill from s3s-project/s3s.

    311 GitHub stars~789 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Project Status

    bactopia/bactopia

    Show a live snapshot of the Bactopia project state — component counts, GroovyDoc coverage, nf-test coverage, and structural issues.

    522 GitHub stars~787 tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Check Coverage

    ldayton/Dippy

    Ensure comprehensive test coverage for a CLI handler. An agent skill from ldayton/Dippy.

    243 GitHub stars~403 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Guarding Destructive Operations

    kajisho5/ffmpeg-skill

    Add and review preconditions on operations that delete, overwrite, rewrite history, or resolve a caller-supplied name to a filesystem path — refusing instead of warning, placing the guard ahead of…

    1.9k GitHub stars~2.6k tokensUpdated 4 days ago
    Testing & QAAuto-check passed

More from adrianco/retort

  • Compare Runs

    adrianco/retort

    Compare evaluated runs in a retort experiment along factor dimensions.

    207 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Evaluate Run

    adrianco/retort

    Evaluate a single retort experiment run. An agent skill from adrianco/retort.

    207 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • File Run Issues

    adrianco/retort

    Aggregate a retort run's findings.jsonl into a machine-readable assessment.json summary with severity counts, penalty score, requirement coverage, and top findings.

    207 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Run Summary

    adrianco/retort

    Summarize the architecture of code generated by a single retort run.

    207 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Update Optimal Blog

    adrianco/retort

    Refresh the data tables in optimal-blog.md from master.db. An agent skill from adrianco/retort.

    207 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Categories

Questions about Diagnose Failed Run

What does Diagnose Failed Run do?

Determine the TRUE cause of a failed retort run before attributing it. Diagnose Failed Run is an agent skill from adrianco/retort. Determine the TRUE cause of a failed retort run before attributing it.

When should I use Diagnose Failed Run?

Diagnose Failed Run fits situations like: A run is recorded failed (testcoverage=0 / gate fail); A pass rate looks low; youre deciding whether a failure is real before reporting.

How do I install Diagnose Failed Run in Claude Code?

Run `npx skills add adrianco/retort --skill diagnose-failed-run -a claude-code`. Or copy the skill folder (skills/diagnose-failed-run in adrianco/retort) into .claude/skills/diagnose-failed-run in your project. Claude Code loads it when a task matches its description.

How do I install Diagnose Failed Run in Codex?

Run `npx skills add adrianco/retort --skill diagnose-failed-run -a codex`. Or copy the skill folder (skills/diagnose-failed-run in adrianco/retort) into .agents/skills/diagnose-failed-run in your project. Codex loads it when a task matches its description.

Can I use Diagnose Failed Run in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adrianco/retort --skill diagnose-failed-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diagnose-failed-run, .gemini/skills/diagnose-failed-run, .github/skills/diagnose-failed-run and .opencode/skills/diagnose-failed-run in your project.

What does Diagnose Failed Run need to run?

Going by SKILL.md and its folder, Diagnose Failed Run needs the command-line tools its instructions call (cargo, go and mvn).

Does Diagnose Failed Run access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Diagnose Failed Run safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Diagnose Failed Run use?

Diagnose Failed Run is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Diagnose Failed Run use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Diagnose Failed Run?

Skills that share tags, products or a category with Diagnose Failed Run: Requirements (rizsotto/Bear, 6.5k stars), Crap Analysis (ardalis/RiverBooks, 135 stars), Code Coverage (s3s-project/s3s, 311 stars) and Project Status (bactopia/bactopia, 522 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Diagnose Failed Run?

adrianco (a GitHub user) maintains it in adrianco/retort, which has 207 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 9, 2026.

Source: adrianco/retort on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.