Agent skill

Failure Diagnosis Loop

by Totoro-jam in Totoro-jam/battle-tested-patterns

Walks the agent through a fixed loop for failing tests and build errors: reproduce, isolate, hypothesize, instrument, fix, verify, then add a regression test.

MITAuto-check passedDevelopment

Install Failure Diagnosis Loop

skills CLI
$ npx skills add Totoro-jam/battle-tested-patterns --skill diagnose -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Totoro-jam/battle-tested-patterns diagnose --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Totoro-jam/battle-tested-patterns.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/diagnose .claude/skills/diagnose && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
diagnose
GitHub stars
345
Token cost
~319 tokens
SKILL.md length
137 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Walks the agent through a fixed loop for failing tests and build errors: reproduce, isolate, hypothesize, instrument, fix, verify, then add a regression test.

  • Works in 7 steps: Reproduce → Isolate → Hypothesize → …
  • A test fails and the cause is not obvious
  • Calls pnpm
  • A build breaks after a change

What it does

The skill forbids jumping to a fix before the failure has been reproduced. The agent runs the failing command, such as `pnpm test`, `pnpm test:rust` or `pnpm test:go`, and captures the exact error. It then narrows the problem to a test file, a test case and an assertion, and states a one-sentence hypothesis before changing any code.

To confirm or reject that hypothesis, it adds minimal logging or assertions and leaves production code alone. Only then does it apply the smallest fix it can, followed by `pnpm test && pnpm build` across the whole suite, not just the failing test. A failure sends the agent back to the first step with what it learned, and a non-trivial bug gets a regression test. The commands match the battle-tested-patterns repository.

When your agent uses it

  • A test fails and the cause is not obvious
  • A build breaks after a change
  • Working through a failing exercise test in the TypeScript, Rust, Go or Python suites

Example prompts

  • “The Go exercise tests are failing, so diagnose why before changing anything.”
  • “pnpm build breaks after my last commit. Find the cause step by step.”
  • “Reproduce the failing Rust test, isolate the case and add a regression test once it is fixed.”

Requirements

  • pnpm and the repository's test and build scripts

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Reproduce
  2. Isolate
  3. Hypothesize
  4. Instrument
  5. Fix
  6. Verify
  7. Regression Test

What it can do on your machine

Read from SKILL.md and the folder at commit 660e0e9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Failure Diagnosis Loop loads about 319 tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 137 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~319

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Totoro-jam/battle-tested-patterns at commit 660e0e9, republished under its MIT licence (© Totoro-jam). 137 words, ~319 tokens.

Download SKILL.mdSave it as .claude/skills/diagnose/SKILL.md (or your agent's skills folder).
name
diagnose
description
Structured debugging loop for exercise tests or build failures. Reproduce → isolate → hypothesize → fix → verify.

Diagnose a Failure

You are debugging a test failure or build error in this project. Follow the structured loop — do NOT jump to a fix without reproducing first.

Loop

1. Reproduce

Run the failing command and capture the exact error:

bash
pnpm test          # All tests (docs + exercises in TS/Rust/Go/Python)
pnpm test:rust     # Rust only
pnpm test:go       # Go only
pnpm test:python   # Python only (auto-finds Python ≥ 3.10)
pnpm build         # VitePress
2. Isolate

Narrow down to the smallest failing unit:

  • Which test file?
  • Which test case?
  • Which assertion?
3. Hypothesize

State your hypothesis in one sentence before changing any code.

4. Instrument

Add minimal logging or assertions to confirm/deny your hypothesis. Do NOT change production code yet.

5. Fix

Apply the minimal fix. Change as little as possible.

6. Verify

Run the full test suite, not just the failing test:

bash
pnpm test && pnpm build

If it still fails, return to step 1 with what you learned.

7. Regression Test

If the bug was non-trivial, add a test that would have caught it.

© Totoro-jam, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/diagnose of Totoro-jam/battle-tested-patterns.

Open the folder on GitHubat commit 660e0e9

Compare with similar skills

Failure Diagnosis Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Failure Diagnosis Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Failure Diagnosis Loop this skillTotoro-jam/battle-tested-patterns345—~319Automated safety check: PassMIT
Minimal Code Fixcobusgreyling/loop-engineering11k1 repos~671Automated safety check: NotesMIT
Root Cause Debuggingjsmastery-pro/skills1.5k—~1.8kAutomated safety check: NotesMIT
Superpowers Systematic Debuggingchristopherarter/superpowers-reasonix102—~2kAutomated safety check: PassMIT
Hypothesis-Driven DebuggingLichAmnesia/lich-skills234—~2.5kAutomated safety check: PassMIT
Systematic Debuggingcbrock84/headcount2k—~657Automated safety check: PassMIT

Similar skills

  • Minimal Code Fix

    cobusgreyling/loop-engineering

    Makes the smallest code change that fixes one well-scoped problem, such as a CI failure, review comment or typo, without refactoring anything unrelated.

    11k GitHub starsUsed in 1 repo~671 tokens
    DevelopmentAuto-check: notes
  • Root Cause Debugging

    jsmastery-pro/skills

    Runs a reproduce, localize, hypothesize, test, fix and verify loop to find a bug's root cause, applies the minimal fix and hands off a regression test.

    1.5k GitHub stars~1.8k tokensUpdated 2 mo ago
    DevelopmentAuto-check: notes
  • Superpowers Systematic Debugging

    christopherarter/superpowers-reasonix

    Any bug, failing or flaky test, or surprise behavior?. An agent skill from christopherarter/superpowers-reasonix.

    102 GitHub stars~2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Hypothesis-Driven Debugging

    LichAmnesia/lich-skills

    Replaces trial-and-error fixing with an observe, hypothesize, experiment and conclude loop kept in DEBUG.md, where no fix is allowed before evidence supports a cause.

    234 GitHub stars~2.5k tokensUpdated 4 mo ago
    DevelopmentAuto-check passed
  • Systematic Debugging

    cbrock84/headcount

    Finds the root cause of a bug, test failure, or unexpected behavior before proposing any fix.

    2k GitHub stars~657 tokensUpdated 23 days ago
    DevelopmentAuto-check passed
  • Test Guided Bug Detector

    ArabelaTso/Skills-4-SE

    Analyze failing tests to detect functional bugs in code. An agent skill from ArabelaTso/Skills-4-SE.

    253 GitHub stars~2.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from Totoro-jam/battle-tested-patterns

  • Adopt a Design Pattern

    Totoro-jam/battle-tested-patterns

    Matches a coding problem to one of 46 documented systems patterns, checks that it really fits, then adapts it into your codebase with a test for its invariant.

    345 GitHub stars~4.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Pattern Conformance Audit

    Totoro-jam/battle-tested-patterns

    Audits a codebase's existing patterns, such as rate limiters, circuit breakers and caches, against canonical invariants and flags mislabeled or divergent ones.

    345 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • New Pattern Authoring Workflow

    Totoro-jam/battle-tested-patterns

    Step-by-step workflow for adding a new pattern to the battle-tested-patterns repo: validate the topic, verify sources, write the doc, code and exercises.

    345 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Verify Source

    Totoro-jam/battle-tested-patterns

    Verify all production proof source links in pattern documents.

    345 GitHub stars~930 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Failure Diagnosis Loop

What does Failure Diagnosis Loop do?

Walks the agent through a fixed loop for failing tests and build errors: reproduce, isolate, hypothesize, instrument, fix, verify, then add a regression test. The skill forbids jumping to a fix before the failure has been reproduced. The agent runs the failing command, such as `pnpm test`, `pnpm test:rust` or `pnpm test:go`, and captures the exact error.

When should I use Failure Diagnosis Loop?

Failure Diagnosis Loop fits situations like: A test fails and the cause is not obvious; A build breaks after a change; working through a failing exercise test in the TypeScript, Rust, Go or Python suites.

How do I install Failure Diagnosis Loop in Claude Code?

Run `npx skills add Totoro-jam/battle-tested-patterns --skill diagnose -a claude-code`. Or copy the skill folder (.claude/skills/diagnose in Totoro-jam/battle-tested-patterns) into .claude/skills/diagnose in your project. Claude Code loads it when a task matches its description.

How do I install Failure Diagnosis Loop in Codex?

Run `npx skills add Totoro-jam/battle-tested-patterns --skill diagnose -a codex`. Or copy the skill folder (.claude/skills/diagnose in Totoro-jam/battle-tested-patterns) into .agents/skills/diagnose in your project. Codex loads it when a task matches its description.

Can I use Failure Diagnosis Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Totoro-jam/battle-tested-patterns --skill diagnose -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diagnose, .gemini/skills/diagnose, .github/skills/diagnose and .opencode/skills/diagnose in your project.

What does Failure Diagnosis Loop need to run?

Going by SKILL.md and its folder, Failure Diagnosis Loop needs the command-line tools its instructions call (pnpm). Our summary lists: pnpm and the repository's test and build scripts.

Does Failure Diagnosis Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Failure Diagnosis Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Failure Diagnosis Loop use?

Failure Diagnosis Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Failure Diagnosis Loop use?

About 319 tokens (SKILL.md is roughly 1.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Failure Diagnosis Loop?

Skills that share tags, products or a category with Failure Diagnosis Loop: Minimal Code Fix (cobusgreyling/loop-engineering, 11k stars), Root Cause Debugging (jsmastery-pro/skills, 1.5k stars), Superpowers Systematic Debugging (christopherarter/superpowers-reasonix, 102 stars) and Hypothesis-Driven Debugging (LichAmnesia/lich-skills, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Failure Diagnosis Loop?

Totoro-jam (a GitHub user) maintains it in Totoro-jam/battle-tested-patterns, which has 345 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 4, 2026.

Source: Totoro-jam/battle-tested-patterns on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.