Agent skill

Fault Diagnosis

by NoobyGains in NoobyGains/godmode

A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes

MITAuto-check passedTesting & QA

Install Fault Diagnosis

skills CLI
$ npx skills add NoobyGains/godmode --skill fault-diagnosis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NoobyGains/godmode fault-diagnosis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NoobyGains/godmode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/fault-diagnosis .claude/skills/fault-diagnosis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fault-diagnosis
GitHub stars
109
Token cost
~2.5k tokens
SKILL.md length
1,235 words
Files
6
Skills in repo
34
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes

  • Works in 4 steps: Root Cause Investigation → Pattern Analysis → Hypothesis and Testing → …
  • Encountering any bug
  • SKILL.md covers Overview, The Prime Directive, When to Use and The Four Phases, plus 7 more sections
  • Runs TypeScript and Shell scripts from its folder

What it does

Fault Diagnosis is an agent skill from NoobyGains/godmode. Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `condition-based-waiting-example.ts`, `condition-based-waiting.md` and `defense-in-depth.md`).

It sits in Testing & QA, covering Failing and flaky tests. The repository describes itself as: The AI development framework that thinks before it builds. 36 composable skills for Claude Code, Cursor, Codex, and OpenCode. The licence is MIT.

When your agent uses it

  • Encountering any bug
  • Unexpected behavior
  • Before proposing fixes

Example prompts

  • “/fault-diagnosis”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Root Cause Investigation
  2. Pattern Analysis
  3. Hypothesis and Testing
  4. Implementation

What it can do on your machine

Read from SKILL.md and the folder at commit 441103a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fault Diagnosis loads about 2.5k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 1,235 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NoobyGains/godmode at commit 441103a, republished under its MIT licence (© NoobyGains). 1,235 words, ~2,535 tokens.

Download SKILL.mdSave it as .claude/skills/fault-diagnosis/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
fault-diagnosis
description
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

Fault Diagnosis

Overview

Guessing at fixes wastes time and introduces new defects. Quick patches mask underlying problems.

Core principle: ALWAYS identify root cause before attempting any fix. Treating symptoms is failure.

No exceptions. No workarounds. No shortcuts.

The Prime Directive

NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST

If you have not completed Phase 1, you are not authorized to propose fixes.

When to Use

Apply to ANY technical issue:

  • Test failures
  • Production bugs
  • Unexpected behavior
  • Performance degradation
  • Build failures
  • Integration breakdowns

Especially important when:

  • Under time pressure (urgency makes guessing tempting)
  • "Just one quick fix" seems obvious
  • You have already attempted multiple fixes
  • A previous fix did not resolve the issue
  • You do not fully understand the problem

Do not skip when:

  • The issue appears simple (simple bugs have root causes too)
  • You are in a hurry (systematic investigation is faster than flailing)
  • Someone wants it resolved NOW (methodical work is faster than thrashing)

The Four Phases

You MUST complete each phase before advancing to the next.

Phase 1: Root Cause Investigation

BEFORE attempting ANY fix:

  1. Read Error Messages Thoroughly

    • Do not skip past errors or warnings
    • They frequently contain the exact answer
    • Read stack traces completely
    • Note line numbers, file paths, error codes
  2. Reproduce Reliably

    • Can you trigger it consistently?
    • What are the exact reproduction steps?
    • Does it happen every time?
    • If not reproducible, gather more data -- do not guess
  3. Examine Recent Changes

    • What changed that could cause this?
    • Git diff, recent commits
    • New dependencies, configuration changes
    • Environmental differences
  4. Gather Evidence in Multi-Component Systems

    WHEN the system has multiple components (CI -> build -> signing, API -> service -> database):

    BEFORE proposing fixes, add diagnostic instrumentation:

    For EACH component boundary:
      - Log what data enters the component
      - Log what data exits the component
      - Verify environment/config propagation
      - Check state at each layer
    
    Run once to collect evidence showing WHERE it breaks
    THEN analyze evidence to identify the failing component
    THEN investigate that specific component

    Example (multi-layer system):

    bash
    # Layer 1: Orchestrator
    echo "=== Orchestrator state: ==="
    echo "TOKEN: ${TOKEN:+SET}${TOKEN:-UNSET}"
    
    # Layer 2: Build script
    echo "=== Build environment: ==="
    env | grep TOKEN || echo "TOKEN not in environment"
    
    # Layer 3: Signing module
    echo "=== Certificate state: ==="
    security list-keychains
    security find-identity -v
    
    # Layer 4: Actual operation
    codesign --sign "$IDENTITY" --verbose=4 "$ARTIFACT"

    This reveals: Which layer fails (secrets -> orchestrator OK, orchestrator -> build FAIL)

  5. Trace Data Flow

    WHEN the error is deep in the call stack:

    See root-cause-tracing.md in this directory for the complete backward tracing method.

    Short version:

    • Where does the bad value originate?
    • What called this function with the bad value?
    • Keep tracing upward until you find the source
    • Fix at the source, not at the symptom
Phase 2: Pattern Analysis

Find the pattern before fixing:

  1. Locate Working Examples

    • Find similar working code in the same codebase
    • What works that resembles what is broken?
  2. Compare Against References

    • If implementing a pattern, read the reference implementation COMPLETELY
    • Do not skim -- read every line
    • Understand the pattern fully before applying
  3. Identify Differences

    • What differs between working and broken?
    • List every difference, no matter how small
    • Do not assume "that cannot matter"
  4. Understand Dependencies

    • What other components does this require?
    • What settings, configuration, environment?
    • What assumptions does it make?
Phase 3: Hypothesis and Testing

Scientific method:

  1. Form a Single Hypothesis

    • State clearly: "I believe X is the root cause because Y"
    • Write it down
    • Be specific, not vague
  2. Test Minimally

    • Make the SMALLEST possible change to test the hypothesis
    • One variable at a time
    • Do not fix multiple things simultaneously
  3. Verify Before Continuing

    • Did it work? Yes -> Phase 4
    • Did not work? Form a NEW hypothesis
    • DO NOT pile additional fixes on top
  4. When You Do Not Know

    • Say "I do not understand X"
    • Do not pretend to know
    • Ask for help
    • Research further
Phase 4: Implementation

Fix the root cause, not the symptom:

  1. Create a Failing Test Case

    • Simplest possible reproduction
    • Automated test if possible
    • One-off test script if no framework available
    • MUST exist before fixing
    • Use the godmode:test-first skill for writing proper failing tests
  2. Implement a Single Fix

    • Address the root cause identified
    • ONE change at a time
    • No "while I'm here" improvements
    • No bundled refactoring
  3. Verify the Fix

    • Test passes now?
    • No other tests broken?
    • Issue actually resolved?
  4. If the Fix Does Not Work

    • STOP
    • Count: How many fixes have you attempted?
    • If < 3: Return to Phase 1, re-analyze with new information
    • If >= 3: STOP and question the architecture (step 5 below)
    • DO NOT attempt fix #4 without architectural discussion
  5. If 3+ Fixes Failed: Question Architecture

    Pattern indicating an architectural problem:

    • Each fix reveals new shared state/coupling/problems in different locations
    • Fixes require "massive refactoring" to implement
    • Each fix creates new symptoms elsewhere

    STOP and question fundamentals:

    • Is this pattern fundamentally sound?
    • Are we persisting through sheer inertia?
    • Should we refactor the architecture vs. continue fixing symptoms?

    Discuss with your human partner before attempting more fixes

    This is NOT a failed hypothesis -- this is a flawed architecture.

Show full SKILL.md (503 more words)Show less

Guardrails - STOP and Follow Process

If you catch yourself thinking:

  • "Quick fix for now, investigate later"
  • "Just try changing X and see what happens"
  • "Apply multiple changes, run tests"
  • "Skip the test, I'll verify manually"
  • "It's probably X, let me fix that"
  • "I don't fully understand but this might work"
  • "Pattern says X but I'll adapt differently"
  • "Here are the main problems: [lists fixes without investigation]"
  • Proposing solutions before tracing data flow
  • "One more fix attempt" (when already tried 2+)
  • Each fix reveals new problems in different places

ALL of these mean: STOP. Return to Phase 1.

If 3+ fixes failed: Question the architecture (see Phase 4, step 5)

Human Partner Signals You Are Off Track

Watch for these redirections:

  • "Is that not happening?" - You assumed without verifying
  • "Will it show us...?" - You should have added evidence gathering
  • "Stop guessing" - You are proposing fixes without understanding
  • "Think deeper" - Question fundamentals, not just symptoms
  • "We're stuck?" (frustrated) - Your approach is not working

When you see these: STOP. Return to Phase 1.

Cognitive Traps

RationalizationWhat Is Actually True
"Issue is simple, process not needed"Simple issues have root causes too. The process is fast for simple bugs.
"Emergency, no time for process"Systematic diagnosis is FASTER than guess-and-check flailing.
"Just try this first, then investigate"The first fix sets the pattern. Do it right from the start.
"I'll write the test after confirming the fix works"Untested fixes do not hold. Test first proves it.
"Multiple fixes at once saves time"Cannot isolate what worked. Creates new bugs.
"Reference too long, I'll adapt the pattern"Partial understanding guarantees bugs. Read it completely.
"I see the problem, let me fix it"Seeing symptoms is not the same as understanding root cause.
"One more fix attempt" (after 2+ failures)3+ failures = architectural problem. Question the pattern, do not fix again.

Quick Reference

PhaseKey ActivitiesSuccess Criteria
1. Root CauseRead errors, reproduce, check changes, gather evidenceUnderstand WHAT and WHY
2. PatternFind working examples, compareIdentify differences
3. HypothesisForm theory, test minimallyConfirmed or new hypothesis
4. ImplementationCreate test, fix, verifyBug resolved, tests pass

When Investigation Reveals No Root Cause

If systematic investigation reveals the issue is truly environmental, timing-dependent, or external:

  1. You have completed the process
  2. Document what you investigated
  3. Implement appropriate handling (retry, timeout, error message)
  4. Add monitoring/logging for future investigation

But: 95% of "no root cause" cases are incomplete investigation.

Supporting Methods

These methods are part of fault diagnosis and available in this directory:

  • root-cause-tracing.md - Trace bugs backward through the call stack to find the original trigger
  • defense-in-depth.md - Add validation at multiple layers after finding root cause
  • condition-based-waiting.md - Replace arbitrary timeouts with condition polling

Related skills:

  • godmode:test-first - For creating failing test case (Phase 4, Step 1)
  • godmode:completion-gate - Verify fix worked before declaring success

Real-World Impact

From diagnosis sessions:

  • Systematic approach: 15-30 minutes to resolution
  • Random fix approach: 2-3 hours of flailing
  • First-attempt fix rate: 95% vs 40%
  • New bugs introduced: Near zero vs common

© NoobyGains, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/fault-diagnosis of NoobyGains/godmode.

  • SKILL.md
  • condition-based-waiting-example.ts
  • condition-based-waiting.md
  • defense-in-depth.md
  • find-polluter.sh
  • root-cause-tracing.md

Open the folder on GitHubat commit 441103a

Compare with similar skills

Fault Diagnosis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fault Diagnosis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fault Diagnosis this skillNoobyGains/godmode109—~2.5kAutomated safety check: PassMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Triage CI FailureDataDog/datadog-agent3.8k—~2.3kAutomated safety check: PassApache-2.0
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Fix Ready PRsfastrepl/anarlog9.5k—~1.4kAutomated safety check: PassMIT
Trx Analysismicrosoft/vstest969—~1.8kAutomated safety check: PassMIT

Similar skills

  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Triage CI Failure

    DataDog/datadog-agent

    Official

    Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

    3.8k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Fix Ready PRs

    fastrepl/anarlog

    Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.

    9.5k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Trx Analysis

    microsoft/vstest

    Official

    Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.

    969 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    204 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from NoobyGains/godmode

All 34 skills in this repo
  • Activation

    NoobyGains/godmode

    A skill your agent uses when starting any conversation - establishes how to locate and invoke skills, mandating Skill tool usage before ANY response including clarifying questions

    109 GitHub stars~2.4k tokensUpdated 7 mo ago
    Auto-check passed
  • Agent Messaging

    NoobyGains/godmode

    A skill your agent uses when dispatching subagents, composing prompts for teammates, structuring handoff reports, or managing context boundaries between agents.

    109 GitHub stars~3k tokensUpdated 7 mo ago
    Auto-check passed
  • Codebase Research

    NoobyGains/godmode

    A skill your agent uses when building ANY feature within an existing project - search the current codebase for existing patterns, conventions, similar implementations, and established approaches…

    109 GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check: notes
  • Completion Gate

    NoobyGains/godmode

    A skill your agent uses when about to declare work done, fixed, or passing, before committing or opening PRs - demands executing verification commands and reading their output before making any…

    109 GitHub stars~1.6k tokensUpdated 7 mo ago
    Auto-check passed
  • Comprehension Check

    NoobyGains/godmode

    A skill your agent uses when implementing any substantial feature, multi-file modification, or architectural change - produces a plain-language walkthrough of every alteration so the developer can…

    109 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Delegated Execution

    NoobyGains/godmode

    A skill your agent uses when executing implementation plans with independent tasks in the current session

    109 GitHub stars~2.4k tokensUpdated 7 mo ago
    Auto-check passed

Categories

Questions about Fault Diagnosis

What does Fault Diagnosis do?

A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes. Fault Diagnosis is an agent skill from NoobyGains/godmode.

When should I use Fault Diagnosis?

Fault Diagnosis fits situations like: encountering any bug; unexpected behavior; before proposing fixes.

How do I install Fault Diagnosis in Claude Code?

Run `npx skills add NoobyGains/godmode --skill fault-diagnosis -a claude-code`. Or copy the skill folder (skills/fault-diagnosis in NoobyGains/godmode) into .claude/skills/fault-diagnosis in your project. Claude Code loads it when a task matches its description.

How do I install Fault Diagnosis in Codex?

Run `npx skills add NoobyGains/godmode --skill fault-diagnosis -a codex`. Or copy the skill folder (skills/fault-diagnosis in NoobyGains/godmode) into .agents/skills/fault-diagnosis in your project. Codex loads it when a task matches its description.

Can I use Fault Diagnosis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NoobyGains/godmode --skill fault-diagnosis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fault-diagnosis, .gemini/skills/fault-diagnosis, .github/skills/fault-diagnosis and .opencode/skills/fault-diagnosis in your project.

What does Fault Diagnosis need to run?

Going by SKILL.md and its folder, Fault Diagnosis needs TypeScript and a shell for the scripts in its folder. Our summary lists: Node.js; A Bash shell.

Does Fault Diagnosis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Fault Diagnosis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fault Diagnosis use?

Fault Diagnosis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fault Diagnosis use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fault Diagnosis?

Skills that share tags, products or a category with Fault Diagnosis: Swig Test (swig/swig, 6.3k stars), Triage CI Failure (DataDog/datadog-agent, 3.8k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars) and Fix Ready PRs (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fault Diagnosis?

NoobyGains (a GitHub user) maintains it in NoobyGains/godmode, which has 109 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on March 9, 2026.

Source: NoobyGains/godmode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.