Agent skill

Hypothesis-Driven Debugging

by LichAmnesia in LichAmnesia/lich-skills

Replaces trial-and-error fixing with an observe, hypothesize, experiment and conclude loop kept in DEBUG.md, where no fix is allowed before evidence supports a cause.

MITAuto-check passedDevelopment

Install Hypothesis-Driven Debugging

skills CLI
$ npx skills add LichAmnesia/lich-skills --skill debug-hypothesis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LichAmnesia/lich-skills debug-hypothesis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LichAmnesia/lich-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/debug-hypothesis .claude/skills/debug-hypothesis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-hypothesis
GitHub stars
234
Token cost
~2.5k tokens
SKILL.md length
1,280 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Replaces trial-and-error fixing with an observe, hypothesize, experiment and conclude loop kept in DEBUG.md, where no fix is allowed before evidence supports a cause.

  • Works in 4 steps: OBSERVE → HYPOTHESIZE → EXPERIMENT → …
  • A failing test whose cause is not obvious
  • SKILL.md covers When to Use, When NOT to Use, The Debug Loop and Phase 1: OBSERVE, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Any non-trivial bug goes through a four-phase investigation: wrong output, a crash, a flaky test, a performance regression or a failure that shows up only in CI. The central rule is that no fix code is written until evidence supports a hypothesis. Each phase has a goal, hard rules and a table of the excuses an agent tends to invent for skipping it.

All reasoning is written to a DEBUG.md file so context compaction cannot erase it. Observe means reproducing the bug, finding a minimal reproduction, recording the environment and noting what still works. Even a gut feeling must be written down as a hypothesis and tested, and each experiment may change at most five lines, otherwise the hypothesis is split. It is not meant for typos, missing imports or errors whose cause the compiler already states, and it takes over once a bug survives one fix attempt.

When your agent uses it

  • A failing test whose cause is not obvious
  • A bug that was fixed twice and keeps returning
  • Behavior that differs between local and CI or between dev and prod
  • An agent stuck applying the same wrong fix again and again

Example prompts

  • “The checkout test fails only in CI. Debug it with the hypothesis loop and keep notes in DEBUG.md.”
  • “Your first fix did not work. Stop and investigate properly before changing anything else.”
  • “Latency doubled after yesterday's deploy. Find the cause without guessing.”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. OBSERVE
  2. HYPOTHESIZE
  3. EXPERIMENT
  4. CONCLUDE

What it can do on your machine

Read from SKILL.md and the folder at commit ebbc355. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hypothesis-Driven Debugging loads about 2.5k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 1,280 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LichAmnesia/lich-skills at commit ebbc355, republished under its MIT licence (© LichAmnesia). 1,280 words, ~2,532 tokens.

Download SKILL.mdSave it as .claude/skills/debug-hypothesis/SKILL.md (or your agent's skills folder).
name
debug-hypothesis
description
Use when debugging any non-trivial bug — wrong output, crash, flaky test, performance regression, or "it works locally but not in CI." Forces a scientific-method loop (Observe → Hypothesize → Experiment → Conclude) so the agent stops guessing and starts reasoning. Prevents the

Hypothesis-Driven Debugging

A four-phase loop that turns debugging from "try random fixes and hope" into a disciplined investigation. Each phase has a goal, hard rules, and a rationalization table for the excuses an agent will invent to skip it.

The core principle: you may not write a fix until you have evidence that your hypothesis is correct. Guessing is not debugging.

When to Use

  • A test fails and the cause is not immediately obvious
  • The same bug has been "fixed" twice and came back
  • The agent tried a fix that didn't work — stop it from trying another
  • A crash or error message you haven't seen before
  • Performance regression with no obvious culprit
  • Behavior differs between environments (local vs CI, dev vs prod)
  • The agent is stuck in a loop, applying the same wrong fix

When NOT to Use

  • Typos, missing imports, or syntax errors — just fix them
  • Build failures with an obvious single-line cause
  • Compiler/linter messages that tell you exactly what and where
  • You already know the root cause and just need to write the fix

If the bug survived one fix attempt, switch to this skill immediately.

The Debug Loop

  OBSERVE  ──▶  HYPOTHESIZE  ──▶  EXPERIMENT  ──▶  CONCLUDE
     │              │                  │               │
     ▼              ▼                  ▼               ▼
  Gather          List 3-5         One minimal       Root cause
  symptoms,       possible         test per          confirmed
  reproduce       causes +         hypothesis,       or loop
  reliably        evidence         max 5 lines       back
     │              │                  │               │
     └──────────────┴──────────────────┴───────────────┘
              write everything to DEBUG.md

Hard rules:

  1. Everything gets written to DEBUG.md. Context compaction will eat your reasoning if it only lives in the conversation.
  2. You may not write fix code during Observe, Hypothesize, or Experiment.
  3. You may not skip Hypothesize. "I think I know what it is" is a hypothesis — write it down and test it like the others.
  4. Each experiment changes at most 5 lines. If your experiment needs more, your hypothesis is too vague — split it.

Phase 1: OBSERVE

Goal. Collect raw facts. Reproduce the bug. Separate what you know from what you assume.

Steps.

  1. Reproduce the bug. Get the exact error message, stack trace, or wrong output. If you cannot reproduce it, that is your first finding.
  2. Find the minimal reproduction. Strip away unrelated code until the bug still appears.
  3. Record the environment: OS, runtime version, dependencies, config.
  4. Note what does work. The boundary between working and broken is where the bug lives.
  5. Write all observations to DEBUG.md under ## Observations.

Exit criteria.

  • Bug reproduced (or documented as non-reproducible with conditions)
  • Exact error message or wrong behavior recorded
  • Minimal reproduction identified
  • Observations written to DEBUG.md

Common Rationalizations

ExcuseReality
"I already know what's wrong"Then write it as a hypothesis and prove it. If you're right, it takes 2 minutes.
"Let me just try this quick fix first"That's how you end up 45 minutes deep with 6 failed attempts.
"The error message is clear enough"Error messages describe symptoms, not causes. NullPointerException tells you what died, not why.
"I don't need to reproduce it, I can see the bug in the code"Can you? Then why hasn't it been fixed yet?

Phase 2: HYPOTHESIZE

Goal. Generate 3-5 possible root causes. For each, list supporting and conflicting evidence from Phase 1. Rank by likelihood.

Steps.

  1. List 3-5 hypotheses. Not 1. Not "I think it's X." Three minimum. Think across categories:
    • Data: wrong input, missing field, type mismatch, encoding
    • Logic: wrong condition, off-by-one, race condition, wrong order
    • Environment: config, version, dependency, permissions
    • State: stale cache, leaked state, initialization order
  2. For each hypothesis, write:
    • Supports: evidence from observations that backs this theory
    • Conflicts: evidence that argues against it
    • Test: the minimal experiment that would prove or disprove it
  3. Mark the ROOT HYPOTHESIS — the one with supporting evidence and no conflicting evidence. If multiple qualify, pick the easiest to test.
  4. Write everything to DEBUG.md under ## Hypotheses.

Example format in DEBUG.md:

markdown
## Hypotheses

### H1: Race condition in session middleware (ROOT HYPOTHESIS)
- Supports: only happens under concurrent requests, timing-dependent
- Conflicts: none yet
- Test: add mutex lock around session read, check if bug disappears

### H2: Stale cache returning expired token
- Supports: works after restart (cache cleared)
- Conflicts: cache TTL is 5min, bug appears within 30s
- Test: disable cache, reproduce

### H3: Wrong env variable in CI
- Supports: works locally, fails in CI
- Conflicts: env diff shows identical values
- Test: print actual runtime value in CI logs

Exit criteria.

  • At least 3 hypotheses written
  • Each has supporting/conflicting evidence
  • Each has a specific, minimal test
  • ROOT HYPOTHESIS identified
  • All written to DEBUG.md

Common Rationalizations

ExcuseReality
"I only have one theory"You have one favorite theory. Think harder. What if it's not that?
"Writing this down is slow"Debugging without writing is slower. You'll forget hypothesis 2 after compaction eats it.
"The first hypothesis is obviously right"Then proving it takes 2 minutes. If you skip proof, you'll spend 30 minutes when it turns out wrong.
"I don't have conflicting evidence"That means you haven't looked hard enough, or it really is the root cause. Either way, test it.
Show full SKILL.md (570 more words)Show less

Phase 3: EXPERIMENT

Goal. Test the ROOT HYPOTHESIS with the smallest possible change. You are a scientist — you are trying to falsify, not confirm.

Steps.

  1. Write the experiment before running it. What will you change? What result confirms the hypothesis? What result rejects it?
  2. Make the change. Maximum 5 lines. If you need more, your hypothesis is too vague.
  3. Run the reproduction from Phase 1.
  4. Record the result in DEBUG.md under ## Experiments.

Experiment rules.

  • One variable at a time. Do not combine two fixes "to save time."
  • Do not write production fix code. Write diagnostic code: log statements, assertions, simplified logic, hardcoded values.
  • Revert the experiment after recording results. Keep the tree clean.
  • If the experiment is inconclusive, that's a result — record it and test the next hypothesis.

Exit criteria.

  • Experiment executed (one change, one variable)
  • Result recorded: confirmed, rejected, or inconclusive
  • Experimental code reverted
  • Results written to DEBUG.md

Common Rationalizations

ExcuseReality
"Let me just fix it instead of testing"Fixing without confirming the cause is how you ship a wrong fix that breaks something else.
"I'll test two things at once to save time"When both change and the bug disappears, which one fixed it? Now you have to test again.
"5 lines isn't enough"5 lines is enough to add a log, an assertion, a hardcoded value, or a short-circuit. If it isn't, your hypothesis is "something is wrong somewhere" — not a hypothesis.
"I don't need to revert, the fix is basically the experiment"The experiment is diagnostic. The fix is production code. They have different quality bars.

Phase 4: CONCLUDE

Goal. Confirm root cause, write the real fix, and add a regression test.

Steps.

  1. If ROOT HYPOTHESIS confirmed:
    • Write the root cause in one sentence in DEBUG.md.
    • Now — and only now — write production fix code.
    • Add a regression test that fails without the fix and passes with it.
    • Commit fix and test together.
  2. If ROOT HYPOTHESIS rejected:
    • Record the rejection and evidence in DEBUG.md.
    • Promote the next hypothesis to ROOT. Return to Phase 3.
    • If all hypotheses rejected, return to Phase 1 with new observations.
  3. Update DEBUG.md with the final ## Root Cause and ## Fix sections.

Exit criteria.

  • Root cause identified and written in one sentence
  • Fix committed
  • Regression test committed
  • DEBUG.md complete with full investigation trail
  • Original reproduction case now passes

Common Rationalizations

ExcuseReality
"I don't need a regression test, it's a simple fix"Simple fixes for simple bugs don't need this skill. You're here because it wasn't simple. Add the test.
"The DEBUG.md is just for debugging, I'll delete it"Keep it. Future-you debugging the same area will thank present-you.
"All hypotheses failed, I'm stuck"Go back to Observe. You missed something. The bug exists, therefore a cause exists.

The Anti-Bulldozer Rule

The #1 failure mode of AI debugging: the agent forms a theory, writes 150 lines of "fix" code, it doesn't work, so it writes another 150 lines going deeper into the same wrong theory.

This skill exists to prevent that. If you catch yourself or the agent:

  • Writing more than 5 lines before confirming a hypothesis → STOP. Back to Phase 2.
  • Trying the same approach a second time → STOP. The hypothesis is rejected. Next one.
  • Ignoring conflicting evidence → STOP. Write it down. Re-rank hypotheses.
  • Feeling "almost there" after 3 failed attempts → STOP. You are bulldozing.

Write it down. Test it. Prove it. Then fix it.

© LichAmnesia, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/debug-hypothesis of LichAmnesia/lich-skills.

Open the folder on GitHubat commit ebbc355

Compare with similar skills

Hypothesis-Driven Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hypothesis-Driven Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hypothesis-Driven Debugging this skillLichAmnesia/lich-skills234—~2.5kAutomated safety check: PassMIT
Root Cause Debuggingjsmastery-pro/skills1.4k—~1.8kAutomated safety check: NotesMIT
Superpowers Systematic Debuggingchristopherarter/superpowers-reasonix102—~2kAutomated safety check: PassMIT
Minimal Code Fixcobusgreyling/loop-engineering11k1 repos~345Automated safety check: NotesMIT
Failure Diagnosis LoopTotoro-jam/battle-tested-patterns343—~319Automated safety check: PassMIT
Systematic Debuggingcbrock84/headcount2k—~657Automated safety check: PassMIT

Similar skills

  • Root Cause Debugging

    jsmastery-pro/skills

    Runs a reproduce, localize, hypothesize, test, fix and verify loop to find a bug's root cause, applies the minimal fix and hands off a regression test.

    1.4k GitHub stars~1.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Superpowers Systematic Debugging

    christopherarter/superpowers-reasonix

    Any bug, failing or flaky test, or surprise behavior?. An agent skill from christopherarter/superpowers-reasonix.

    102 GitHub stars~2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Minimal Code Fix

    cobusgreyling/loop-engineering

    Makes the smallest code change that fixes one well-scoped problem, such as a CI failure, review comment or typo, without refactoring anything unrelated.

    11k GitHub starsUsed in 1 repo~345 tokens
    DevelopmentAuto-check: notes
  • Failure Diagnosis Loop

    Totoro-jam/battle-tested-patterns

    Walks the agent through a fixed loop for failing tests and build errors: reproduce, isolate, hypothesize, instrument, fix, verify, then add a regression test.

    343 GitHub stars~319 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Systematic Debugging

    cbrock84/headcount

    Finds the root cause of a bug, test failure, or unexpected behavior before proposing any fix.

    2k GitHub stars~657 tokensUpdated 19 days ago
    DevelopmentAuto-check passed
  • Test Guided Bug Detector

    ArabelaTso/Skills-4-SE

    Analyze failing tests to detect functional bugs in code. An agent skill from ArabelaTso/Skills-4-SE.

    253 GitHub stars~2.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from LichAmnesia/lich-skills

All 11 skills in this repo
  • Google Analytics 4 Analysis

    LichAmnesia/lich-skills

    Pulls Google Analytics 4 data through the Data API with TypeScript scripts and turns it into a daily SEO report or prioritized traffic and bounce-rate recommendations.

    234 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check: notes
  • Nano Banana Image Generator

    LichAmnesia/lich-skills

    Generates or edits PNG images with Google's Nano Banana 2 model through a small script, with a choice of 512, 1K, 2K or 4K output.

    234 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check: notes
  • Spec-Driven Development v2

    LichAmnesia/lich-skills

    Organizes long-running agent work into a Project, Sprint and Task hierarchy with per-task state files, isolated worktrees, review loops and script-checked rules.

    234 GitHub stars~3.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Tavily Web Search

    LichAmnesia/lich-skills

    Runs headless web searches and single-page extraction through the Tavily API from a Python script, returning cited, summarized results without a browser.

    234 GitHub stars~1k tokensUpdated 3 mo ago
    Auto-check: notes
  • Build Until Pass Loop

    LichAmnesia/lich-skills

    Drives a failing build, typecheck, lint or test command to a passing exit code through small, one-fix-at-a-time rounds, stopping at a hard attempt cap instead of looping forever.

    234 GitHub stars~2.8k tokensUpdated 3 mo ago
    Auto-check passed
  • Spec-Driven Development

    LichAmnesia/lich-skills

    Runs a gated Spec, Plan, Build, Test, Review, Ship workflow so non-trivial changes are specified, verified and reviewed before they ship, with a named artifact per phase.

    234 GitHub stars~3.5k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Hypothesis-Driven Debugging

What does Hypothesis-Driven Debugging do?

Replaces trial-and-error fixing with an observe, hypothesize, experiment and conclude loop kept in DEBUG.md, where no fix is allowed before evidence supports a cause. Any non-trivial bug goes through a four-phase investigation: wrong output, a crash, a flaky test, a performance regression or a failure that shows up only in CI. The central rule is that no fix code is written until evidence supports a hypothesis.

When should I use Hypothesis-Driven Debugging?

Hypothesis-Driven Debugging fits situations like: A failing test whose cause is not obvious; A bug that was fixed twice and keeps returning; behavior that differs between local and CI or between dev and prod; an agent stuck applying the same wrong fix again and again.

How do I install Hypothesis-Driven Debugging in Claude Code?

Run `npx skills add LichAmnesia/lich-skills --skill debug-hypothesis -a claude-code`. Or copy the skill folder (skills/debug-hypothesis in LichAmnesia/lich-skills) into .claude/skills/debug-hypothesis in your project. Claude Code loads it when a task matches its description.

How do I install Hypothesis-Driven Debugging in Codex?

Run `npx skills add LichAmnesia/lich-skills --skill debug-hypothesis -a codex`. Or copy the skill folder (skills/debug-hypothesis in LichAmnesia/lich-skills) into .agents/skills/debug-hypothesis in your project. Codex loads it when a task matches its description.

Can I use Hypothesis-Driven Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LichAmnesia/lich-skills --skill debug-hypothesis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-hypothesis, .gemini/skills/debug-hypothesis, .github/skills/debug-hypothesis and .opencode/skills/debug-hypothesis in your project.

What does Hypothesis-Driven Debugging need to run?

SKILL.md names no scripts, command-line tools or credentials: Hypothesis-Driven Debugging is instructions for the agent only.

Does Hypothesis-Driven Debugging access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hypothesis-Driven Debugging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hypothesis-Driven Debugging use?

Hypothesis-Driven Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hypothesis-Driven Debugging use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hypothesis-Driven Debugging?

Skills that share tags, products or a category with Hypothesis-Driven Debugging: Root Cause Debugging (jsmastery-pro/skills, 1.4k stars), Superpowers Systematic Debugging (christopherarter/superpowers-reasonix, 102 stars), Minimal Code Fix (cobusgreyling/loop-engineering, 11k stars) and Failure Diagnosis Loop (Totoro-jam/battle-tested-patterns, 343 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hypothesis-Driven Debugging?

LichAmnesia (a GitHub user) maintains it in LichAmnesia/lich-skills, which has 234 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on June 9, 2026.

Source: LichAmnesia/lich-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.