Agent skill

Systematic Debugging

by BlackBeltTechnology in BlackBeltTechnology/pi-agent-dashboard

Root-cause a bug already in front of you, instead of guessing at fixes.

MITAuto-check passedDevelopment

Install Systematic Debugging

skills CLI
$ npx skills add BlackBeltTechnology/pi-agent-dashboard --skill systematic-debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BlackBeltTechnology/pi-agent-dashboard systematic-debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BlackBeltTechnology/pi-agent-dashboard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/eng-disciplines/.pi/skills/systematic-debugging .claude/skills/systematic-debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
systematic-debugging
GitHub stars
316
Token cost
~1.9k tokens
SKILL.md length
966 words
Files
1
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

Root-cause a bug already in front of you, instead of guessing at fixes.

  • Works in 4 steps: Root Cause → Pattern → Hypothesis → …
  • Like root cause this
  • SKILL.md covers Overview, When to Use, The Four Phases and The Tight Feedback Loop, plus 3 more sections
  • Calls npm

What it does

Systematic Debugging is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Root-cause a bug already in front of you, instead of guessing at fixes. Use on triggers like "root cause this", "why is this failing", "debug systematically", "this test is flaky", "it works locally but not in CI", or when a fix attempt has already failed once. Enforces a phased evidence-first process before any code change. Not a feature-build or ship workflow.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Debugging and Root cause analysis. The repository describes itself as: Real-time web dashboard for pi coding-agent sessions. Multi-session view, live chat mirroring, integrated terminal, diff viewer, pi-flows execution, and mobile-first remote… The licence is MIT.

When your agent uses it

  • Like root cause this
  • Why is this failing
  • Debug systematically
  • This test is flaky

Example prompts

  • “root cause this”
  • “why is this failing”
  • “debug systematically”
  • “/systematic-debugging”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Root Cause
  2. Pattern
  3. Hypothesis
  4. Implementation

What it can do on your machine

Read from SKILL.md and the folder at commit e23e533. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Systematic Debugging loads about 1.9k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 966 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BlackBeltTechnology/pi-agent-dashboard at commit e23e533, republished under its MIT licence (© BlackBeltTechnology). 966 words, ~1,852 tokens.

Download SKILL.mdSave it as .claude/skills/systematic-debugging/SKILL.md (or your agent's skills folder).
name
systematic-debugging
description
Root-cause a bug already in front of you, instead of guessing at fixes. Use on triggers like "root cause this", "why is this failing", "debug systematically", "this test is flaky", "it works locally but not in CI", or when a fix attempt has already failed once. Enforces a phased evidence-first process before any code change. Not a feature-build or ship workflow.
related_skills
doubt-driven-review, review-code, observability-instrumentation

Systematic Debugging

Overview

A bug is a gap between what the code does and what you believe it does. Guessing at fixes closes the gap by accident, if at all — and each blind edit adds a new variable that hides the real cause. Systematic debugging is the discipline of gathering evidence until the cause is known, then changing exactly one thing.

The failure mode this skill prevents: reading a stack trace, forming an instant theory, editing code to match the theory, re-running, and repeating. That loop feels like progress and usually is not — it mutates the system faster than it explains it.

When to Use

  • A test fails and the message doesn't immediately tell you why
  • Behaviour differs between environments ("works locally, fails in CI", "works in dev mode, not production")
  • A fix you already tried didn't work (you are now on attempt ≥ 2 — stop guessing)
  • An intermittent / flaky failure you can't reproduce on demand
  • A regression: something that worked now doesn't, and you don't know which change broke it

When NOT to use:

  • The cause is already obvious and proven (typo, off-by-one you can see, wrong constant)
  • You are building a new feature, not diagnosing existing behaviour
  • The "bug" is actually a missing requirement — that's a spec conversation, not a debug session

The Four Phases

Each phase has a success criterion. Do not advance until it is met. Skipping ahead is the whole antipattern.

text
Phase 1  ROOT CAUSE      gather evidence ─▶ criterion: you can state the cause in one sentence
   │                                        with evidence, not a guess
   ▼
Phase 2  PATTERN         is this cause elsewhere? ─▶ criterion: you've searched for sibling
   │                                                 instances of the same class of bug
   ▼
Phase 3  HYPOTHESIS      change ONE variable ─▶ criterion: a prediction that, if wrong,
   │                                            disproves your theory (a real test)
   ▼
Phase 4  IMPLEMENTATION  fix + regression test ─▶ criterion: a test that fails before the fix
                                                  and passes after
Phase 1 — Root Cause

Collect evidence before forming a theory. The goal of this phase is a sentence of the form "X fails because Y, and here is the observation that shows Y."

  • Read the full error and stack trace, not the first line. The frame that matters is often three deep.
  • Reproduce deterministically. A bug you can't reproduce, you can't verify you fixed.
  • Capture the actual state at the failure point. When console.log can't reach the state (closure variables, a paused async frame, the Electron main process, WebSocket server internals), reach for the node-inspect-debugger skill — real breakpoints and a scope-chain dump beat sprinkled logs.
  • Distinguish symptom from cause. "Returns undefined" is a symptom; "the map key is built from a stale closure value" is a cause.

Success criterion: you can name the cause in one sentence, backed by an observation. If you can only say "I think it's the cache," you are not done with Phase 1.

Phase 2 — Pattern

A bug is rarely unique. Before fixing this instance, ask whether the same class of mistake exists elsewhere.

  • Grep for the same call shape / the same missing guard.
  • If one handler forgot to await, check its siblings.
  • Note candidates; you are cataloguing, not fixing them all now.

Success criterion: you've searched for sibling instances and know whether the fix is one-site or systemic.

Phase 3 — Hypothesis

State a hypothesis that could be wrong — and how you'd know. A theory that can't be disproven isn't a diagnosis, it's a belief.

  • Change one variable at a time. If you change three things and it works, you've learned nothing about which mattered — and added two new liabilities.
  • Predict the outcome before you run: "If the cause is the stale key, forcing a fresh key makes the failure disappear; if it still fails, my theory is wrong."
  • A failed prediction is a result, not a setback — it eliminates a branch.

Success criterion: you have a one-variable change and a falsifiable prediction.

Show full SKILL.md (403 more words)Show less
Phase 4 — Implementation

Only now do you write the fix.

  • Write (or update) a regression test first that reproduces the bug — it must FAIL before the fix. This is the RED step; it makes the bug concrete and proves the fix later.
  • Apply the minimal change that makes it pass.
  • Re-run the reproduction. Then run the wider suite to confirm no collateral damage.

Success criterion: a test that fails before the fix and passes after, plus a green wider suite.

The Tight Feedback Loop

Fast, captured feedback is what makes evidence cheap. Use this repo's documented convention — run once, capture, then grep the file instead of re-running to see errors:

bash
npm test 2>&1 | tee /tmp/pi-test.log     # run once, capture everything
grep -nE 'FAIL|Error|✗|✘' /tmp/pi-test.log   # find failures
grep -n -A 20 'FAIL ' /tmp/pi-test.log        # failure + context

Never rerun npm test just to re-read an error you already produced — grep the captured log. Each unnecessary rerun is latency between you and the cause.

The Rule of Three

After three failed fixes, STOP. Three misses means your model of the system is wrong, not that the fourth edit is the charm. Continuing to patch against a broken model deepens the hole.

When you hit three:

  1. Stop editing.
  2. Hand off to the doubt-driven-review skill — spawn a fresh-context adversarial reviewer to cross-examine the architecture and your assumptions, not just this line. The premise you've been protecting for three attempts is the thing to doubt.
  3. Reconcile its findings, then re-enter Phase 1 with a corrected model.

The Rule of Three is a circuit breaker against sunk-cost debugging. Honour it.

Red Flags

  • Editing code before you can state the cause in one sentence (skipped Phase 1)
  • Changing more than one variable per attempt (Phase 3 violated — you won't know what worked)
  • Reading only the first line of a stack trace
  • "Let me just try X" three times in a row without a new hypothesis (Rule of Three breach)
  • Fixing the symptom (if (x == null) return) without knowing why x is null
  • Declaring victory with no regression test — you can't prove it's fixed or stays fixed
  • Re-running the suite to re-read an error instead of grepping the captured log

Verification

  • The cause was stated in one evidence-backed sentence before any fix (Phase 1)
  • Sibling instances of the bug class were searched for (Phase 2)
  • Each attempt changed exactly one variable against a falsifiable prediction (Phase 3)
  • A regression test failed before the fix and passes after (Phase 4)
  • If three fixes failed, control was handed to doubt-driven-review rather than a fourth blind attempt

© BlackBeltTechnology, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/eng-disciplines/.pi/skills/systematic-debugging of BlackBeltTechnology/pi-agent-dashboard.

Open the folder on GitHubat commit e23e533

Compare with similar skills

Systematic Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Systematic Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Systematic Debugging this skillBlackBeltTechnology/pi-agent-dashboard316—~1.9kAutomated safety check: PassMIT
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT
Systematic DebuggingChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNone

Similar skills

  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 8 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Systematic Debugging

    ChrisWiles/claude-code-showcase

    Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

    6.1k GitHub starsUsed in 3 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    102k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed

More from BlackBeltTechnology/pi-agent-dashboard

All 66 skills in this repo
  • Browser

    BlackBeltTechnology/pi-agent-dashboard

    Browser automation via the agent-browser CLI. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    316 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • CI Troubleshoot

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose failed GitHub Actions runs for pi-agent-dashboard: the 11-file workflow taxonomy, affected-test selection, the release pipeline, known failure modes, and how to read gh run logs and…

    316 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Debug Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose problems in the running pi-agent-dashboard system: server.log, /api/health, bridge WebSocket connectivity, vitest triage, known-issue FAQ entries.

    316 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Implement

    BlackBeltTechnology/pi-agent-dashboard

    Disciplined implementation in pi-agent-dashboard: the rebuild matrix (extension→reload, server→restart, client→build+restart, openspec-apply→full rebuild) plus the project's code discipline rules.

    316 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Pi Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Monitor and control the pi-dashboard server. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    316 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Session To Guideline

    BlackBeltTechnology/pi-agent-dashboard

    Turn a pi session into a Markdown "how-we-did-it" collaboration guideline: reads the session's JSONL transcript and synthesizes a reusable playbook of which prompts worked, what had to be steered…

    316 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Systematic Debugging

What does Systematic Debugging do?

Root-cause a bug already in front of you, instead of guessing at fixes. Systematic Debugging is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Root-cause a bug already in front of you, instead of guessing at fixes.

When should I use Systematic Debugging?

Systematic Debugging fits situations like: like root cause this; why is this failing; debug systematically; this test is flaky.

How do I install Systematic Debugging in Claude Code?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill systematic-debugging -a claude-code`. Or copy the skill folder (packages/eng-disciplines/.pi/skills/systematic-debugging in BlackBeltTechnology/pi-agent-dashboard) into .claude/skills/systematic-debugging in your project. Claude Code loads it when a task matches its description.

How do I install Systematic Debugging in Codex?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill systematic-debugging -a codex`. Or copy the skill folder (packages/eng-disciplines/.pi/skills/systematic-debugging in BlackBeltTechnology/pi-agent-dashboard) into .agents/skills/systematic-debugging in your project. Codex loads it when a task matches its description.

Can I use Systematic Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill systematic-debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/systematic-debugging, .gemini/skills/systematic-debugging, .github/skills/systematic-debugging and .opencode/skills/systematic-debugging in your project.

What does Systematic Debugging need to run?

Going by SKILL.md and its folder, Systematic Debugging needs the command-line tools its instructions call (npm).

Does Systematic Debugging access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Systematic Debugging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Systematic Debugging use?

Systematic Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Systematic Debugging use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Systematic Debugging?

Skills that share tags, products or a category with Systematic Debugging: OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars), Root Cause Debugging (garrytan/gstack, 136k stars) and Graph-Based Bug Tracing (tirth8205/code-review-graph, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Systematic Debugging?

BlackBeltTechnology (a GitHub organization) maintains it in BlackBeltTechnology/pi-agent-dashboard, which has 316 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 6, 2026.

Source: BlackBeltTechnology/pi-agent-dashboard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.