Agent skill

Debugging

by romiluz13 in romiluz13/cc10x

Debugging discipline: feedback loop FIRST, root cause before fix, blast radius after fix.

MITAuto-check: notesDevelopment

Install Debugging

skills CLI
$ npx skills add romiluz13/cc10x --skill debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install romiluz13/cc10x debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/cc10x/skills/debugging .claude/skills/debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debugging
GitHub stars
164
Token cost
~3k tokens
SKILL.md length
1,706 words
Files
6 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Debugging discipline: feedback loop FIRST, root cause before fix, blast radius after fix.

  • Works in 4 steps: Root Cause Investigation → Pattern Analysis → Hypothesis and Testing → …
  • Tasks that involve Debugging
  • SKILL.md covers Reference Files, Feedback Loop FIRST (Before…, LSP-Powered Root Cause Tracing and The Four Phases, plus 6 more sections
  • Calls git

What it does

Debugging is an agent skill from romiluz13/cc10x. Debugging discipline: feedback loop FIRST, root cause before fix, blast radius after fix. Covers the 10-rung construction ladder, LSP-powered tracing, hypothesis quality criteria, and four-phase investigation. Loaded by bug-investigator.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `evals/pressure-test-1-outage.md`, `evals/pressure-test-2-sunk-cost.md` and `evals/pressure-test-3-authority.md`).

It sits in Development, covering Debugging and Root cause analysis. The repository describes itself as: The Loop Engine for Claude Code — engineer the loop, not the prompt. 1 router · 9 agents · 16 skills · 4 workflows. Fail-closed gates, test honesty, anti-anchored review. The licence is MIT.

When your agent uses it

  • Tasks that involve Debugging
  • Tasks that involve Root cause analysis

Example prompts

  • “/debugging”

Requirements

  • Pre-approved tools (allowed-tools): Read, Edit, Bash, Grep, Glob, LSP

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Root Cause Investigation
  2. Pattern Analysis
  3. Hypothesis and Testing
  4. Implementation

What it can do on your machine

Read from SKILL.md and the folder at commit 891acf0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Edit
    • Bash
    • Grep
    • Glob
    • LSP

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debugging loads about 3k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 1,706 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Edit, Bash, Grep, Glob, LSP

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from romiluz13/cc10x at commit 891acf0, republished under its MIT licence (© romiluz13). 1,706 words, ~3,004 tokens.

Download SKILL.mdSave it as .claude/skills/debugging/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
debugging
description
Debugging discipline: feedback loop FIRST, root cause before fix, blast radius after fix. Covers the 10-rung construction ladder, LSP-powered tracing, hypothesis quality criteria, and four-phase investigation. Loaded by bug-investigator.
allowed-tools
Read, Edit, Bash, Grep, Glob, LSP
user-invocable
false

Debugging

Feedback Loop FIRST: No hypothesis without a repro loop. No fix without root cause. No fix without blast radius scan.

Reference Files

  • references/investigation-hygiene.md — investigation discipline, evidence handling
  • references/root-cause-playbooks.md — scenario-specific debugging playbooks

Feedback Loop FIRST (Before Any Hypothesis)

A hypothesis without a repro loop is a guess. Before H1, build a fast, deterministic, agent-runnable signal that turns red on the bug.

Construction Ladder (try in rank order, stop at first that works — ordered by loop tightness: earlier rungs are faster and more deterministic)
  1. Failing automated test (unit/integration) — best: lives at a seam, reusable as RED
  2. curl/HTTP request with asserted response
  3. CLI snapshot diff (run command, diff stdout/stderr/exit)
  4. Headless browser script (real DOM/runtime crash)
  5. Trace replay (recorded request/log/event re-run)
  6. Throwaway harness (tiny script calling the suspect function)
  7. Property/fuzz check (when failing input is unknown)
  8. git bisect run (regression with existing test)
  9. Differential old-vs-new (last-good vs HEAD behavior diff)
  10. Human-in-the-loop (LAST resort: scripted manual steps)

Tighten the loop — treat it as a product. Once you have a loop, keep tightening:

  • Faster? Cache setup, skip unrelated init, narrow the test scope — sub-second beats sub-minute.
  • Sharper signal? Assert the exact failing fact, not a noisy superset — never just "didn't crash".
  • More deterministic? Pin time, seed RNG, isolate filesystem, freeze network — same input → same red, no drift. A 30-second flaky loop is barely better than none; a 2-second deterministic one is a debugging superpower.

Red-capable completion criteria — the loop is done when you can name one command (a script path, a test invocation, a curl) that you have already run at least once (paste the invocation and its output), and it is:

  • Red-capable — drives the actual bug code path and asserts the user's exact symptom (can go red on this bug, green once fixed). Not "runs without erroring".
  • Deterministic — same verdict every run (flaky bugs: a pinned, high reproduction rate).
  • Fast — seconds, not minutes.
  • Agent-runnable — you can run it unattended. No red-capable command, no hypothesis phase. If you catch yourself reading code to build a theory before this command exists, STOP.

Flaky bugs: run in a tight loop (for i in $(seq 1 N); do ...; done), record hit rate (e.g. 3/50), treat raising that rate as loop iteration.

When You Genuinely Cannot Build a Loop

STOP. Do NOT advance to hypothesis. Return BLOCKED with:

  • What was tried: each rung attempted and why it failed
  • Concrete ask: the one thing that would unblock (env/credential access, captured artifact, permission for temporary instrumentation)

LSP-Powered Root Cause Tracing

Use LSP to trace root causes through the codebase:

  • Go to Definition — follow the call chain to where the value is actually set
  • Find References — find all callers of a suspect function (blast radius)
  • Go to Type Definition — check if the type allows the failing value
  • Hover — check types and signatures at the failure site

Don't guess where a value comes from — trace it with LSP. Don't grep for a function name — use Find References to get every caller with type info.

The Four Phases

Phase 1: Root Cause Investigation
  1. Understand — expected vs actual, when did it start?
  2. Git History — git log --oneline -20 -- <files>, git blame, git diff BASE..HEAD
  3. Compounded knowledge — if docs/solutions/debugging/ exists, check for a prior write-up matching this symptom before starting fresh investigation
  4. LOG FIRST — collect error logs, stack traces, run failing commands. The error text is the highest-density evidence you will ever get; acting first destroys or masks it.
  5. Before-capture — persist step 4's capture while the bug is reproducing: save the red output (log excerpt, command output, screenshot) to a durable artifact BEFORE writing the fix. It is the "before" half of the evidence pair; the fixed state alone proves less, and the failing state is cheapest to capture while it is red.
  6. Feedback Loop — build repro signal (construction ladder above). No loop → fail closed.
  7. Variant Scan — identify which variant dimensions must keep working (locale, config, env, platform, data shape, concurrency). A fix verified on one variant routinely breaks a sibling variant.

Repro Minimisation: After reproducing the bug, shrink to the smallest scenario that still goes red before forming hypotheses. Cut inputs, callers, config, and environment one at a time. Re-run after each cut. Every remaining element is load-bearing — removing it should make the bug disappear.

Why: A minimal repro shrinks the hypothesis space. The fewer moving parts, the fewer places the bug could hide.

Phase 2: Pattern Analysis
  1. Read the code around the failure — not just the failing line, the surrounding logic
  2. Check for recent changes — git diff the files involved
  3. Look for similar patterns — grep for the same anti-pattern elsewhere
  4. Identify the mechanism — not "what's wrong" but "how does the wrong thing happen"
Phase 3: Hypothesis and Testing

Generate 3-5 ranked hypotheses (H1, H2, ...) with 0-100 confidence BEFORE testing any of them — fewer than 3 means you anchored. Rank by explanatory power — which hypothesis explains the most symptoms with the fewest assumptions. Testing the first plausible hypothesis anchors you; generating multiple first prevents anchoring bias and surfaces connections between hypotheses. Proceed to fix only when one reaches 80+.

Hypothesis Quality Criteria:

  • States a specific mechanism ("X returns null because Y is not set when Z")
  • Predicts a specific test outcome ("if I set Y, X returns the correct value")
  • Is falsifiable ("if Y is already set, this hypothesis is wrong")
  • Explains ALL observed symptoms, not just the primary one

Instrumentation preference: Each probe must map to a specific prediction. Change one variable at a time. Tool preference:

  1. Debugger / REPL inspection if the env supports it — one breakpoint beats ten logs.
  2. Targeted logs at the boundaries that distinguish hypotheses.
  3. Never "log everything and grep". Tag every debug log with a unique prefix (e.g. [DEBUG-a4f2]) so cleanup is a single grep.

Performance branch: For performance regressions, logs are usually wrong. Establish a baseline measurement first (timing harness, performance.now, profiler, query plan), then bisect. Measure first, fix second.

Hypothesis Confidence Scoring:

ScoreMeaning
90-100Verified: traced with LSP, reproduces the bug, fix resolves it
80-89Strong: consistent with all evidence, mechanism is clear
60-79Plausible: fits some evidence but gaps remain — investigate more
<60Speculative: do not act — gather more evidence

When to Restart Investigation: If 3 hypotheses fail, you're pattern-matching, not investigating. Re-read the loop output. Re-trace with LSP. Consider you're looking at the wrong layer. If two or more fixes that share one premise failed the same gate, write the premise down — the premise, not the next fix variant, is now the suspect. Ask what the premise predicts beyond that gate: where the failure should concentrate, which variant should stay green. Test the prediction with the repro loop before the next fix; a prediction the observation contradicts disposes of the premise, not just the fix.

Show full SKILL.md (572 more words)Show less
Phase 4: Implementation
  1. RED — failing regression test reproducing the bug (must fail before fix)
  2. GREEN — minimal fix (smallest diff, no hardcoding)
  3. Blast Radius Scan — search same file for identical anti-patterns, adjacent files for same signature
  4. Verify — regression test passes + relevant suite passes
  5. Prevention — recommend lint rule, test, type guard, or monitoring. Recommend after the fix is in, not before — you have more information now than when you started.
  6. Cleanup — delete throwaway harnesses/prototypes, or move them to a clearly-marked debug location. Grep-remove all [DEBUG-...] tagged instrumentation (must return nothing). Confirm the repro loop no longer fires.

Scenario Playbooks

Read references/root-cause-playbooks.md for scenario-specific guidance:

  • State machine bugs (dead states, missing transitions)
  • Race conditions (timing-dependent failures)
  • Data corruption (cascading from wrong input)
  • Performance degradation (regression after change)
  • Integration failures (contract mismatch between services)

Debug Attempt Tracking

Track failed hypotheses: [DEBUG-N]: {what was tried} → {result}. After 3 failed hypotheses, set NEEDS_EXTERNAL_RESEARCH: true. After research files provided and still stuck, return BLOCKED.

Causal Chain Gate

Do not propose a fix until you can explain the full causal chain from trigger to symptom with no gaps. "Somehow X leads to Y" is a gap, not an explanation.

Predictions for uncertain links: When the causal chain has an uncertain link, form a prediction — something in a different code path that must also be true if your hypothesis is correct. If the prediction is wrong but the fix "works," you found a symptom fix, not the root cause.

Rationalization Table

ExcuseReality
"Emergency, no time for process"Systematic debugging is FASTER than guess-and-check thrashing (15-30 min vs 2-3 hours)
"I'll just try changing X and see"Random changes destroy evidence. Form a hypothesis first.
"Quick fix for now, investigate later""Later" never comes. Ship the real fix now.
"This should work" (without prediction)If you can't predict what will happen, you don't understand the bug.
"It worked before"Something changed. Find what. git bisect or git log --oneline -20 -- <files>.
"The tests pass so it's fixed"Tests can pass for the wrong reason. Verify the test actually exercises the bug path.
"I'm confident this is the cause"Confidence without a prediction is a feeling, not evidence.
"Let me just add a try/catch"Catching the error hides the bug. Find the root cause first.

Red Flags — STOP and Reconsider

  • You're about to make a change without a hypothesis
  • You're about to add a try/catch to suppress an error
  • You're about to hardcode a value to make a test pass
  • You've tried 3 fixes and none worked — you're pattern-matching, not debugging
  • You're considering skipping the feedback loop because "the bug is obvious"
  • You're about to mark FIXED without a regression test that was RED first
  • You're considering weakening an assertion to make the test pass
  • You're about to revert a fix and "try something else" without understanding why the fix failed

Pressure Testing

The gates in this skill must hold under pressure — deadline, complexity, "obvious bug" overconfidence. Before trusting a debug cycle:

  • Would this gate hold if the user said "just fix it now"?
  • Would this gate hold if the bug seemed obvious? The feedback loop gate exists precisely because "obvious" bugs are often wrong diagnoses.
  • Would this gate hold at 3am with no sleep? Rationalization tables exist because tired engineers skip process.

Pressure is exactly when these gates pay for themselves — a gate that can be talked out of under pressure was never protecting you.

© romiluz13, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in plugins/cc10x/skills/debugging of romiluz13/cc10x.

  • SKILL.md
  • evals/pressure-test-1-outage.md
  • evals/pressure-test-2-sunk-cost.md
  • evals/pressure-test-3-authority.md
  • references/investigation-hygiene.md
  • references/root-cause-playbooks.md

Open the folder on GitHubat commit 891acf0

Compare with similar skills

Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debugging this skillromiluz13/cc10x164—~3kAutomated safety check: NotesMIT
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT
Systematic DebuggingChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNone

Similar skills

  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Systematic Debugging

    ChrisWiles/claude-code-showcase

    Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

    6.1k GitHub starsUsed in 3 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    102k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed

More from romiluz13/cc10x

All 22 skills in this repo
  • Diff Driven Docs

    romiluz13/cc10x

    A skill your agent uses when a BUILD phase completes, a commit is staged, or a PR is about to be created, and the diff has not yet been reflected in documentation.

    164 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check: notes
  • Cc10x Guide

    romiluz13/cc10x

    Answers questions about cc10x itself — what it is, how to install and configure it, how the router, workflows, memory, and hooks operate, and how to troubleshoot.

    164 GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • Cc10x Router

    romiluz13/cc10x

    THE ONLY ENTRY POINT FOR CC10X. An agent skill from romiluz13/cc10x.

    164 GitHub stars~17k tokensUpdated 7 days ago
    Auto-check passed
  • Codebase Design

    romiluz13/cc10x

    Canonical deep-module vocabulary (module, interface, depth, seam, adapter, leverage, locality) for designing a module's shape — a lot of behaviour behind a small interface at a clean seam, testable…

    164 GitHub stars~1.9k tokensUpdated 7 days ago
    Auto-check passed
  • MCP CLI

    romiluz13/cc10x

    A skill your agent uses when you need a one-off MCP server capability during research or debugging without permanently mounting it as a context-polluting integration.

    164 GitHub stars~739 tokensUpdated 7 days ago
    Auto-check: notes
  • A skill your agent uses when a git merge or rebase reports conflicts and the operation is in progress.

    164 GitHub stars~709 tokensUpdated 7 days ago
    Auto-check: notes

Categories

Questions about Debugging

What does Debugging do?

Debugging discipline: feedback loop FIRST, root cause before fix, blast radius after fix. Debugging is an agent skill from romiluz13/cc10x. Debugging discipline: feedback loop FIRST, root cause before fix, blast radius after fix.

When should I use Debugging?

Debugging fits situations like: tasks that involve Debugging; tasks that involve Root cause analysis.

How do I install Debugging in Claude Code?

Run `npx skills add romiluz13/cc10x --skill debugging -a claude-code`. Or copy the skill folder (plugins/cc10x/skills/debugging in romiluz13/cc10x) into .claude/skills/debugging in your project. Claude Code loads it when a task matches its description.

How do I install Debugging in Codex?

Run `npx skills add romiluz13/cc10x --skill debugging -a codex`. Or copy the skill folder (plugins/cc10x/skills/debugging in romiluz13/cc10x) into .agents/skills/debugging in your project. Codex loads it when a task matches its description.

Can I use Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add romiluz13/cc10x --skill debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging, .gemini/skills/debugging, .github/skills/debugging and .opencode/skills/debugging in your project.

What does Debugging need to run?

Going by SKILL.md and its folder, Debugging needs the command-line tools its instructions call (git). Its frontmatter pre-approves these tools: Read, Edit, Bash, Grep, Glob, LSP.

Does Debugging access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Debugging safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Debugging use?

Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debugging use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Debugging?

Skills that share tags, products or a category with Debugging: OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars), Root Cause Debugging (garrytan/gstack, 136k stars) and Graph-Based Bug Tracing (tirth8205/code-review-graph, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debugging?

romiluz13 (a GitHub user) maintains it in romiluz13/cc10x, which has 164 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on September 30, 2026.

Source: romiluz13/cc10x on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.