Agent skill

Debugging

by WrongStack in WrongStack/WrongStack

A skill your agent uses when something is broken and the cause is unknown — a failing test, a crash, an error message, wrong output, a regression, a hang, or behaviour that differs between…

MITAuto-check passedDevelopment

Install Debugging

skills CLI
$ npx skills add WrongStack/WrongStack --skill debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install WrongStack/WrongStack debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/WrongStack/WrongStack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/core/skills/debugging .claude/skills/debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debugging
GitHub stars
370
Token cost
~2.3k tokens
SKILL.md length
1,184 words
Files
2
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when something is broken and the cause is unknown — a failing test, a crash, an error message, wrong output, a regression, a hang, or behaviour that differs between…

  • Works in 8 steps: Reproduce before fixing. Get a command,… → Read the whole error. The top of a stack… → Form a hypothesis that explains every… → …
  • Something is broken and the cause is unknown — a failing test
  • SKILL.md covers Overview, Rules, Workflow and Proof-grade reproduction, plus 6 more sections
  • Calls git

What it does

Debugging is an agent skill from WrongStack/WrongStack. Use this skill when something is broken and the cause is unknown — a failing test, a crash, an error message, wrong output, a regression, a hang, or behaviour that differs between environments — and it has to be found and fixed at the root. Also use it to build the failing proof for a suspected defect before fixing it, as in a proof-driven bug hunt. Triggers: user says "debug", "broken", "doesn't work", "failing", "error", "exception", "stack trace", "crash", "regression", "it used to work", "hangs", "why does"…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `SKILL.save.md`).

It sits in Development, covering Debugging, Failing and flaky tests and Root cause analysis. The repository describes itself as: An AI coding agent that reads your code, edits files, runs commands, and reasons through bugs — across a terminal REPL, a full-screen TUI, and a browser UI, while you keep your… The licence is MIT.

When your agent uses it

  • Something is broken and the cause is unknown — a failing test
  • An error message
  • Behaviour that differs between environments — and it has to be found and fixed at the root
  • Build the failing proof for a suspected defect before fixing it

Example prompts

  • “broken”
  • “t work”
  • “failing”
  • “/debugging”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Reproduce before fixing. Get a command, test, or input that shows the failure
  2. Read the whole error. The top of a stack trace is where the failure surfaced;
  3. Form a hypothesis that explains every symptom, then run the cheapest
  4. Change one thing at a time, and keep a short log of what was tried and what
  5. Fix the cause, not the symptom. Swallowing the error, widening a type, adding
  6. Prove the fix: the reproduction passes, a regression test fails without the
  7. Remove temporary instrumentation before finishing.
  8. After three disproven hypotheses, step back. Re-check the assumptions — right

What it can do on your machine

Read from SKILL.md and the folder at commit 57f6018. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debugging loads about 2.3k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 1,184 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from WrongStack/WrongStack at commit 57f6018, republished under its MIT licence (© WrongStack). 1,184 words, ~2,289 tokens.

Download SKILL.mdSave it as .claude/skills/debugging/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
debugging
description
Use this skill when something is broken and the cause is unknown — a failing test, a crash, an error message, wrong output, a regression, a hang, or behaviour that differs between environments — and it has to be found and fixed at the root. Also use it to build the failing proof for a suspected defect before fixing it, as in a proof-driven bug hunt. Triggers: user says "debug", "broken", "doesn't work", "failing", "error", "exception", "stack trace", "crash", "regression", "it used to work", "hangs", "why does", "root cause", "works on my machine", "reproduce", "prove it".
version
1.1.0
required-capabilities
filesystem.read
optional-capabilities
execution.shell, verification.run, code.inspect, version-control.manage

Debugging

Overview

Debugging is a search for the first point where reality diverges from expectation. It goes fast when every step narrows the search with evidence, and slowly when it guesses and patches symptoms. The deliverable is a fix at the cause, a test that fails without it, and an explanation that accounts for every observation.

This skill starts from an observed failure, or from a suspected defect that bug-hunter selected and that now has to be proven. Scanning code for defects nobody has reported yet is bug-hunter.

Rules

  1. Reproduce before fixing. Get a command, test, or input that shows the failure on demand. If that isn't possible yet, the first job is making it reproducible — logs, inputs, environment — not changing code. If it never becomes reproducible, production code stays untouched and the report says what was tried.
  2. Read the whole error. The top of a stack trace is where the failure surfaced; the cause is often further down: a Caused by, the first frame in project code, or the earliest error in the log.
  3. Form a hypothesis that explains every symptom, then run the cheapest experiment that could prove it wrong.
  4. Change one thing at a time, and keep a short log of what was tried and what it showed.
  5. Fix the cause, not the symptom. Swallowing the error, widening a type, adding a retry, or special-casing the failing input is only right when that really is the correct behaviour.
  6. Prove the fix: the reproduction passes, a regression test fails without the fix, and neighbouring tests still pass.
  7. Remove temporary instrumentation before finishing.
  8. After three disproven hypotheses, step back. Re-check the assumptions — right file, right build, right branch, right environment — or bisect.

Workflow

  1. Capture the exact error text, command, input, environment, and what changed recently (history, dependency bumps, configuration).
  2. Reproduce with the smallest command that fails — ideally one test.
  3. Localize:
    • follow the stack trace to the first project frame and read it;
    • trace a wrong value backwards to where it first became wrong;
    • with the codebase index, the codebase-incoming-calls and codebase-outgoing-calls tools show how the failing code is reached, and codebase-context finds code from a description of the behaviour;
    • if it used to work, bisect with the reproduction as the test;
    • if the search space is large, halve it: add a checkpoint in the middle, or disable half of the suspects.
  4. Explain the cause in one or two sentences that cover every symptom, including "why only sometimes" and "why only here".
  5. Fix at the cause with a minimal diff.
  6. Verify: the reproduction passes; the regression test goes red then green; covering tests (the codebase-targeted-test tool finds them) and the type checker pass.

Proof-grade reproduction

A reproduction is evidence only if it could not have failed for any other reason. Before trusting a red run:

  • It runs the production path. Call the real implementation; mock only the boundaries you don't own (network, clock, randomness, third-party services). A copied algorithm or a mock of the code under suspicion proves nothing.
  • It asserts the contract. The assertion encodes the expected behaviour and its basis — a documented contract, a caller's requirement, or an established test — not whatever the code happens to return.
  • It fails on that assertion. An import error, a setup failure, a missing dependency, or a timeout is blocked verification, not a red proof. Fix the harness until the failure message is about the defect.
  • It has a control. Pair it with a neighbouring valid input or an unaffected path that passes before and after the fix, so a broken harness cannot masquerade as the bug.
  • It is deterministic. Fake timers, seeded randomness, controlled async ordering, fixed fixtures. For anything timing-related, run it several times.
  • It runs the code you think it runs. Check whether the runner loads source or built output. In a monorepo a stale build of a dependency makes red and green compare the wrong code; rebuild the chain or point the runner at source.
  • It is recorded before any fix. Keep the working directory, command, exit code, and the meaningful failure lines; the same command must later go green with unchanged assertions and fixtures.

Keep throwaway proof scripts and fixtures in a scratch directory you own, and remove only that directory once the evidence is in the report.

Show full SKILL.md (459 more words)Show less
Shared checkouts

Other people or processes may be editing the same working tree. Record the starting revision and dirty paths first. If the affected source changes between the red run and the fix, re-read it and reproduce again. Never revert, stash, or check out over someone else's changes to manufacture a baseline — copy the unfixed implementation to an isolated location instead.

First moves by symptom

SymptomFirst moves
Exception with a stack traceFirst project frame; the values at that frame; where a null or undefined came from
Wrong output, no errorTrace the value back; check intermediate values; compare against a known-good input
It used to workHistory of the touched files; git bisect run <reproduction command>
Only in CI or on one machineDiff the environment: runtime and dependency versions, env vars, OS path case, line endings, timezone, locale, parallelism
IntermittentRaces, shared state, timers, test order, retries hiding errors; run it many times, alone, and shuffled
HangUnresolved promise, deadlock, missing timeout, waiting on stdin; take a stack dump or log progress at each boundary
Slower than beforeProfile before and after; don't guess
Build or type errorRead the full error chain; fix the value or the declaration at its source

Instrumentation

  • Prefer a focused test or a debugger over prints scattered across files.
  • Tag temporary logs with a unique marker so every one is found and removed, and log values with their types, not "got here".
  • Never add logging that writes secrets or personal data.

Report

text
## Root cause
The retry wrapper reuses the AbortSignal from the first attempt; once it times
out, every retry is aborted immediately, so the job fails after 30 s instead of
retrying for 2 min.

## Evidence
- `pnpm test retry.test.ts -t "retries after timeout"` failed with AbortError before the fix
- src/net/retry.ts:41 passes `options.signal` to every attempt

## Fix
src/net/retry.ts:41 — create a fresh timeout signal per attempt, combined with the caller's signal.

## Verification
- Regression test red before, green after
- net package tests and type check pass

Anti-patterns

  • Shotgun edits — changing several things until the error disappears.
  • "Fixed" without reproducing — the failure may simply not have occurred this run.
  • Catching and ignoring the exception.
  • Sleeps and retries for an intermittent failure whose cause is unknown.
  • Blaming the framework, compiler, or cache before ruling out the project's own code — and clearing caches only when evidence points there.
  • Naming a cause from the error message alone without reading the code.
  • Proof theatre — a "reproduction" that fails on setup, a mock, or an artificial throw, or that re-implements the logic instead of calling it.
  • Moving the goalposts — changing the proof's assertions or fixtures after the fix without re-running the changed proof against the unfixed code.

Before returning

  • Failure reproduced, or it is stated plainly why it couldn't be
  • Reproduction calls the production path, fails on the contract assertion, and has a passing control
  • Root cause stated, consistent with every symptom
  • Fix at the cause; no swallowed errors or special-casing
  • Regression test seen red, then green; related suites pass
  • Temporary instrumentation removed

Skills in scope

  • testing — for the regression test that pins the fix
  • bug-hunter — for scanning nearby code for the same defect pattern
  • verify-before-done — for the final proof before reporting
  • git-flow — for bisecting and for committing the fix

© WrongStack, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in packages/core/skills/debugging of WrongStack/WrongStack.

  • SKILL.md
  • SKILL.save.md

Open the folder on GitHubat commit 57f6018

Compare with similar skills

Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debugging this skillWrongStack/WrongStack370—~2.3kAutomated safety check: PassMIT
Systematic DebuggingChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNone
Debugging and Error Recoveryaddyosmani/agent-skills103k1 repos~2.6kAutomated safety check: PassMIT
Systematic Debugginged3dai/ed3d-plugins2503 repos~2.4kAutomated safety check: PassNone
Debugging And Error Recoveryabashev/vfs-s31066 repos~2.6kAutomated safety check: PassApache-2.0
Veomni DebugByteDance-Seed/VeOmni2.2k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • Systematic Debugging

    ChrisWiles/claude-code-showcase

    Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

    6.1k GitHub starsUsed in 3 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    103k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed
  • Systematic Debugging

    ed3dai/ed3d-plugins

    A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes - four-phase framework (root cause investigation, pattern analysis, hypothesis…

    250 GitHub starsUsed in 3 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Guides systematic root-cause debugging. An agent skill from abashev/vfs-s3.

    106 GitHub starsUsed in 6 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Veomni Debug

    ByteDance-Seed/VeOmni

    A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

    2.2k GitHub stars~2.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Root Cause Debugging

    jsmastery-pro/skills

    Runs a reproduce, localize, hypothesize, test, fix and verify loop to find a bug's root cause, applies the minimal fix and hands off a regression test.

    1.4k GitHub stars~1.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes

More from WrongStack/WrongStack

All 38 skills in this repo
  • Design Craft

    WrongStack/WrongStack

    Design or substantially improve user-facing interfaces with a product-specific visual direction, content hierarchy, and rendered critique.

    370 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Design Critique

    WrongStack/WrongStack

    A skill your agent uses to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states…

    370 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Mailbox Bridge

    WrongStack/WrongStack

    A skill your agent uses when external coding agents (Claude Code, Aider, custom scripts) need to participate in the project's shared WrongStack mailbox, or when a user asks to "expose the mailbox"…

    370 GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Multi Agent

    WrongStack/WrongStack

    A skill your agent uses whenever work can be split across multiple AI agents running in parallel, or when orchestrating leader/worker patterns in WrongStack.

    370 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Web Platform Baseline

    WrongStack/WrongStack

    Use this skill before asserting that a CSS, HTML or accessibility capability is available, unavailable, or the right tool — it carries dated, refreshable platform facts and refuses to let stale…

    370 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Wrongstack Mailbox

    WrongStack/WrongStack

    A skill your agent uses when the user wants to communicate with WrongStack's shared project mailbox from outside WrongStack — read messages sent by WrongStack agents, send replies, broadcast to all…

    370 GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Categories

Questions about Debugging

What does Debugging do?

A skill your agent uses when something is broken and the cause is unknown — a failing test, a crash, an error message, wrong output, a regression, a hang, or behaviour that differs between…. Debugging is an agent skill from WrongStack/WrongStack. Use this skill when something is broken and the cause is unknown — a failing test, a crash, an error message, wrong output, a regression, a hang, or behaviour that differs between environments — and it has to be found and fixed at the root.

When should I use Debugging?

Debugging fits situations like: something is broken and the cause is unknown — a failing test; an error message; behaviour that differs between environments — and it has to be found and fixed at the root; build the failing proof for a suspected defect before fixing it.

How do I install Debugging in Claude Code?

Run `npx skills add WrongStack/WrongStack --skill debugging -a claude-code`. Or copy the skill folder (packages/core/skills/debugging in WrongStack/WrongStack) into .claude/skills/debugging in your project. Claude Code loads it when a task matches its description.

How do I install Debugging in Codex?

Run `npx skills add WrongStack/WrongStack --skill debugging -a codex`. Or copy the skill folder (packages/core/skills/debugging in WrongStack/WrongStack) into .agents/skills/debugging in your project. Codex loads it when a task matches its description.

Can I use Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add WrongStack/WrongStack --skill debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging, .gemini/skills/debugging, .github/skills/debugging and .opencode/skills/debugging in your project.

What does Debugging need to run?

Going by SKILL.md and its folder, Debugging needs the command-line tools its instructions call (git).

Does Debugging access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Debugging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debugging use?

Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debugging use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debugging?

Skills that share tags, products or a category with Debugging: Systematic Debugging (ChrisWiles/claude-code-showcase, 6.1k stars), Debugging and Error Recovery (addyosmani/agent-skills, 103k stars), Systematic Debugging (ed3dai/ed3d-plugins, 250 stars) and Debugging And Error Recovery (abashev/vfs-s3, 106 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debugging?

WrongStack (a GitHub organization) maintains it in WrongStack/WrongStack, which has 370 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on October 7, 2026.

Source: WrongStack/WrongStack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.