Agent skill

Structured Debugging

by QwenLM in QwenLM/qwen-code

Hypothesis-driven debugging methodology for hard bugs. An agent skill from QwenLM/qwen-code.

Apache-2.0Auto-check passedDevelopment

Install Structured Debugging

skills CLI
$ npx skills add QwenLM/qwen-code --skill structured-debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install QwenLM/qwen-code structured-debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/structured-debugging .claude/skills/structured-debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
structured-debugging
GitHub stars
28k
Token cost
~2k tokens
SKILL.md length
1,161 words
Files
2
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Hypothesis-driven debugging methodology for hard bugs. An agent skill from QwenLM/qwen-code.

  • Works in 6 steps: Hypothesize → Design Instrumentation → Verify Data Collection → …
  • Youre investigating non-trivial bugs
  • SKILL.md covers The Cycle, Failure Modes to Avoid, Persistent State: A Special… and When to Exit the Cycle, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Structured Debugging is an agent skill from QwenLM/qwen-code. Hypothesis-driven debugging methodology for hard bugs. Use this skill whenever you're investigating non-trivial bugs, unexpected behavior, flaky tests, or tracing issues through complex systems. Activate proactively when debugging requires more than a quick glance — especially when the first attempt at a fix didn't work, when behavior seems "impossible", or when you're tempted to blame an external system (model, API, library) without evidence.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `examples/headless-bg-agent-empty-stdout.md`).

It sits in Development, covering Debugging and Failing and flaky tests. The repository describes itself as: An open-source AI coding agent that lives in your terminal. The licence is Apache-2.0.

When your agent uses it

  • Youre investigating non-trivial bugs
  • Unexpected behavior
  • Tracing issues through complex systems

Example prompts

  • “t work, when behavior seems”
  • “, or when you”
  • “/structured-debugging”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Hypothesize
  2. Design Instrumentation
  3. Verify Data Collection
  4. Run and Observe
  5. Document Findings
  6. Iterate

What it can do on your machine

Read from SKILL.md and the folder at commit 4970bfa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Structured Debugging loads about 2k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 1,161 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from QwenLM/qwen-code at commit 4970bfa, republished under its Apache-2.0 licence (© QwenLM). 1,161 words, ~2,017 tokens.

Download SKILL.mdSave it as .claude/skills/structured-debugging/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
structured-debugging
description
Hypothesis-driven debugging methodology for hard bugs. Use this skill whenever you're investigating non-trivial bugs, unexpected behavior, flaky tests, or tracing issues through complex systems. Activate proactively when debugging requires more than a quick glance — especially when the first attempt at a fix didn't work, when behavior seems "impossible", or when you're tempted to blame an external system (model, API, library) without evidence.

Structured Debugging

When debugging hard issues, the natural instinct is to form a theory and immediately apply a fix. This fails more often than it works. The fix addresses the wrong cause, adds complexity, creates false confidence, and obscures the real issue. Worse, after several failed attempts you lose track of what's been tried and start guessing randomly.

This methodology replaces guessing with a disciplined cycle that converges on the root cause. Each iteration narrows the search space. It's slower per attempt but dramatically faster overall because you stop wasting runs on wrong theories.

The Cycle

1. Hypothesize

Before touching code, write down what you think is happening and why. Be specific about the expected state at each step in the execution path.

Bad: "Something is wrong with the wait loop." Good: "The leader hangs because hasActiveTeammates() returns true after all agents have reported completed, likely because terminal status isn't being set on the agent object after the backend process exits."

For bugs you expect to take more than one round, create a side note file for the investigation in whichever location the project uses for such notes.

Write your hypothesis there. This file persists across conversation turns and even across sessions — it's your investigation journal.

2. Design Instrumentation

Add targeted debug logs or assertions at the exact decision points that would confirm or reject your hypothesis. Think about what data you need to see.

Don't scatter console.log everywhere. Identify the 2-3 places where your hypothesis makes a testable prediction, and instrument those.

Prefer logging values (return codes, payload contents, stream types, message bodies, env state) over presence checks ("was this function called?", "was this branch taken?"). Code-path traces tell you what ran; data traces tell you what it ran on. Most non-trivial bugs are correct code processing wrong data.

Ask yourself: "If my hypothesis is correct, what will I see at point X? If it's wrong, what will I see instead?"

3. Verify Data Collection

Before running, confirm that your instrumentation output will actually be captured and accessible.

Common traps:

  • stderr discarded by 2>/dev/null in the test command
  • Process killed before flush (logs lost)
  • Logging to a file in a directory that doesn't exist
  • Output piped through something that truncates it
  • Looking at log files from a previous run, not the current one

A test run that produces no data is wasted.

4. Run and Observe

Execute the test. Read the actual output — every line of it. Don't assume what it says.

When the data contradicts your hypothesis, believe the data. Don't rationalize it away. The whole point of this step is to let reality override your theory.

5. Document Findings

Update the side note with:

  • What the data showed (quote specific log lines)
  • What was confirmed vs. disproved
  • Updated hypothesis for the next iteration

This is critical for not losing context across attempts. Hard bugs typically take 3-5 rounds. Without notes, you'll forget what you ruled out and waste runs re-checking things.

6. Iterate

Update the hypothesis based on the new evidence. Go back to step 2. Each round should narrow the search space.

If you're not making progress after 3 rounds, step back and question your assumptions. The bug might be in a layer you haven't considered.

Failure Modes to Avoid

These are the specific traps this methodology is designed to prevent. When you notice yourself drifting toward any of them, stop and return to the cycle.

Jumping to fixes without evidence

The most common failure. You have a plausible theory, so you "fix" it and run again. If the theory was wrong, you've added complexity, wasted a test run, and possibly introduced a new bug. The side note should always show "hypothesis verified by [specific data]" before any fix is applied.

Blaming external systems

"The model is hallucinating." "The API is flaky." "The library has a bug." These conclusions feel satisfying because they put the problem outside your control. They're also usually wrong.

Before blaming an external system, inspect what it actually received. A model that appears to hallucinate may be responding rationally to stale data you didn't know was there. An API that appears flaky may be receiving malformed requests. Look at the inputs, not just the outputs.

Show full SKILL.md (458 more words)Show less
Inspecting code paths but not data

You instrument the code and prove it executes correctly — the right functions are called, in the right order, with no errors. But the bug persists. Why?

Because the code can work perfectly while processing garbage input. A function that correctly reads an inbox, correctly delivers messages, and correctly formats output is still broken if the inbox contains stale messages from a previous run.

Always inspect the content flowing through the code, not just whether the code runs. Check payloads, message contents, file data, and database state.

Reframing the user's report instead of investigating it

When the user reports a symptom your own run doesn't reproduce, the contradiction is the evidence — the two environments differ in some way you haven't identified yet. The wrong move is to reframe their report ("they must be on a stale SHA", "they must be confused about what they saw", "must be a flake") so that your run becomes the ground truth. Once you do that, every later piece of evidence gets bent to defend the reframing, and the actual bug stays hidden.

The right move: catalogue what differs between their environment and yours (TTY vs pipe, terminal emulator, shell, locale, env vars, prior state, build artifacts) before forming any hypothesis. For ambiguous symptoms ("no output", "it's slow", "it's wrong") ask one disambiguating question first — e.g., "does does it hang or exit cleanly?" That prunes the hypothesis space before any test run.

Losing context across attempts

After several debugging rounds, you start forgetting what you already tried and what you ruled out. You re-check things, go in circles, or abandon a promising line of investigation because you lost track of where it was heading.

This is why the side note file exists. Update it after every run. When you start a new round, re-read it first.

Persistent State: A Special Category

Features that persist data across runs — caches, session recordings, message queues, temp files, and database rows often cause "impossible" bugs. The current run's behavior is contaminated by leftover state from previous runs.

When behavior seems irrational, always check:

  • Is there persistent state that carries across runs?
  • Was it cleared before this run?
  • Is the system responding to stale data rather than current data?

This is easy to miss because the code is correct — it's the data that's wrong.

When to Exit the Cycle

Apply the fix only when you can point to specific data from your instrumentation that confirms the root cause. Write in the side note:

Root cause: [specific mechanism]
Evidence: [specific log lines / data that confirm it]
Fix: [what you're changing and why it addresses the root cause]

Then apply the fix, remove instrumentation, and verify with a clean run.

Worked examples

  • examples/headless-bg-agent-empty-stdout.md — pipe-captured runs all passed; the user's TTY printed nothing. The contradiction was the bug. Illustrates reproduction contradiction is data and instrument data, not code paths.

© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .qwen/skills/structured-debugging of QwenLM/qwen-code.

  • SKILL.md
  • examples/headless-bg-agent-empty-stdout.md

Open the folder on GitHubat commit 4970bfa

Compare with similar skills

Structured Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Structured Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Structured Debugging this skillQwenLM/qwen-code28k—~2kAutomated safety check: PassApache-2.0
React Router Bug Fix Workflowremix-run/react-router57k—~1.3kAutomated safety check: PassMIT
Runtime DebugOpenHikmah/openhikmah-web181—~618Automated safety check: PassGPL-3.0
Root Cause Debuggingjsmastery-pro/skills1.4k—~1.8kAutomated safety check: NotesMIT
Superpowers Systematic Debuggingchristopherarter/superpowers-reasonix102—~2kAutomated safety check: PassMIT
Fix The Classjoetawil7/first-pass112—~1.5kAutomated safety check: PassMIT

Similar skills

  • React Router Bug Fix Workflow

    remix-run/react-router

    Fixes a React Router bug reported in a GitHub issue end to end: fetching the issue, validating the reproduction, writing a failing test and implementing the fix on a new branch.

    57k GitHub stars~1.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Runtime Debug

    OpenHikmah/openhikmah-web

    Debug and verification workflow for runtime-bundle and module-resolution regressions.

    181 GitHub stars~618 tokensUpdated today
    DevelopmentAuto-check passed
  • Root Cause Debugging

    jsmastery-pro/skills

    Runs a reproduce, localize, hypothesize, test, fix and verify loop to find a bug's root cause, applies the minimal fix and hands off a regression test.

    1.4k GitHub stars~1.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Superpowers Systematic Debugging

    christopherarter/superpowers-reasonix

    Any bug, failing or flaky test, or surprise behavior?. An agent skill from christopherarter/superpowers-reasonix.

    102 GitHub stars~2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Fix The Class

    joetawil7/first-pass

    Bug-fix routine that fixes the whole class of bug, not just the reported instance.

    112 GitHub stars~1.5k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Minimal Code Fix

    cobusgreyling/loop-engineering

    Makes the smallest code change that fixes one well-scoped problem, such as a CI failure, review comment or typo, without refactoring anything unrelated.

    11k GitHub starsUsed in 1 repo~345 tokens
    DevelopmentAuto-check: notes

More from QwenLM/qwen-code

All 41 skills in this repo
  • Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.

    28k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.

    28k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.

    28k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Questions about Structured Debugging

What does Structured Debugging do?

Hypothesis-driven debugging methodology for hard bugs. An agent skill from QwenLM/qwen-code. Structured Debugging is an agent skill from QwenLM/qwen-code. Hypothesis-driven debugging methodology for hard bugs.

When should I use Structured Debugging?

Structured Debugging fits situations like: youre investigating non-trivial bugs; unexpected behavior; tracing issues through complex systems.

How do I install Structured Debugging in Claude Code?

Run `npx skills add QwenLM/qwen-code --skill structured-debugging -a claude-code`. Or copy the skill folder (.qwen/skills/structured-debugging in QwenLM/qwen-code) into .claude/skills/structured-debugging in your project. Claude Code loads it when a task matches its description.

How do I install Structured Debugging in Codex?

Run `npx skills add QwenLM/qwen-code --skill structured-debugging -a codex`. Or copy the skill folder (.qwen/skills/structured-debugging in QwenLM/qwen-code) into .agents/skills/structured-debugging in your project. Codex loads it when a task matches its description.

Can I use Structured Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill structured-debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/structured-debugging, .gemini/skills/structured-debugging, .github/skills/structured-debugging and .opencode/skills/structured-debugging in your project.

What does Structured Debugging need to run?

SKILL.md names no scripts, command-line tools or credentials: Structured Debugging is instructions for the agent only.

Does Structured Debugging access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Structured Debugging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Structured Debugging use?

Structured Debugging is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Structured Debugging use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Structured Debugging?

Skills that share tags, products or a category with Structured Debugging: React Router Bug Fix Workflow (remix-run/react-router, 57k stars), Runtime Debug (OpenHikmah/openhikmah-web, 181 stars), Root Cause Debugging (jsmastery-pro/skills, 1.4k stars) and Superpowers Systematic Debugging (christopherarter/superpowers-reasonix, 102 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Structured Debugging?

QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,337 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 7, 2026.

Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.