Agent skill

Test Guided Bug Detector

by ArabelaTso in ArabelaTso/Skills-4-SE

Analyze failing tests to detect functional bugs in code. An agent skill from ArabelaTso/Skills-4-SE.

Apache-2.0Auto-check passedDevelopment

Install Test Guided Bug Detector

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill test-guided-bug-detector -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE test-guided-bug-detector --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-guided-bug-detector .claude/skills/test-guided-bug-detector && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-guided-bug-detector
GitHub stars
253
Token cost
~2.8k tokens
SKILL.md length
586 words
Files
4 (incl. references)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze failing tests to detect functional bugs in code. An agent skill from ArabelaTso/Skills-4-SE.

  • Works in 7 steps: Parse Test Failure → Understand Test Intent → Trace Execution Path → …
  • Debugging test failures
  • SKILL.md covers Overview, Bug Detection Workflow, Analysis Process and Common Bug Patterns, plus 9 more sections
  • Calls pytest

What it does

Test Guided Bug Detector is an agent skill from ArabelaTso/Skills-4-SE. Analyze failing tests to detect functional bugs in code. Takes repository and failing test output as input, analyzes execution behavior, assertions, and stack traces to identify suspicious code regions and root causes. Use when debugging test failures, investigating regression bugs, or understanding why tests fail. Explains the bug mechanism, identifies affected code, and suggests fixes based on test expectations vs actual behavior.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/analysis_strategies.md`, `references/bug_patterns.md` and `references/failure_types.md`).

It sits in Development, covering Failing and flaky tests, Debugging and Root cause analysis. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • Debugging test failures
  • Investigating regression bugs
  • Understanding why tests fail

Example prompts

  • “/test-guided-bug-detector”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Parse Test Failure
  2. Understand Test Intent
  3. Trace Execution Path
  4. Identify Discrepancy
  5. Analyze Suspicious Code
  6. Explain Bug Mechanism
  7. Suggest Fix

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Guided Bug Detector loads about 2.8k tokens when it runs, and up to ~8.3k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 586 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 586 words, ~2,755 tokens.

Download SKILL.mdSave it as .claude/skills/test-guided-bug-detector/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
test-guided-bug-detector
description
Analyze failing tests to detect functional bugs in code. Takes repository and failing test output as input, analyzes execution behavior, assertions, and stack traces to identify suspicious code regions and root causes. Use when debugging test failures, investigating regression bugs, or understanding why tests fail. Explains the bug mechanism, identifies affected code, and suggests fixes based on test expectations vs actual behavior.

Test-Guided Bug Detector

Analyze failing tests to detect and explain functional bugs in code.

Overview

When tests fail, they provide valuable clues about bugs in the code. This skill analyzes:

  1. Test failure output - Error messages, stack traces, assertion failures
  2. Test expectations - What the test expects to happen
  3. Actual behavior - What actually happened
  4. Code execution path - Which code was executed
  5. Suspicious patterns - Common bug patterns that match the failure

The goal is to identify the root cause bug and explain why the test exposes it.

Bug Detection Workflow

Failing Test Output
    ↓
Parse Failure Information
    ↓
Identify Test Expectations
    ↓
Trace Execution Path
    ↓
Analyze Discrepancy
    ↓
Identify Suspicious Code
    ↓
Explain Bug Mechanism
    ↓
Suggest Fix

Analysis Process

Step 1: Parse Test Failure

Extract key information from test output:

What to extract:

  • Test name and location
  • Failure type (assertion, exception, timeout, etc.)
  • Expected vs actual values
  • Stack trace
  • Error messages

Example:

FAILED tests/test_calculator.py::test_divide - AssertionError: assert 0 == 5
Expected: 5
Actual: 0

Stack trace:
  File "tests/test_calculator.py", line 15, in test_divide
    assert divide(10, 2) == 5
  File "src/calculator.py", line 8, in divide
    return a // b
Step 2: Understand Test Intent

Determine what the test is trying to verify:

Questions to answer:

  • What functionality is being tested?
  • What are the inputs?
  • What is the expected output?
  • What properties should hold?

Example:

python
def test_divide():
    # Intent: Verify division returns correct result
    result = divide(10, 2)
    assert result == 5  # Expects 10 / 2 = 5
Step 3: Trace Execution Path

Follow the code path from test to failure:

Trace elements:

  • Function calls in stack trace
  • Control flow decisions
  • Data transformations
  • Return values

Example trace:

test_divide()
  → divide(10, 2)
    → return a // b  (integer division)
    → returns 5
  → assert 5 == 5  ✓ Should pass!
Step 4: Identify Discrepancy

Find where expected and actual diverge:

Common discrepancies:

  • Wrong operator (// vs /)
  • Off-by-one errors
  • Null/None handling
  • Type mismatches
  • Logic errors

Example:

python
# Expected: 10 / 2 = 5.0
# Actual: 10 // 2 = 5 (but test got 0?)
# Discrepancy: Something else is wrong!
Step 5: Analyze Suspicious Code

Examine code for bug patterns:

Bug patterns to check:

  • Uninitialized variables
  • Wrong operators
  • Missing return statements
  • Incorrect conditions
  • Edge case handling

Example analysis:

python
def divide(a, b):
    result = 0  # BUG: Initialized but never updated!
    return a // b  # This line is unreachable? No, wait...
    # Actually, this returns correctly, but...
Step 6: Explain Bug Mechanism

Describe how the bug causes the failure:

Explanation structure:

  1. What the code does
  2. What it should do
  3. Why there's a mismatch
  4. How the test exposes it
Step 7: Suggest Fix

Propose concrete fix with explanation:

Fix components:

  • Code change
  • Why it fixes the bug
  • How to verify the fix

Common Bug Patterns

For detailed bug patterns and detection strategies, see references/bug_patterns.md.

Categories include:

  • Logic errors (wrong operators, conditions)
  • State management (uninitialized, stale state)
  • Boundary conditions (off-by-one, edge cases)
  • Type errors (implicit conversions, null handling)
  • Concurrency bugs (race conditions, deadlocks)
Show full SKILL.md (245 more words)Show less

Failure Type Analysis

For analyzing different types of test failures, see references/failure_types.md.

Failure types:

  • Assertion failures
  • Exceptions and errors
  • Timeouts
  • Unexpected behavior
  • Flaky tests

Example Analysis

Input: Failing test

python
# Test file: tests/test_list_utils.py
def test_remove_duplicates():
    input_list = [1, 2, 2, 3, 3, 3, 4]
    result = remove_duplicates(input_list)
    assert result == [1, 2, 3, 4]
    assert input_list == [1, 2, 2, 3, 3, 3, 4]  # Original unchanged

# Test output:
# FAILED - AssertionError: assert [1, 2, 3, 4] == [1, 2, 2, 3, 3, 3, 4]
# The second assertion failed!

# Implementation: src/list_utils.py
def remove_duplicates(lst):
    seen = set()
    i = 0
    while i < len(lst):
        if lst[i] in seen:
            lst.pop(i)  # BUG: Modifies input list!
        else:
            seen.add(lst[i])
            i += 1
    return lst

Output: Bug analysis

markdown
# Bug Analysis Report

## Test Failure Summary

**Test:** test_remove_duplicates
**Location:** tests/test_list_utils.py:2
**Failure Type:** Assertion failure
**Failed Assertion:** `assert input_list == [1, 2, 2, 3, 3, 3, 4]`

## Expected vs Actual

**Expected:** Original list unchanged: `[1, 2, 2, 3, 3, 3, 4]`
**Actual:** Original list modified: `[1, 2, 3, 4]`

## Root Cause

**Bug Location:** src/list_utils.py:7
**Bug Type:** Unintended side effect (input mutation)

**Problematic Code:**
```python
lst.pop(i)  # Modifies the input list directly

Bug Mechanism

  1. What happens: The function modifies the input list in-place using lst.pop(i)
  2. Why it's wrong: The test expects the original list to remain unchanged
  3. How test exposes it: Second assertion checks that input_list is unmodified
  4. Why it fails: Since Python passes lists by reference, modifications to lst affect the original input_list

Execution Trace

test_remove_duplicates()
  input_list = [1, 2, 2, 3, 3, 3, 4]
  ↓
  remove_duplicates(input_list)  # lst points to same list as input_list
    i=0: lst[0]=1, not in seen, add to seen, i=1
    i=1: lst[1]=2, not in seen, add to seen, i=2
    i=2: lst[2]=2, in seen, lst.pop(2)  # Removes from input_list!
    # Now lst = input_list = [1, 2, 3, 3, 3, 4]
    i=2: lst[2]=3, not in seen, add to seen, i=3
    i=3: lst[3]=3, in seen, lst.pop(3)  # Removes from input_list!
    # Now lst = input_list = [1, 2, 3, 3, 4]
    i=3: lst[3]=3, in seen, lst.pop(3)  # Removes from input_list!
    # Now lst = input_list = [1, 2, 3, 4]
    i=3: lst[3]=4, not in seen, add to seen, i=4
    return lst  # Returns [1, 2, 3, 4]
  ↓
  result = [1, 2, 3, 4]  ✓ First assertion passes
  input_list = [1, 2, 3, 4]  ✗ Second assertion fails!

Suspicious Code Regions

Primary Suspect: src/list_utils.py:7
python
lst.pop(i)  # Direct mutation of input

Suspicion Level: HIGH Reason: Modifies input list, violating immutability expectation

Secondary Suspect: src/list_utils.py:11
python
return lst  # Returns reference to modified input

Suspicion Level: MEDIUM Reason: Returns same object as input, not a new list

Option 1: Create a copy (Recommended)

python
def remove_duplicates(lst):
    result = []  # Create new list
    seen = set()
    for item in lst:
        if item not in seen:
            seen.add(item)
            result.append(item)
    return result

Why this fixes it:

  • Creates new list instead of modifying input
  • Original list remains unchanged
  • Clearer intent

Option 2: Explicit copy

python
def remove_duplicates(lst):
    lst = lst.copy()  # Work on a copy
    seen = set()
    i = 0
    while i < len(lst):
        if lst[i] in seen:
            lst.pop(i)
        else:
            seen.add(lst[i])
            i += 1
    return lst

Why this fixes it:

  • lst.copy() creates a shallow copy
  • Modifications don't affect original
  • Preserves original algorithm structure

Verification

To verify the fix:

  1. Run the failing test: pytest tests/test_list_utils.py::test_remove_duplicates
  2. Both assertions should pass
  3. Add additional test for immutability:
python
def test_remove_duplicates_immutable():
    original = [1, 2, 2, 3]
    original_copy = original.copy()
    result = remove_duplicates(original)
    assert original == original_copy  # Verify no mutation

This bug could affect:

  • Any code that assumes remove_duplicates doesn't modify input
  • Functions that reuse the input list after calling remove_duplicates
  • Concurrent code where multiple threads access the same list

## Analysis Strategies

For detailed analysis strategies by language and framework, see [references/analysis_strategies.md](references/analysis_strategies.md).

Strategies include:
- Python (pytest, unittest)
- JavaScript (Jest, Mocha)
- Java (JUnit)
- C/C++ (Google Test)
- Go (testing package)

## Best Practices

1. **Start with the failure message** - It often points directly to the bug
2. **Understand test intent** - Know what should happen
3. **Trace execution carefully** - Follow the actual code path
4. **Look for common patterns** - Many bugs follow known patterns
5. **Consider edge cases** - Bugs often hide at boundaries
6. **Check assumptions** - Verify what the code assumes
7. **Explain clearly** - Make the bug mechanism understandable

## Red Flags

Watch for these suspicious patterns:

**High-priority red flags:**
- Uninitialized variables
- Missing return statements
- Wrong operators (==  vs =, // vs /)
- Off-by-one errors (< vs <=)
- Null/None without checks
- Mutable default arguments
- Side effects in pure functions

**Medium-priority warnings:**
- Complex conditionals
- Nested loops with breaks
- Exception swallowing
- Type conversions
- Global state access

## Report Template

```markdown
# Bug Analysis Report

## Test Failure Summary
- Test name and location
- Failure type
- Failed assertion/error

## Expected vs Actual
- What should happen
- What actually happened

## Root Cause
- Bug location (file:line)
- Bug type
- Problematic code snippet

## Bug Mechanism
- Step-by-step explanation
- Why it's wrong
- How test exposes it

## Execution Trace
- Detailed trace from test to failure
- Variable values at key points

## Suspicious Code Regions
- Primary suspects with evidence
- Secondary suspects

## Recommended Fix
- Proposed code change
- Explanation of why it fixes the bug
- How to verify

## Related Issues
- Other code that might be affected

Additional Resources

For detailed guidance:

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/test-guided-bug-detector of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • references/analysis_strategies.md
  • references/bug_patterns.md
  • references/failure_types.md

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Test Guided Bug Detector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Guided Bug Detector compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Guided Bug Detector this skillArabelaTso/Skills-4-SE253—~2.8kAutomated safety check: PassApache-2.0
Root Cause Debuggingjsmastery-pro/skills1.4k—~1.8kAutomated safety check: NotesMIT
Superpowers Systematic Debuggingchristopherarter/superpowers-reasonix102—~2kAutomated safety check: PassMIT
Minimal Code Fixcobusgreyling/loop-engineering11k1 repos~345Automated safety check: NotesMIT
Failure Diagnosis LoopTotoro-jam/battle-tested-patterns344—~319Automated safety check: PassMIT
Hypothesis-Driven DebuggingLichAmnesia/lich-skills234—~2.5kAutomated safety check: PassMIT

Similar skills

  • Root Cause Debugging

    jsmastery-pro/skills

    Runs a reproduce, localize, hypothesize, test, fix and verify loop to find a bug's root cause, applies the minimal fix and hands off a regression test.

    1.4k GitHub stars~1.8k tokensUpdated 2 mo ago
    DevelopmentAuto-check: notes
  • Superpowers Systematic Debugging

    christopherarter/superpowers-reasonix

    Any bug, failing or flaky test, or surprise behavior?. An agent skill from christopherarter/superpowers-reasonix.

    102 GitHub stars~2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Minimal Code Fix

    cobusgreyling/loop-engineering

    Makes the smallest code change that fixes one well-scoped problem, such as a CI failure, review comment or typo, without refactoring anything unrelated.

    11k GitHub starsUsed in 1 repo~345 tokens
    DevelopmentAuto-check: notes
  • Failure Diagnosis Loop

    Totoro-jam/battle-tested-patterns

    Walks the agent through a fixed loop for failing tests and build errors: reproduce, isolate, hypothesize, instrument, fix, verify, then add a regression test.

    344 GitHub stars~319 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Hypothesis-Driven Debugging

    LichAmnesia/lich-skills

    Replaces trial-and-error fixing with an observe, hypothesize, experiment and conclude loop kept in DEBUG.md, where no fix is allowed before evidence supports a cause.

    234 GitHub stars~2.5k tokensUpdated 4 mo ago
    DevelopmentAuto-check passed
  • Systematic Debugging

    cbrock84/headcount

    Finds the root cause of a bug, test failure, or unexpected behavior before proposing any fix.

    2k GitHub stars~657 tokensUpdated 21 days ago
    DevelopmentAuto-check passed

More from ArabelaTso/Skills-4-SE

All 150 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Test Guided Bug Detector

What does Test Guided Bug Detector do?

Analyze failing tests to detect functional bugs in code. An agent skill from ArabelaTso/Skills-4-SE. Test Guided Bug Detector is an agent skill from ArabelaTso/Skills-4-SE. Analyze failing tests to detect functional bugs in code.

When should I use Test Guided Bug Detector?

Test Guided Bug Detector fits situations like: debugging test failures; investigating regression bugs; understanding why tests fail.

How do I install Test Guided Bug Detector in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill test-guided-bug-detector -a claude-code`. Or copy the skill folder (skills/test-guided-bug-detector in ArabelaTso/Skills-4-SE) into .claude/skills/test-guided-bug-detector in your project. Claude Code loads it when a task matches its description.

How do I install Test Guided Bug Detector in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill test-guided-bug-detector -a codex`. Or copy the skill folder (skills/test-guided-bug-detector in ArabelaTso/Skills-4-SE) into .agents/skills/test-guided-bug-detector in your project. Codex loads it when a task matches its description.

Can I use Test Guided Bug Detector in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill test-guided-bug-detector -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-guided-bug-detector, .gemini/skills/test-guided-bug-detector, .github/skills/test-guided-bug-detector and .opencode/skills/test-guided-bug-detector in your project.

What does Test Guided Bug Detector need to run?

Going by SKILL.md and its folder, Test Guided Bug Detector needs the command-line tools its instructions call (pytest). Our summary lists: Python 3.

Does Test Guided Bug Detector access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Guided Bug Detector safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Guided Bug Detector use?

Test Guided Bug Detector is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Guided Bug Detector use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to Test Guided Bug Detector?

Skills that share tags, products or a category with Test Guided Bug Detector: Root Cause Debugging (jsmastery-pro/skills, 1.4k stars), Superpowers Systematic Debugging (christopherarter/superpowers-reasonix, 102 stars), Minimal Code Fix (cobusgreyling/loop-engineering, 11k stars) and Failure Diagnosis Loop (Totoro-jam/battle-tested-patterns, 344 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Guided Bug Detector?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.