Official agent skill

Diagnosing Bugs

by getsentry in getsentry/sentry-dart

A discipline for hard bugs, flaky tests, CI hangs, and performance regressions in this SDK.

OfficialMITAuto-check passedTesting & QA

Install Diagnosing Bugs

skills CLI
$ npx skills add getsentry/sentry-dart --skill diagnosing-bugs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install getsentry/sentry-dart diagnosing-bugs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/getsentry/sentry-dart.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/diagnosing-bugs .claude/skills/diagnosing-bugs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
diagnosing-bugs
GitHub stars
873
Token cost
~1.7k tokens
SKILL.md length
1,014 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

A discipline for hard bugs, flaky tests, CI hangs, and performance regressions in this SDK.

  • Works in 6 steps: Build a feedback loop → Reproduce and minimise → Hypothesise → …
  • The user says diagnose
  • SKILL.md covers Phase 1 — Build a feedback loop, Phase 2 — Reproduce and minimise, Phase 3 — Hypothesise and Phase 4 — Instrument, plus 2 more sections
  • Calls flutter, dart and git

What it does

Diagnosing Bugs is an agent skill from getsentry/sentry-dart, published by the product's own GitHub organization. A discipline for hard bugs, flaky tests, CI hangs, and performance regressions in this SDK. Use when the user says "diagnose" or "debug this", reports something broken/throwing/failing/hanging/slow, or a test is flaky. Builds a tight, red-capable feedback loop before hypothesizing.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests and Cross-platform mobile apps. It works with Sentry and Flutter. The repository describes itself as: Sentry SDK for Dart and Flutter. The licence is MIT.

When your agent uses it

  • The user says diagnose
  • Reports something broken/throwing/failing/hanging/slow
  • A test is flaky

Example prompts

  • “diagnose”
  • “debug this”
  • “/diagnosing-bugs”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Build a feedback loop
  2. Reproduce and minimise
  3. Hypothesise
  4. Instrument
  5. Fix and regression test
  6. Cleanup and post-mortem

What it can do on your machine

Read from SKILL.md and the folder at commit 96d6367. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • flutter
    • dart
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Diagnosing Bugs loads about 1.7k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 1,014 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from getsentry/sentry-dart at commit 96d6367, republished under its MIT licence (© getsentry). 1,014 words, ~1,714 tokens.

Download SKILL.mdSave it as .claude/skills/diagnosing-bugs/SKILL.md (or your agent's skills folder).
name
diagnosing-bugs
description
A discipline for hard bugs, flaky tests, CI hangs, and performance regressions in this SDK. Use when the user says "diagnose" or "debug this", reports something broken/throwing/failing/hanging/slow, or a test is flaky. Builds a tight, red-capable feedback loop before hypothesizing.

A discipline for hard bugs. Skip a phase only when you can say why.

The whole skill is Phase 1: get a loop that goes red on this bug. Everything after is mechanical once you have it. If you catch yourself reading code to form a theory before that loop exists, stop — jumping to a hypothesis without a red-capable loop is the exact failure this prevents.

For trivial bugs with an obvious fix and an obvious test, skip this and just write the failing test — this skill is for the bugs that resist.

Phase 1 — Build a feedback loop

A tight loop is fast, deterministic, and goes red on this bug. Build one and the bug is 90% solved; bisection, hypothesis-testing, and instrumentation all just consume it. Be aggressive and creative here — spend disproportionate effort.

Ways to construct one — try roughly in this order
  1. Failing test at the seam — dart test path/to/test.dart or flutter test, unit or widget, at whatever seam reaches the bug. Tightest loop there is.
  2. fakeAsync harness — for timer, microtask, and timing-dependent bugs (most flakes). Pin the clock via options.clock, elapse time deterministically.
  3. testWidgets harness — for Flutter UI bugs: pump the widget tree, drive the gesture, assert on the tree (e.g. a hit-test or navigation regression).
  4. Integration test — flutter test integration_test/... when the bug crosses the native (JNI/FFI) boundary, which cannot be faked (see packages/flutter/AGENTS.md).
  5. Replay a captured payload — save a real envelope / event / span / network payload to disk and replay it through the code path in isolation.
  6. Throwaway harness — a minimal void main() that exercises the bug path with one call and mocked deps.
  7. Stress/repetition loop — for non-deterministic bugs: run the trigger 100×, parallelise, narrow timing windows, inject delays. The goal is a higher reproduction rate, not a clean repro.
  8. Differential loop — run the same input through two versions or two configs (e.g. before/after a dependency bump) and diff the output. For regressions that "appeared after X".
  9. Bisection harness — if it appeared between two commits, script the red/green check and git bisect run it.
Tighten the loop

Once you have a loop, treat it as a product and tighten it:

  • Faster — narrow the test scope, skip unrelated init.
  • Sharper — assert the exact symptom the user described, not "didn't crash".
  • More deterministic — pin the clock (options.clock), use fakeAsync, seed any randomness, avoid real network/filesystem (per test-guidelines). A 2-second deterministic loop beats a 30-second flaky one.

For non-deterministic bugs the goal is a high enough reproduction rate to debug against — a 50% flake is debuggable, 1% is not. Keep raising the rate.

Completion criterion

Phase 1 is done when you can name one command — a test invocation or script path — that you have already run at least once (paste the invocation and its output), and that is:

  • Red-capable — drives the actual bug path and asserts the user's exact symptom, so it goes red now and green once fixed
  • Deterministic — same verdict every run (flaky bugs: a pinned, high reproduction rate)
  • Fast — seconds, not minutes

No red-capable command, no Phase 2.

When you genuinely cannot build a loop

Say so explicitly, list what you tried, and ask the user for: access to an environment that reproduces it, a captured artifact (envelope dump, log, CI run, screen recording with timestamps), or permission for temporary instrumentation. Do not hypothesise without a loop.

Phase 2 — Reproduce and minimise

Run the loop, watch it go red. Confirm it produces the failure mode the user described — not a nearby one (wrong bug = wrong fix). Then shrink the repro to the smallest scenario that still goes red: cut inputs, callers, config, and steps one at a time, re-running after each cut. Done when every remaining element is load-bearing — removing any one turns it green. The minimal repro shrinks the hypothesis space and becomes the regression test.

Show full SKILL.md (374 more words)Show less

Phase 3 — Hypothesise

Generate 3–5 ranked, falsifiable hypotheses before testing any — a single hypothesis anchors you on the first plausible idea. Each must state its prediction: "If X is the cause, then changing Y makes the bug vanish." If you can't state the prediction, it's a vibe — sharpen or discard it. Show the ranked list to the user before testing; they often re-rank it instantly ("we just changed #3"). Don't block on it if they're away.

Phase 4 — Instrument

Each probe maps to one prediction from Phase 3. Change one variable at a time.

  • Prefer a debugger / breakpoint over logs where the env supports it. One breakpoint beats ten logs.
  • Otherwise add targeted logs at the boundaries that distinguish hypotheses — never "log everything and grep".
  • Tag every debug log with a unique prefix like [DEBUG-a4f2] so cleanup is one grep.
  • Performance regressions: logs are the wrong tool. Establish a baseline measurement (Stopwatch, a timing harness, DevTools, the observatory), then bisect. Measure first, fix second.

Phase 5 — Fix and regression test

Write the regression test before the fix — but only at a correct seam, one that exercises the real bug pattern as it occurs at the call site (per test-guidelines). A too-shallow seam (a unit test that can't replicate the chain that triggered it) gives false confidence. If no correct seam exists, that is itself the finding — note it; the architecture is preventing the bug from being locked down.

With a correct seam: turn the minimised repro into a failing test, watch it fail, apply the fix, watch it pass, then re-run the Phase 1 loop against the original un-minimised scenario.

Phase 6 — Cleanup and post-mortem

Before declaring done:

  • Original repro no longer reproduces (re-run the Phase 1 loop)
  • Regression test passes (or the absence of a seam is documented)
  • All [DEBUG-…] instrumentation removed (grep the prefix)
  • Throwaway harnesses deleted
  • The correct hypothesis is stated in the commit / PR message, so the next debugger learns

Then ask: what would have prevented this bug? If the answer is architectural — no good test seam, tangled callers, a shallow module hiding the real bug behind it — hand off to the design-first skill with the specifics. Make that call after the fix is in, when you know the most.

© getsentry, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/diagnosing-bugs of getsentry/sentry-dart.

Open the folder on GitHubat commit 96d6367

Compare with similar skills

Diagnosing Bugs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Diagnosing Bugs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Diagnosing Bugs this skillgetsentry/sentry-dart873—~1.7kAutomated safety check: PassMIT
Diagnosing Bugsgetsentry/sentry-react-native1.8k—~1.4kAutomated safety check: PassMIT
Flutter TestingMADTeacher/mad-agents-skills110—~1.5kAutomated safety check: PassMIT
Testingevanca/flutter-ai-rules647—~1.1kAutomated safety check: PassMIT
Test Guidelinesgetsentry/sentry-react-native1.8k—~1.3kAutomated safety check: PassMIT
Frb Fix CIfzyzcjy/flutter_rust_bridge5.4k—~4.4kAutomated safety check: PassMIT

Similar skills

  • Diagnosing Bugs

    getsentry/sentry-react-native

    Official

    A discipline for hard bugs, flaky tests, CI hangs, native crashes, and performance regressions in this SDK.

    1.8k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Flutter Testing

    MADTeacher/mad-agents-skills

    Write, fix, review, debug, and validate Flutter tests for apps, packages, and plugins.

    110 GitHub stars~1.5k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Testing

    evanca/flutter-ai-rules

    A skill your agent uses when writing or reviewing Flutter/Dart tests (unit, widget, golden), fixing flaky tests, adding coverage, or choosing between unit and widget tests.

    647 GitHub stars~1.1k tokensUpdated 23 days ago
    Testing & QAAuto-check passed
  • Test Guidelines

    getsentry/sentry-react-native

    Official

    Enforce Sentry React Native SDK test conventions for naming, structure, mocking, and fixtures with Jest.

    1.8k GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Frb Fix CI

    fzyzcjy/flutter_rust_bridge

    A skill your agent uses when CI fails in flutterrustbridge - before deep investigation

    5.4k GitHub stars~4.4k tokensUpdated yesterday
    MobileAuto-check passed
  • Green Gate

    VeryGoodOpenSource/vgv-ai-flutter-plugin

    Drives a Dart or Flutter package fully green through a verify-fix-rerun loop across four gates: analyze, format, test, coverage.

    169 GitHub stars~4.7k tokensUpdated yesterday
    MobileAuto-check: notes

More from getsentry/sentry-dart

  • Code Guidelines

    getsentry/sentry-dart

    Official

    Enforce Sentry Dart/Flutter SDK code guidelines for implementation, refactoring, and review.

    873 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Test Guidelines

    getsentry/sentry-dart

    Official

    Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures.

    873 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Design First

    getsentry/sentry-dart

    Official

    Shape non-trivial work before writing it — decide the modules, the seams, and the public API surface up front.

    873 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Review

    getsentry/sentry-dart

    Official

    Three-axis review of the branch diff — Standards (this repo's documented standards + public API surface), Spec (the originating Linear issue / PR), and Correctness (runtime bugs + the SDK threat…

    873 GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Diagnosing Bugs

What does Diagnosing Bugs do?

A discipline for hard bugs, flaky tests, CI hangs, and performance regressions in this SDK. Diagnosing Bugs is an agent skill from getsentry/sentry-dart, published by the product's own GitHub organization. A discipline for hard bugs, flaky tests, CI hangs, and performance regressions in this SDK.

When should I use Diagnosing Bugs?

Diagnosing Bugs fits situations like: the user says diagnose; reports something broken/throwing/failing/hanging/slow; A test is flaky.

How do I install Diagnosing Bugs in Claude Code?

Run `npx skills add getsentry/sentry-dart --skill diagnosing-bugs -a claude-code`. Or copy the skill folder (.agents/skills/diagnosing-bugs in getsentry/sentry-dart) into .claude/skills/diagnosing-bugs in your project. Claude Code loads it when a task matches its description.

How do I install Diagnosing Bugs in Codex?

Run `npx skills add getsentry/sentry-dart --skill diagnosing-bugs -a codex`. Or copy the skill folder (.agents/skills/diagnosing-bugs in getsentry/sentry-dart) into .agents/skills/diagnosing-bugs in your project. Codex loads it when a task matches its description.

Can I use Diagnosing Bugs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add getsentry/sentry-dart --skill diagnosing-bugs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diagnosing-bugs, .gemini/skills/diagnosing-bugs, .github/skills/diagnosing-bugs and .opencode/skills/diagnosing-bugs in your project.

What does Diagnosing Bugs need to run?

Going by SKILL.md and its folder, Diagnosing Bugs needs the command-line tools its instructions call (flutter, dart and git).

Does Diagnosing Bugs access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Diagnosing Bugs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Diagnosing Bugs use?

Diagnosing Bugs is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Diagnosing Bugs use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Diagnosing Bugs?

Skills that share tags, products or a category with Diagnosing Bugs: Diagnosing Bugs (getsentry/sentry-react-native, 1.8k stars), Flutter Testing (MADTeacher/mad-agents-skills, 110 stars), Testing (evanca/flutter-ai-rules, 647 stars) and Test Guidelines (getsentry/sentry-react-native, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Diagnosing Bugs?

getsentry (a GitHub organization, an official publisher) maintains it in getsentry/sentry-dart, which has 873 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: getsentry/sentry-dart on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.