Topic · Development
Best root cause analysis skills for Claude Code, Codex and other agents.
- skills
- 610
- official
- 56
Root cause analysis skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers. | obra/ | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 2 | Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs. | cursor/ | 10k | 9 repos | ~2.6k | Automated safety check: Pass | No licence | yesterday |
| 3 | Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations. | PowerShell/ | 56k | — | ~5.1k | Automated safety check: Pass | MIT | yesterday |
| 4 | Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant. | AprilNEA/ | 23k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 4 days ago |
| 5 | Forms a two-agent committee with contrasting profiles to analyze a stuck problem in parallel, reconcile their views and return a consensus plan without editing files. | getpaseo/ | 20k | 1 repo | ~496 | Automated safety check: Pass | Unknown | yesterday |
| 6 | Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals. | dotnet/ | 23k | — | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 7 | Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report. | appsmithorg/ | 41k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code. | saadeghi/ | 43k | — | ~2.3k | Automated safety check: Pass | MIT | 8 days ago |
| 9 | Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written. | garrytan/ | 136k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 10 | Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions. | appsmithorg/ | 41k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget. | tirth8205/ | 32k | 1 repo | ~287 | Automated safety check: Pass | MIT | yesterday |
| 12 | Answers questions about a past agent run from its recording, using causal graphs and replay, instead of reconstructing events from memory. | iflytek/ | 5.2k | 4 repos | ~3k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 13 | Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent. | Chachamaru127/ | 3.2k | 1 repo | ~1.1k | Automated safety check: Notes | MIT | 3 days ago |
| 14 | Triages findings from a Strix pentest by severity, fixes each root cause with a minimal change, and re-runs Strix to confirm the exploit no longer works. | usestrix/ | 67k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated. | pingcap/ | 41k | — | ~498 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 16 | Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level. | loopx-project/ | 6.2k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 17 | 17.Reverse Flow Guided reverse engineering workflow for binaries, firmware, mobile apps, scripts, document samples, protocol captures, and unknown artifacts. | lingbol088-spec/ | 935 | — | ~2.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 18 | Fix or implement a tracker issue end to end from a single command — takes an issue id or a plain problem description (filed first via om-prepare-issue), classifies, then drives the bug autofix chain… | go-musicfox/ | 2.6k | 1 repo | ~5k | Automated safety check: Notes | GPL-3.0 | 1 mo ago |
| 19 | Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop. | aiming-lab/ | 15k | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 20 | Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written. | ChrisWiles/ | 6.1k | 3 repos | ~1.2k | Automated safety check: Pass | No licence | 9 mo ago |
| 21 | Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses. | addyosmani/ | 102k | 1 repo | ~2.6k | Automated safety check: Pass | MIT | 4 days ago |
| 22 | Pushes an agent to keep verifying and changing approach after repeated failures, using a diagnosis line, evidence-based completion and confirmation before risky edits. | tanweai/ | 20k | — | ~502 | Automated safety check: Pass | MIT | 28 days ago |
| 23 | Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time. | kubeshark/ | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 24 | Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code. | GreptimeTeam/ | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 25 | Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests. | comet-ml/ | 22k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 26 | Master systematic debugging techniques, profiling tools, and root cause analysis to efficiently track down bugs across any codebase or technology stack. | sangrokjung/ | 849 | 12 repos | ~3.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 27 | Debugging guide for Extempore covering its three layers, compilation paths, startup sequence and the batch, eval and interactive modes used to isolate JIT problems. | digego/ | 1.5k | — | ~4.4k | Automated safety check: Pass | No licence | 13 days ago |
| 28 | Conduct evidence-backed Happier code, plan-completeness, session, worktree, feature, commit, branch, PR, codebase, and release-readiness reviews with affected-corridor analysis, high-confidence… | happier-dev/ | 1.9k | — | ~4.5k | Automated safety check: Pass | MIT | today |
| 29 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | yesterday |
| 30 | 30.Fablize A harness that makes Opus (or any Claude model) behave like Fable — it enforces seeing a task through to the end, with evidence and verification, as procedure. | fivetaku/ | 895 | — | ~1.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 31 | 31.Debug Agent Systematic evidence-based debugging using runtime logs. An agent skill from millionco/expect. | millionco/ | 3.6k | — | ~2.6k | Automated safety check: Pass | Unknown | 5 mo ago |
| 32 | Fixes an OpenROAD bug from a GitHub issue or error code: finds the root cause, implements the fix, adds a regression test and prepares a signed-off commit. | The-OpenROAD-Project/ | 3.2k | — | ~784 | Automated safety check: Pass | BSD-3-Clause | yesterday |
| 33 | A skill your agent uses when asked to run Design Error Detection (quick defect scan), find design errors in a Simulink model, perform root cause analysis on DED findings, fix division-by-zero… | matlab/ | 1.2k | — | ~2.1k | Automated safety check: Pass | Unknown | 7 days ago |
| 34 | 34.Post Mortem Write the canonical engineering record of a fixed bug — root cause, mechanism, fix, validation, and how it slipped through. | thananon/ | 3.2k | — | ~3.4k | Automated safety check: Pass | No licence | 3 mo ago |
| 35 | Walks through open Sentry issues for the tooll3 project, latest first, proposing a fix for each and committing them one at a time with your review between. | tixl3d/ | 5.1k | — | ~2.1k | Automated safety check: Notes | MIT | yesterday |
| 36 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes - four-phase framework (root cause investigation, pattern analysis, hypothesis… | ed3dai/ | 250 | 3 repos | ~2.4k | Automated safety check: Pass | No licence | 1 mo ago |
| 37 | Guides systematic root-cause debugging. An agent skill from abashev/vfs-s3. | abashev/ | 106 | 6 repos | ~2.6k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 38 | Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons. | rampstackco/ | 935 | 1 repo | ~2.5k | Automated safety check: Pass | MIT | today |
| 39 | Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label. | bgauryy/ | 946 | — | ~1.5k | Automated safety check: Pass | MIT | 4 days ago |
| 40 | 40.CI Triage Triage failing GitHub PR checks: list failures with gh, fetch capped Actions logs, skip non-Actions checks, and summarize root cause. | Mentra-Community/ | 2.4k | — | ~582 | Automated safety check: Pass | Apache-2.0 | today |
| 41 | Analyze MSBuild binary logs to diagnose build failures. An agent skill from microsoft/testfx. | microsoft/ | 1k | 3 repos | ~730 | Automated safety check: Pass | MIT | today |
| 42 | Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest. | microsoft/ | 969 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 43 | 矛盾分析法:把复杂问题拆成若干对立面,找出规定其他矛盾的主要矛盾及其主要方面,判定对抗性 / 非对抗性,并据此选择处理方式。当问题头绪多、多个因素互相牵制、优先级不清、根因不明、反复修不好、trade-off 说不清时触发;直接执行类任务或用户已定方案时不触发。 | HughYau/ | 3.8k | — | ~485 | Automated safety check: Pass | MIT | 6 days ago |
| 44 | After a bug is found, traces its root cause and feeds a new testable invariant back into the project spec so the bug class can't recur. | JuliusBrussee/ | 1.1k | — | ~653 | Automated safety check: Pass | MIT | 1 mo ago |
| 45 | Forces a one-sentence, evidence-backed root cause before any fix is applied, and gates when a diagnosis session is even allowed to touch code. | tw93/ | 7.2k | — | ~4.3k | Automated safety check: Pass | MIT | yesterday |
| 46 | 46.Veomni Debug A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 47 | Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments. | alibaba/ | 412 | — | ~1.9k | Automated safety check: Pass | Unknown | 14 days ago |
| 48 | 48.Diagnose Investigate unexpected behavior and mysterious bugs. An agent skill from avibebuilder/claude-prime. | avibebuilder/ | 120 | 1 repo | ~1.2k | Automated safety check: Pass | MIT | 4 mo ago |
Questions, answered from the data.
What is the best root cause analysis skill?
Diagnosing Superpowers Sessions from obra/superpowers ranks first of the 610 root cause analysis skills listed here, with the highest score: its repository has 296k GitHub stars, 3 other GitHub owners carry a copy, its SKILL.md loads about 1.7k tokens and it passes the automated safety check with no findings. Next come Code Design Rationale Investigator and Pester Failure Analysis.
Which root cause analysis skills are official?
56 of the 610 root cause analysis skills are official, published by the vendor's own GitHub organization: Code Design Rationale Investigator, Copilot Session Failure Analysis, Binlog Failure Analysis, Trx Analysis, Triagebot Action Bug Triage and 51 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in Development
- Pull requests1,927
- Refactoring1,154
- Debugging1,154
- Changelog and release notes1,147
- Code review1,053
- Technical documentation998
- Linting and formatting830
- Git workflow809
- Diagrams801
- Project scaffolding742
- Code quality579
- Architecture decision records562
- Commit messages556
- Git worktrees528
- Performance optimization440
- Design patterns408
- Error handling317
- Type safety301
- Dependency management285
- Software architecture280
- Monorepo tooling267
- Async programming243
- Issue triage234
- Codebase onboarding232