Topic · Testing & QA
Best failing and flaky tests skills, page 5
Failing and flaky tests skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 193 | 193.PR Babysitter Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes. | mblode/ | 143 | — | ~3.4k | Automated safety check: Pass | MIT | 2 days ago |
| 194 | Skill to automatically triage failing tests by reading test-metadata status files, matching GitHub issues, and generating a multi-section failure report. | GoogleCloudPlatform/ | 974 | — | ~1k | Automated safety check: Pass | Unknown | today |
| 195 | How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky. | bikeindex/ | 308 | — | ~5k | Automated safety check: Pass | AGPL-3.0 | today |
| 196 | A skill your agent uses to capture visual artefacts from a device for test failures, golden image generation, QA repro, and demo videos. | skydoves/ | 333 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 197 | Coordinate with other agents/humans via the optional Interlinked MCP Server, and use local checkpoints & file reservations. | QuentinCody/ | 178 | — | ~2.9k | Automated safety check: Pass | MIT | 6 days ago |
| 198 | Runs a hypothesis-driven debugging loop for crashes, hangs and silent failures in any language, grounding every claim in runtime evidence and locking the fix with a test. | code-yeongyu/ | 70k | — | ~3.2k | Automated safety check: Pass | Unknown | today |
| 199 | 199.Start Work Sets up a druxt.js change before the first edit, from the issue to a feature branch off develop, a short spec and a failing test. | druxt/ | 114 | — | ~887 | Automated safety check: Pass | MIT | today |
| 200 | 200.Investigate Systematic debugging skill. An agent skill from blueberrycongee/termcanvas. | blueberrycongee/ | 406 | — | ~562 | Automated safety check: Pass | MIT | 4 mo ago |
| 201 | 201.PR Ready Clear everything blocking an existing PR/MR from merging — rebase onto the default branch, get CI green, and resolve review threads (Copilot and human). | OutThisLife/ | 199 | — | ~2.5k | Automated safety check: Pass | MIT | 2 days ago |
| 202 | 202.Testing Marchat Writes and runs marchat tests with race detection, coverage, and dialect smoke patterns. | Cod-e-Codes/ | 137 | — | ~802 | Automated safety check: Pass | MIT | 5 days ago |
| 203 | 203.Novu Prepare PR Post-implementation PR prep for Novu feature branches — quality passes, commit/PR hygiene, CI triage, and review feedback. | novuhq/ | 40k | — | ~1.5k | Automated safety check: Pass | Unknown | today |
| 204 | 204.Absolute Deflake Flaky test fixes: detect nondeterministic tests empirically (repeat/shuffle/parallel runs), diagnose the root cause, fix it — never retry/skip/sleep — and verify across many randomized runs. | maddhruv/ | 218 | 1 repo | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 205 | Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format. | Azure/ | 134 | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 206 | Finds real correctness bugs in code changes. An agent skill from getsentry/warden. | getsentry/ | 414 | — | ~1.9k | Automated safety check: Pass | Unknown | 8 days ago |
| 207 | 207.Continue PR Continue work on an existing PR - resolve conflicts, fix CI failures, address reviewer feedback, and push updates. | ClickHouse/ | 50k | — | ~5.5k | Automated safety check: Notes | Apache-2.0 | today |
| 208 | Unit/integration testing standards for RedisInsight using Jest and Testing Library: test structure, the renderComponent helper, faker for test data, mocking patterns, and waitFor instead of fixed… | redis/ | 8.9k | — | ~3.3k | Automated safety check: Pass | Unknown | 3 days ago |
| 209 | 209.Test Fixing Run tests and systematically fix all failing tests using smart error grouping. | davila7/ | 32k | 8 repos | ~739 | Automated safety check: Pass | MIT | today |
| 210 | 210.Make Git Escrow Create a new git escrow bounty for a test suite. An agent skill from internet-court/internet-court-skill. | internet-court/ | 6.4k | 2 repos | ~922 | Automated safety check: Notes | MIT | 1 mo ago |
| 211 | Finds which commit broke a failing Swift package suite in the cmux repo by bisecting on CI, then judges per test whether it went stale or the code regressed. | manaflow-ai/ | 28k | — | ~1.5k | Automated safety check: Pass | Unknown | today |
| 212 | 212.CI Debugging Systematic CI/CD failure diagnosis using hypothesis-first investigation, local reproduction, and environment delta analysis. | citypaul/ | 739 | — | ~1.5k | Automated safety check: Notes | Unknown | 5 days ago |
| 213 | Parse a Prow CI job URL or Jira ticket to extract E2E test failure details including test name, spec file, release branch, platform, and error messages | redhat-developer/ | 172 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | today |
| 214 | 214.Heal PR Heal a Blade PR by fixing CI failures, missing changesets, and sanity issues. | razorpay/ | 656 | — | ~414 | Automated safety check: Pass | MIT | yesterday |
| 215 | Write and evaluate effective Python tests using pytest. An agent skill from iusztinpaul/squid. | iusztinpaul/ | 203 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 216 | Runbook for closing gaps between a Flowfile visual flow's results and its exported Polars or FlowFrame Python code, measured by tests rather than by eye. | Edwardvaneechoud/ | 373 | — | ~7.5k | Automated safety check: Pass | MIT | today |
| 217 | 217.Bug Hunter Find and fix a bug in the octane monorepo. An agent skill from octanejs/octane. | octanejs/ | 1.4k | — | ~707 | Automated safety check: Pass | MIT | today |
| 218 | Fetch and analyze Prow e2e pipeline failures for cloud-provider-azure. | kubernetes-sigs/ | 294 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 219 | Reviews the uncommitted changes in an adk-python working tree and reports correctness, design, public-API stability, test, sample and documentation gaps as a prioritized findings report, fixing them… | google/ | 22k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 220 | Keep iterating on code changes until the tests pass, the build succeeds, or linting is clean. | spencerpauly/ | 842 | — | ~796 | Automated safety check: Pass | CC0-1.0 | 2 mo ago |
| 221 | 221.CI Red Sweep Repairs confirmed CI failures on the base branch first, then prepares each affected PR one at a time with exact-commit evidence, never merging or weakening tests. | diegosouzapw/ | 74k | — | ~868 | Automated safety check: Pass | MIT | today |
| 222 | Audit open "flaky test" GitHub issues and close those whose tests are no longer failing on master. | ClickHouse/ | 50k | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | today |
| 223 | 223.Perf Report Analyze CI performance comparison reports for a ClickHouse PR. | ClickHouse/ | 50k | — | ~2.6k | Automated safety check: Notes | Apache-2.0 | today |
| 224 | Walks the agent through a fixed loop for failing tests and build errors: reproduce, isolate, hypothesize, instrument, fix, verify, then add a regression test. | Totoro-jam/ | 344 | — | ~319 | Automated safety check: Pass | MIT | 1 mo ago |
| 225 | 225.Run Tests Runs .NET tests with dotnet test. An agent skill from runceel/ReactiveProperty. | runceel/ | 944 | — | ~3.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 226 | 226.Monitor PR Monitor PR for CI failures and comments, act on comments if necessary, respond to comments, and resolve comment threads. | doc-detective/ | 135 | — | ~155 | Automated safety check: Pass | AGPL-3.0 | 6 days ago |
| 227 | Analysing GitHub Actions CI failures from a run URL. An agent skill from golemcloud/golem. | golemcloud/ | 1.5k | — | ~862 | Automated safety check: Pass | Unknown | today |
| 228 | 228.Implementation Write application code to make failing tests pass using contract-driven, slice-based architecture. | EmeaAppGbb/ | 100 | — | ~2.8k | Automated safety check: Pass | MIT | 5 mo ago |
| 229 | Expert in diagnosing Playwright failures. Uses the @playwright MCP to verify live DOM state. | Tahanima/ | 113 | — | ~210 | Automated safety check: Pass | MIT | 7 mo ago |
| 230 | Reads CI results for a dotnet/maui pull request and reports in one short comment whether the failures relate to the PR or to the base branch. | dotnet/ | 23k | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 231 | Submits .NET MAUI unit tests to Helix queues from a local machine and monitors job status and per-work-item logs with PowerShell scripts. | dotnet/ | 23k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 232 | Attempts one alternative fix for a bug, runs the given test command against it and reports what happened, always differing from existing PR fixes. | dotnet/ | 23k | — | ~8.4k | Automated safety check: Pass | MIT | today |
| 233 | Confirms that newly added tests actually fail without the fix, auto-detecting UI, device, unit or XAML tests and running the matching runner. | dotnet/ | 23k | — | ~2.7k | Automated safety check: Pass | MIT | today |
| 234 | Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug. | dotnet/ | 23k | — | ~3k | Automated safety check: Pass | MIT | today |
| 235 | Hypothesis-driven debugging methodology for hard bugs. An agent skill from QwenLM/qwen-code. | QwenLM/ | 28k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 236 | 236.Dev Debug This skill should be used when the user asks to "dev-debug", "テストが失敗する", "ビルドエラーを直して", "デバッグ", "debug failing tests", "fix build error", "エラーを修正", "コンパイルエラー", "環境の問題を解決". | classmethod/ | 974 | — | ~1.6k | Automated safety check: Pass | MIT | 2 mo ago |
| 237 | 237.Red Green Fix Bug fix workflow that proves test validity with a red-then-green CI sequence. | Comfy-Org/ | 2.1k | — | ~1.9k | Automated safety check: Pass | GPL-3.0 | today |
| 238 | 238.Writing Tests This skill should be used when writing, running, or fixing tests in the brepjs repository — when a task says "add a test", "write a regression test", "tests are failing", "test timed out", "coverage… | andymai/ | 114 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | today |
| 239 | Replaces trial-and-error fixing with an observe, hypothesize, experiment and conclude loop kept in DEBUG.md, where no fix is allowed before evidence supports a cause. | LichAmnesia/ | 234 | — | ~2.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 240 | 240.Has Review Work Decide whether a GitHub PR has unanswered authorized review comments or new required CI failures worth a follow-up agent. | openshift-eng/ | 120 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |