Topic · Testing & QA
Best failing and flaky tests skills, page 3
Failing and flaky tests skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Fix failing pure-stage simulation tests that use asserttracematch() by running the test, reading the trace diff in the failure output, and updating the expected trace using tm helpers from… | pragma-org/ | 116 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 98 | HashQL testing strategies including compiletest (UI tests), unit tests, and snapshot tests. | hashintel/ | 1.7k | — | ~1.9k | Automated safety check: Pass | AGPL-3.0 | today |
| 99 | Diagnoses and fixes failing GitHub Actions runs by identifying the failure, fetching only the relevant logs, finding the root cause and reproducing it locally. | ruby-git/ | 1.8k | — | ~1.9k | Automated safety check: Pass | MIT | 6 days ago |
| 100 | 100.GitHub PR Images Embed a local image file into an existing GitHub PR — either in the PR body or as a comment. | bikeindex/ | 308 | — | ~1.8k | Automated safety check: Pass | AGPL-3.0 | today |
| 101 | Create C bindings for Apple frameworks in dotnet/macios. An agent skill from dotnet/macios. | dotnet/ | 2.9k | — | ~7.3k | Automated safety check: Pass | Unknown | today |
| 102 | A skill your agent uses when tests have race conditions, timing dependencies, or inconsistent pass/fail behavior - replaces arbitrary timeouts with condition polling to wait for actual state… | sandgardenhq/ | 137 | 3 repos | ~933 | Automated safety check: Pass | Unknown | 17 days ago |
| 103 | Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging. | pytorch/ | 5.1k | — | ~1.8k | Automated safety check: Pass | Unknown | today |
| 104 | 104.CI Diagnostics Diagnose Proton CI failures and performance comparison results from GitHub checks and uploaded reports. | timeplus-io/ | 2.3k | — | ~685 | Automated safety check: Pass | Apache-2.0 | 17 days ago |
| 105 | Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y. | chongdashu/ | 149 | — | ~2.2k | Automated safety check: Pass | No licence | 5 mo ago |
| 106 | 106.CI Prep Prepares the current branch for CI by running the exact same steps locally and fixing issues. | Nimblesite/ | 138 | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 107 | 107.Release App Cut and publish a new NarraCat-app version — bump the version number, build the Windows x64 package in CI, package + sign + notarize the macOS build locally, and publish both platforms into one… | yannikzz/ | 103 | — | ~1.6k | Automated safety check: Notes | AGPL-3.0 | yesterday |
| 108 | Classifies a failing test, typecheck or CI job before any code changes, by recording the failure and running a clean control to show whether it was already broken. | different-ai/ | 24k | — | ~779 | Automated safety check: Pass | Unknown | today |
| 109 | 109.Fix CI Run a local pnpm monorepo CI loop, fix failures, and stop only when the full sequence is green. | gronxb/ | 1.8k | — | ~572 | Automated safety check: Pass | Unknown | today |
| 110 | 110.Functional Tests A skill your agent uses when writing, editing, reviewing, or running functional (end-to-end) tests for the Astronomer APC repository. | astronomer/ | 491 | — | ~2.2k | Automated safety check: Pass | Unknown | today |
| 111 | 111.CI Triage Triage a failed splice GitHub Actions job (cn-test-failures ref) into a reproducible evidence packet - fetch job log and artifact, isolate the flagged lines, check the known flake families for… | canton-network/ | 118 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 112 | 112.Babysit PR A skill your agent uses when monitoring an open GitHub PR for CI failures, review feedback, mergeability, and safe retries or fixes. | jellydn/ | 123 | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 113 | 113.E2E Run, debug, and manage Playwright e2e tests. An agent skill from sendou-ink/sendou.ink. | sendou-ink/ | 297 | — | ~2.1k | Automated safety check: Notes | AGPL-3.0 | today |
| 114 | 114.Fix Diagnose and fix Session Sniffer bugs, errors, tracebacks, logs, lint failures, static-analysis findings, test failures, and IDE-reported problems. | BUZZARDGTA/ | 104 | — | ~2.7k | Automated safety check: Pass | GPL-3.0 | today |
| 115 | 115.CI Triage Fetch and classify a failed GitHub Actions job for this repo — distinguish out-of-memory kills, six-hour timeouts, configure errors and genuine test failures, and identify what was in flight. | fair-acc/ | 115 | — | ~548 | Automated safety check: Pass | LGPL-3.0 | today |
| 116 | Load when investigating a specific flaky test. An agent skill from DataDog/pup. | DataDog/ | 1k | 1 repo | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 117 | 117.Offload Activate when you see offload.toml in a repo, offload referenced in build targets (justfile, Makefile, scripts), or when you need to run a large test suite in parallel. | imbue-ai/ | 125 | — | ~3.1k | Automated safety check: Pass | MIT | 12 days ago |
| 118 | Post-mortem analysis of CI failures across recent PRs in dotnet/macios. | dotnet/ | 2.9k | — | ~7.8k | Automated safety check: Pass | Unknown | today |
| 119 | 119.Babysit PR Monitor and diagnose GitHub Actions checks on ZenUML web-sequence PRs, fixing code-caused CI failures when appropriate. | ZenUml/ | 150 | — | ~871 | Automated safety check: Pass | MIT | 3 days ago |
| 120 | Run unit tests for the Cosmos Explorer project. An agent skill from Azure/cosmos-explorer. | Azure/ | 131 | — | ~638 | Automated safety check: Pass | MIT | today |
| 121 | Any bug, failing or flaky test, or surprise behavior?. An agent skill from christopherarter/superpowers-reasonix. | christopherarter/ | 102 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 122 | 122.Debug Systematic root-cause debugging — reproduce, isolate, fix at the source, prove the fix. | gnomeria/ | 690 | — | ~715 | Automated safety check: Pass | MIT | 1 mo ago |
| 123 | A skill your agent uses to check GitHub Actions build status or diagnose why a workflow run failed and propose a fix. | MartinStyk/ | 369 | — | ~1.9k | Automated safety check: Pass | GPL-3.0 | 2 days ago |
| 124 | 124.Instructor QA Run multi-dimensional quality assurance for InstructorPHP. An agent skill from cognesy/instructor-php. | cognesy/ | 328 | — | ~1.3k | Automated safety check: Pass | MIT | 2 days ago |
| 125 | A skill your agent uses when writing, fixing, extending, or reviewing OPA5 integration tests for SAP Fiori Elements applications - whether the app uses OData V4 (sap.fe.test library) or OData V2… | SAP/ | 158 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 126 | Shepherd the current user's open PR through base updates, CI failures, and review feedback without rewriting history or merging. | trailofbits/ | 765 | — | ~625 | Automated safety check: Pass | Apache-2.0 | today |
| 127 | Runs an end-to-end workflow to diagnose, reproduce, fix and validate a failing integration test in the Datadog Terraform provider, ending with a draft PR. | DataDog/ | 468 | — | ~2.4k | Automated safety check: Notes | MPL-2.0 | today |
| 128 | Diagnoses and fixes flaky Playwright e2e tests by replacing race-prone patterns with retry-safe alternatives. | Comfy-Org/ | 2.1k | — | ~2.8k | Automated safety check: Pass | GPL-3.0 | today |
| 129 | Guides running, reviewing and fixing test failures across grails-core modules with Gradle, including targeted runs and the aggregate HTML and Markdown reports. | apache/ | 2.9k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 130 | 130.Debug CI Reproduce Linux CI failures locally using Docker when the same tests pass on the host, especially Go platform differences and VS Code extension tests requiring xvfb. | web-infra-dev/ | 460 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 131 | Diagnose and fix required CI failures for the current DDNS branch without weakening tests, platform coverage, caches, or repository policy. | NewFuture/ | 4.7k | — | ~327 | Automated safety check: Pass | MIT | yesterday |
| 132 | Describes a decision tree of steps and skills to utilize when, starting from a test failure or the failure of bazel run command using heir-opt, you would like to produce a reproducing input IR that… | google/ | 929 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 133 | Maintain and follow up on a single Docker documentation pull request that you own or are responsible for updating. | docker/ | 4.7k | — | ~825 | Automated safety check: Pass | Apache-2.0 | today |
| 134 | Fixes a bug with a failing regression test first when a cheap, focused test path exists, and falls back to the closest practical check when it does not. | cursor/ | 10k | 8 repos | ~786 | Automated safety check: Pass | No licence | today |
| 135 | Reviews and accepts Insta snapshot changes in the Oxc repo from the command line, one reviewed snapshot at a time, instead of using the interactive UI. | oxc-project/ | 23k | — | ~356 | Automated safety check: Pass | MIT | today |
| 136 | Fix TeamCity project-leak test failures end to end. An agent skill from JetBrains/intellij-community. | JetBrains/ | 21k | — | ~8.3k | Automated safety check: Pass | Unknown | today |
| 137 | A discipline for hard bugs, flaky tests, CI hangs, native crashes, and performance regressions in this SDK. | getsentry/ | 1.8k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 138 | Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures. | sgl-project/ | 37k | 2 repos | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 139 | Explains how to run and extend the clickup-cli tests: unit tests with a mocked client, e2e tests against a real ClickUp workspace, and the fixture data they rely on. | krodak/ | 121 | — | ~1.4k | Automated safety check: Notes | MIT | yesterday |
| 140 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior. | RedWoodOG/ | 177 | 6 repos | ~2.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 141 | 141.Vibe Debug A short debugging procedure for an existing project: reproduce the failure, test one hypothesis at a time, make the smallest fix, add a regression check and re-verify. | KhazP/ | 3.1k | — | ~403 | Automated safety check: Notes | MIT | 3 days ago |
| 142 | A skill your agent uses when running claudikins-kernel:ship, preparing PRs, writing changelogs, deciding merge strategy, or handling CI failures — enforces GRFP-style iterative approval, code… | povvo/ | 128 | — | ~3.2k | Automated safety check: Notes | MIT | 5 mo ago |
| 143 | Playwright E2E testing expert for browser automation, cross-browser testing, visual regression, network interception, and CI integration. | cin12211/ | 224 | — | ~1.3k | Automated safety check: Pass | MIT | 16 days ago |
| 144 | 144.Openclaw Testing Choose proportional OpenClaw tests and checks, diagnose failures, and route environment-sensitive or release proof to its owner. | openclaw/ | 392k | — | ~1.9k | Automated safety check: Pass | MIT | today |