Topic · Testing & QA
Best failing and flaky tests skills for Claude Code, Codex and other agents.
- skills
- 529
- official
- 78
Failing and flaky tests skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way. | openinterpreter/ | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations. | PowerShell/ | 56k | — | ~5.1k | Automated safety check: Pass | MIT | today |
| 3 | Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions. | langgenius/ | 158k | — | ~682 | Automated safety check: Pass | Unknown | today |
| 4 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes | ultralisp/ | 258 | 51 repos | ~2.4k | Automated safety check: Pass | No licence | 24 days ago |
| 5 | Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report. | appsmithorg/ | 41k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Log genuine, recurring repository friction to .agents/PAPERCUTS.md — confusing setup, a flaky repo command or script, a misleading in-repo error, stale generated files, or a non-obvious gotcha that… | every-app/ | 23k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 7 | Runs RustPython tests inside a Linux container built with Apple's container CLI, so macOS users can compare Linux results with their local ones. | RustPython/ | 22k | — | ~467 | Automated safety check: Pass | MIT | today |
| 8 | Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions. | appsmithorg/ | 41k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Iterate on a PR until CI passes. An agent skill from meshery/meshery-operator. | meshery/ | 151 | 7 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 10 | Takes a GitHub or YouTrack issue for the Exposed project through reproduction, a failing test, a fix, validation and a pull request. | JetBrains/ | 9.3k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Downloads Azure Pipelines CI logs for an Ansible pull request or build so the agent can analyze test failures, after asking you first. | ansible/ | 71k | — | ~825 | Automated safety check: Pass | GPL-3.0 | today |
| 12 | Inspect, analyze, troubleshoot, or review Codacy findings and local analyzer/API helpers. | netdata/ | 81k | — | ~2.2k | Automated safety check: Notes | GPL-3.0 | today |
| 13 | 13.Update V86 Build and install v86 (wasm + libv86.js + BIOS) into windows95. | felixrieseberg/ | 24k | — | ~1.7k | Automated safety check: Pass | Unknown | 26 days ago |
| 14 | Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real… | quay/ | 2.8k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent. | Chachamaru127/ | 3.2k | 1 repo | ~1.1k | Automated safety check: Notes | MIT | 2 days ago |
| 16 | Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated. | pingcap/ | 41k | — | ~498 | Automated safety check: Pass | Apache-2.0 | today |
| 17 | 17.Swig Test Run SWIG test suite for specific languages. An agent skill from swig/swig. | swig/ | 6.3k | — | ~2.3k | Automated safety check: Pass | Unknown | today |
| 18 | A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests | payloadcms/ | 45k | — | ~4.4k | Automated safety check: Pass | MIT | today |
| 19 | Fixes a React Router bug reported in a GitHub issue end to end: fetching the issue, validating the reproduction, writing a failing test and implementing the fix on a new branch. | remix-run/ | 57k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 20 | Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written. | ChrisWiles/ | 6.1k | 3 repos | ~1.2k | Automated safety check: Pass | No licence | 9 mo ago |
| 21 | Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses. | addyosmani/ | 102k | 1 repo | ~2.6k | Automated safety check: Pass | MIT | 3 days ago |
| 22 | Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code. | GreptimeTeam/ | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 23 | Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests. | comet-ml/ | 22k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | Upgrades a Python standard library module from CPython into RustPython with update_lib, then triages and marks the tests that still fail. | RustPython/ | 22k | — | ~876 | Automated safety check: Pass | MIT | today |
| 25 | 25.CI Fix Scan all CI builds and tests, find failures, fetch error logs, and fix the code. | FastLED/ | 7.5k | — | ~897 | Automated safety check: Pass | MIT | today |
| 26 | 26.Fix Issue Fix a reported issue in Remix from a GitHub issue. An agent skill from remix-run/remix. | remix-run/ | 33k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 27 | Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm. | guqiong96/ | 464 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 28 | 28.Testing A skill your agent uses for every Kortix test task, behavior change, bug fix, refactor, API route change, CLI change, SDK change, browser journey, test failure, coverage question, local benchmark… | kortix-ai/ | 20k | — | ~3.6k | Automated safety check: Notes | Unknown | today |
| 29 | Diagnose any Quay Prow job failure end to end: prowjob.json - top-level build log - JUnit - resolved failing step - Playwright results.json when the failing step is Playwright, continuing through… | quay/ | 2.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 30 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior during spec-superflow execution, before proposing fixes. | MageByte-Zero/ | 839 | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 5 days ago |
| 31 | Analyze Vortex GitHub Actions CI failures. An agent skill from vortex-data/vortex. | vortex-data/ | 3.2k | — | ~810 | Automated safety check: Pass | Apache-2.0 | today |
| 32 | Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change. | microsoft/ | 22k | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 33 | 33.CI Watchdog Continuously monitor GitHub PR CI checks and automatically fix failures until all checks pass. | latitude-dev/ | 4.7k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 34 | Runs a gated finish-line checklist before committing a PlotJuggler PJ4 change: build proof, red-test triage, hooks, docs freshness and a diff self-review. | PlotJuggler/ | 6.2k | — | ~1.3k | Automated safety check: Pass | MPL-2.0 | 6 days ago |
| 35 | Plans the smallest check that could disprove a code change in the OpenLogi project, then escalates through reproduction, focused tests and a final gate before a push. | AprilNEA/ | 23k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 36 | Investigate and triage CI failures for dotnet/macios from Azure DevOps build URLs. | dotnet/ | 2.9k | — | ~2.3k | Automated safety check: Pass | Unknown | today |
| 37 | Reproduce a GitHub Actions Linux CI failure locally when it does not happen on your machine: a podman/docker image that mirrors the ubuntu-22.04 runner by reusing the real Tools/CI-linux-.sh install… | swig/ | 6.3k | — | ~1.2k | Automated safety check: Pass | Unknown | today |
| 38 | Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests. | DynamoDS/ | 2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 39 | Run and debug Agmente iOS end-to-end tests against a real local Codex CLI app-server instance. | rebornix/ | 545 | — | ~625 | Automated safety check: Pass | MIT | 4 mo ago |
| 40 | 40.PR Review Address review comments and CI failures for the current branch's PR | wysaid/ | 1.9k | — | ~1.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 41 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes - four-phase framework (root cause investigation, pattern analysis, hypothesis… | ed3dai/ | 250 | 3 repos | ~2.4k | Automated safety check: Pass | No licence | 1 mo ago |
| 42 | Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala). | apache/ | 4.6k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 43 | Triage a flaky Playwright test end to end, from a Sippy signal to a written fix proposal: Sippy numbers and failing run URLs, Prow artifacts (or the access gap), the spec, a local reproduction, and… | quay/ | 2.8k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 44 | Guides systematic root-cause debugging. An agent skill from abashev/vfs-s3. | abashev/ | 106 | 6 repos | ~2.6k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 45 | Investigates a failing RustPython test by comparing it with CPython, then either fixes it or gathers the details for an incompatibility report. | RustPython/ | 22k | — | ~467 | Automated safety check: Pass | MIT | today |
| 46 | Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches. | fastrepl/ | 9.4k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 47 | 47.Wdio Testing Write, run, and debug WebDriverIO (WDIO) UI tests for the Ansible VS Code extension. | ansible/ | 486 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 48 | 48.CI Triage Triage failing GitHub PR checks: list failures with gh, fetch capped Actions logs, skip non-Actions checks, and summarize root cause. | Mentra-Community/ | 2.4k | — | ~582 | Automated safety check: Pass | Apache-2.0 | today |
Questions, answered from the data.
What is the best failing and flaky tests skill?
PR Babysitter from openinterpreter/openinterpreter ranks first of the 529 failing and flaky tests skills listed here, with the highest score: its repository has 69k GitHub stars, 3 other GitHub owners carry a copy, its SKILL.md loads about 4.2k tokens and it passes the automated safety check with no findings. Next come Pester Failure Analysis and Cucumber and Playwright E2E Tests.
Which failing and flaky tests skills are official?
78 of the 529 failing and flaky tests skills are official, published by the vendor's own GitHub organization: Exposed Bug Fix Workflow, ONNX Runtime CI Management, Macios CI Failure Inspector, Fix Flakes, Trx Analysis and 73 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.