Topic · Testing & QA

Best failing and flaky tests skills for Claude Code, Codex and other agents.

Skills that diagnose failing, flaky and order-dependent tests and red CI builds and propose fixes.
skills
529
official
78

Failing and flaky tests skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Failing and flaky tests skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

openinterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0today
2

Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

PowerShell/PowerShell56k—~5.1kAutomated safety check: PassMITtoday
3

Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

langgenius/dify158k—~682Automated safety check: PassUnknowntoday
4

A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes

ultralisp/ultralisp25851 repos~2.4kAutomated safety check: PassNo licence24 days ago
5

Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

appsmithorg/appsmith41k—~1.5kAutomated safety check: PassApache-2.0today
6

Log genuine, recurring repository friction to .agents/PAPERCUTS.md — confusing setup, a flaky repo command or script, a misleading in-repo error, stale generated files, or a non-obvious gotcha that…

every-app/open-seo23k1 repo~1.2kAutomated safety check: PassMITtoday
7

Runs RustPython tests inside a Linux container built with Apple's container CLI, so macOS users can compare Linux results with their local ones.

RustPython/RustPython22k—~467Automated safety check: PassMITtoday
8

Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions.

appsmithorg/appsmith41k—~1.3kAutomated safety check: PassApache-2.0today
9

Iterate on a PR until CI passes. An agent skill from meshery/meshery-operator.

meshery/meshery-operator1517 repos~2.2kAutomated safety check: PassApache-2.016 days ago
10

Takes a GitHub or YouTrack issue for the Exposed project through reproduction, a failing test, a fix, validation and a pull request.

JetBrains/Exposed9.3k—~3.8kAutomated safety check: PassApache-2.0today
11

Downloads Azure Pipelines CI logs for an Ansible pull request or build so the agent can analyze test failures, after asking you first.

ansible/ansible71k—~825Automated safety check: PassGPL-3.0today
12

Inspect, analyze, troubleshoot, or review Codacy findings and local analyzer/API helpers.

netdata/netdata81k—~2.2kAutomated safety check: NotesGPL-3.0today
13

Build and install v86 (wasm + libv86.js + BIOS) into windows95.

felixrieseberg/windows9524k—~1.7kAutomated safety check: PassUnknown26 days ago
14

Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…

quay/quay2.8k—~2.2kAutomated safety check: PassApache-2.0today
15

Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent.

Chachamaru127/claude-code-harness3.2k1 repo~1.1kAutomated safety check: NotesMIT2 days ago
16

Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated.

pingcap/tidb41k—~498Automated safety check: PassApache-2.0today
17

Run SWIG test suite for specific languages. An agent skill from swig/swig.

swig/swig6.3k—~2.3kAutomated safety check: PassUnknowntoday
18

A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests

payloadcms/payload45k—~4.4kAutomated safety check: PassMITtoday
19

Fixes a React Router bug reported in a GitHub issue end to end: fetching the issue, validating the reproduction, writing a failing test and implementing the fix on a new branch.

remix-run/react-router57k—~1.3kAutomated safety check: PassMITtoday
20

Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

ChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNo licence9 mo ago
21

Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

addyosmani/agent-skills102k1 repo~2.6kAutomated safety check: PassMIT3 days ago
22

Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

GreptimeTeam/greptimedb6.7k—~4.4kAutomated safety check: PassApache-2.02 days ago
23

Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

comet-ml/opik22k—~1.8kAutomated safety check: PassApache-2.0today
24

Upgrades a Python standard library module from CPython into RustPython with update_lib, then triages and marks the tests that still fail.

RustPython/RustPython22k—~876Automated safety check: PassMITtoday
25

Scan all CI builds and tests, find failures, fetch error logs, and fix the code.

FastLED/FastLED7.5k—~897Automated safety check: PassMITtoday
26

Fix a reported issue in Remix from a GitHub issue. An agent skill from remix-run/remix.

remix-run/remix33k—~1.8kAutomated safety check: PassMITyesterday
27

Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

guqiong96/Lvllm4642 repos~349Automated safety check: PassApache-2.015 days ago
28

A skill your agent uses for every Kortix test task, behavior change, bug fix, refactor, API route change, CLI change, SDK change, browser journey, test failure, coverage question, local benchmark…

kortix-ai/suna20k—~3.6kAutomated safety check: NotesUnknowntoday
29

Diagnose any Quay Prow job failure end to end: prowjob.json - top-level build log - JUnit - resolved failing step - Playwright results.json when the failing step is Playwright, continuing through…

quay/quay2.8k—~2.9kAutomated safety check: PassApache-2.0today
30

A skill your agent uses when encountering any bug, test failure, or unexpected behavior during spec-superflow execution, before proposing fixes.

MageByte-Zero/spec-superflow8391 repo~1.6kAutomated safety check: PassMIT5 days ago
31

Analyze Vortex GitHub Actions CI failures. An agent skill from vortex-data/vortex.

vortex-data/vortex3.2k—~810Automated safety check: PassApache-2.0today
32

Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change.

microsoft/onnxruntime22k—~4.1kAutomated safety check: PassMITtoday
33

Continuously monitor GitHub PR CI checks and automatically fix failures until all checks pass.

latitude-dev/latitude-llm4.7k—~1.6kAutomated safety check: PassMITyesterday
34

Runs a gated finish-line checklist before committing a PlotJuggler PJ4 change: build proof, red-test triage, hooks, docs freshness and a diff self-review.

PlotJuggler/PlotJuggler6.2k—~1.3kAutomated safety check: PassMPL-2.06 days ago
35

Plans the smallest check that could disprove a code change in the OpenLogi project, then escalates through reproduction, focused tests and a final gate before a push.

AprilNEA/OpenLogi23k—~1.4kAutomated safety check: PassApache-2.03 days ago
36

Investigate and triage CI failures for dotnet/macios from Azure DevOps build URLs.

dotnet/macios2.9k—~2.3kAutomated safety check: PassUnknowntoday
37

Reproduce a GitHub Actions Linux CI failure locally when it does not happen on your machine: a podman/docker image that mirrors the ubuntu-22.04 runner by reusing the real Tools/CI-linux-.sh install…

swig/swig6.3k—~1.2kAutomated safety check: PassUnknowntoday
38

Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

DynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0today
39

Run and debug Agmente iOS end-to-end tests against a real local Codex CLI app-server instance.

rebornix/Agmente545—~625Automated safety check: PassMIT4 mo ago
40

Address review comments and CI failures for the current branch's PR

wysaid/android-gpuimage-plus1.9k—~1.4kAutomated safety check: PassMIT2 mo ago
41

A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes - four-phase framework (root cause investigation, pattern analysis, hypothesis…

ed3dai/ed3d-plugins2503 repos~2.4kAutomated safety check: PassNo licence1 mo ago
42

Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

apache/fory4.6k—~2.2kAutomated safety check: PassApache-2.0today
43

Triage a flaky Playwright test end to end, from a Sippy signal to a written fix proposal: Sippy numbers and failing run URLs, Prow artifacts (or the access gap), the spec, a local reproduction, and…

quay/quay2.8k—~2.6kAutomated safety check: PassApache-2.0today
44

Guides systematic root-cause debugging. An agent skill from abashev/vfs-s3.

abashev/vfs-s31066 repos~2.6kAutomated safety check: PassApache-2.07 days ago
45

Investigates a failing RustPython test by comparing it with CPython, then either fixes it or gathers the details for an incompatibility report.

RustPython/RustPython22k—~467Automated safety check: PassMITtoday
46

Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.

fastrepl/anarlog9.4k—~1.4kAutomated safety check: PassMITtoday
47

Write, run, and debug WebDriverIO (WDIO) UI tests for the Ansible VS Code extension.

ansible/vscode-ansible486—~2.2kAutomated safety check: PassMITtoday
48

Triage failing GitHub PR checks: list failures with gh, fetch capped Actions logs, skip non-Actions checks, and summarize root cause.

Mentra-Community/MentraOS2.4k—~582Automated safety check: PassApache-2.0today

Questions, answered from the data.

What is the best failing and flaky tests skill?

PR Babysitter from openinterpreter/openinterpreter ranks first of the 529 failing and flaky tests skills listed here, with the highest score: its repository has 69k GitHub stars, 3 other GitHub owners carry a copy, its SKILL.md loads about 4.2k tokens and it passes the automated safety check with no findings. Next come Pester Failure Analysis and Cucumber and Playwright E2E Tests.

Which failing and flaky tests skills are official?

78 of the 529 failing and flaky tests skills are official, published by the vendor's own GitHub organization: Exposed Bug Fix Workflow, ONNX Runtime CI Management, Macios CI Failure Inspector, Fix Flakes, Trx Analysis and 73 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.