Search
Failing and flaky tests
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Triage a flaky Playwright test end to end, from a Sippy signal to a written fix proposal: Sippy numbers and failing run URLs, Prow artifacts (or the access gap), the spec, a local reproduction, and… | quay/ | 2.8k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 50 | Guides systematic root-cause debugging. An agent skill from abashev/vfs-s3. | abashev/ | 106 | 6 repos | ~2.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 51 | Investigates a failing RustPython test by comparing it with CPython, then either fixes it or gathers the details for an incompatibility report. | RustPython/ | 22k | — | ~467 | Automated safety check: Pass | MIT | today |
| 52 | Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches. | fastrepl/ | 9.5k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 53 | 53.Papercuts Log genuine, recurring repository friction to .agents/PAPERCUTS.md — confusing setup, a flaky repo command or script, a misleading in-repo error, stale generated files, or a non-obvious gotcha that… | every-app/ | 23k | — | ~1.2k | Automated safety check: Pass | MIT | 2 days ago |
| 54 | 54.Wdio Testing Write, run, and debug WebDriverIO (WDIO) UI tests for the Ansible VS Code extension. | ansible/ | 488 | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 55 | Diagnose and fix flaky tests tracked as open kind/flake issues in kubernetes-sigs/agent-sandbox — reproduce the flake, apply a minimal fix, and open a PR linking the issue. | kubernetes-sigs/ | 4.2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 56 | Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest. | microsoft/ | 969 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 57 | Build, compile, run server/client/cluster, execute unit/stateless/SQL tests, verify results, and troubleshoot build/test failures. | timeplus-io/ | 2.3k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 21 days ago |
| 58 | Enforce Sentry React Native SDK test conventions for naming, structure, mocking, and fixtures with Jest. | getsentry/ | 1.8k | — | ~1.3k | Automated safety check: Pass | MIT | 2 days ago |
| 59 | Runs Google Test binaries in parallel with gtest-parallel to speed up single-threaded tests, repeat flaky ones and filter specific tests. | webrtc-sdk/ | 446 | 1 repo | ~432 | Automated safety check: Pass | BSD-3-Clause | 5 days ago |
| 60 | 60.Wio Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health. | workersio/ | 204 | — | ~5.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 61 | After a bug is found, traces its root cause and feeds a new testable invariant back into the project spec so the bug class can't recur. | JuliusBrussee/ | 1.2k | — | ~653 | Automated safety check: Pass | MIT | 1 mo ago |
| 62 | 62.Veomni Debug A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 63 | 63.Skillhone Local Issue, pull-request, and Wiki workbench for agent skills. | Tencent/ | 169 | — | ~3.4k | Automated safety check: Pass | Unknown | 21 days ago |
| 64 | A skill your agent uses when the user asks to fix a bug, references a GitHub issue number, or describes an issue and wants a fix. | brunosabot/ | 269 | — | ~529 | Automated safety check: Pass | MIT | 4 mo ago |
| 65 | Testing CKEditor 5 plugins in the Trilium monorepo. An agent skill from TriliumNext/Trilium. | TriliumNext/ | 38k | — | ~3.3k | Automated safety check: Pass | AGPL-3.0 | today |
| 66 | Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures. | getsentry/ | 873 | — | ~3.1k | Automated safety check: Pass | MIT | yesterday |
| 67 | Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance. | microsoft/ | 22k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 68 | 68.Reviewloop Iteratively improves a PR until all review bots (Greptile, Devin, and others) are satisfied with zero unresolved comments, then fixes any CI failures. | ankitvgupta/ | 495 | — | ~2.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 69 | Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM… | agent-substrate/ | 4.8k | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
| 70 | Analyze current GraalPy periodic job failures for ROTA. An agent skill from oracle/graalpython. | oracle/ | 1.7k | — | ~792 | Automated safety check: Pass | Unknown | 2 days ago |
| 71 | Expertise in analyzing flake github issues in the flutter/flutter repository. | flutter/ | 180k | — | ~1.1k | Automated safety check: Pass | BSD-3-Clause | today |
| 72 | Analyze Azure DevOps CI build failures in dd-trace-dotnet pipeline. | DataDog/ | 572 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 73 | Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2. | CloudAI-X/ | 1.4k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | 5 days ago |
| 74 | 74.MCP Debugger A skill your agent uses when investigating a bug, failing test, or unexpected runtime behavior and the mcp-debugger MCP server is available — drives real step-through debuggers (breakpoints, stack… | debugmcp/ | 174 | — | ~4.2k | Automated safety check: Pass | MIT | today |
| 75 | Writes Vitest tests following project patterns: tests/ directories, vi.mock() for module mocking with vi.hoisted() for test-time factories, global LLM mock from src/test/setup.ts, environment… | caliber-ai-org/ | 1.3k | — | ~3.2k | Automated safety check: Pass | MIT | 17 days ago |
| 76 | 76.Dbg Debug applications using the dbg CLI debugger. An agent skill from theodo-group/debug-that. | theodo-group/ | 158 | — | ~2.5k | Automated safety check: Pass | MIT | 2 days ago |
| 77 | 77.Quicksilver Offload bulk judgment calls to Jev (TypeSafe's fast System One model) so Claude doesn't read, and pay for, content it only needs a verdict on. | UditAkhourii/ | 116 | — | ~2.2k | Automated safety check: Notes | MIT | 15 days ago |
| 78 | Guide for diagnosing GitHub Actions test failures, extracting failed tests from runs, and creating or updating failing-test issues. | microsoft/ | 6.4k | — | ~4.7k | Automated safety check: Pass | MIT | today |
| 79 | Debug Playwright E2E test failures from GitHub Actions CI runs. | quay/ | 2.8k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 80 | 80.CI Triage Triage failing GitHub PR checks: list failures with gh, fetch capped Actions logs, skip non-Actions checks, and summarize root cause. | Mentra-Community/ | 2.4k | — | ~582 | Automated safety check: Pass | Apache-2.0 | today |
| 81 | 81.Releasing Version and release c15t packages with Changesets. An agent skill from c15t/c15t. | c15t/ | 1.9k | — | ~753 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 82 | Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type. | huiliyi37/ | 1.1k | — | ~1k | Automated safety check: Notes | Apache-2.0 | today |
| 83 | Analyzes recent CI failures on pull requests to identify flaky tests, using retry outcomes (failed attempt → green re-run) and cross-PR recurrence as evidence, and maintains a local longitudinal… | opsmill/ | 534 | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 84 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 85 | Investigate and fix flaky/random CI test failures in dotnet/macios. | dotnet/ | 2.9k | — | ~1.3k | Automated safety check: Pass | Unknown | 2 days ago |
| 86 | 86.Mz Debug CI Investigate CI failures on PR via gh + Buildkite MCP or bk CLI. | MaterializeInc/ | 6.4k | — | ~3.9k | Automated safety check: Pass | Unknown | today |
| 87 | Update ALLOWEDKEYWORDALIASES in ClickHouseSqlUtils.java and ENGINETOTABLETYPE in DatabaseMetaDataImpl.java from failing test output. | ClickHouse/ | 1.6k | — | ~457 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 88 | 88.Testing Run tests and add Next.js version coverage for the cache handler. | trieb-work/ | 151 | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 89 | Runs a reproduce, localize, hypothesize, test, fix and verify loop to find a bug's root cause, applies the minimal fix and hands off a regression test. | jsmastery-pro/ | 1.5k | — | ~1.8k | Automated safety check: Notes | MIT | 2 mo ago |
| 90 | Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y. | chongdashu/ | 149 | — | ~2.1k | Automated safety check: Pass | No licence | 5 mo ago |
| 91 | 91.Code Solving Structured coding workflow for non-trivial code work: debug, build features, refactor, optimize, migrate and review code through 7 steps with evidence-based quality gates. | HoangTheQuyen/ | 122 | — | ~3.7k | Automated safety check: Pass | MIT | 2 days ago |
| 92 | Adds dotnet/maui-specific context for investigating failing PR checks and broken nightly builds: pipelines, Helix logs, binlogs and merge-readiness verdicts. | dotnet/ | 23k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 93 | 93.Fix CI Run a local pnpm monorepo CI loop, fix failures, and stop only when the full sequence is green. | gronxb/ | 1.8k | — | ~572 | Automated safety check: Pass | Unknown | today |
| 94 | 94.Review Reviews a GitHub pull request for correctness, architecture, security, backward compatibility, and test coverage. | webern/ | 385 | — | ~2k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 95 | 95.Slow Tests Run the gated test classes that plain pytest skips (slow Docker/sandbox tests, live model-provider API tests, flaky tests, trio variants). | UKGovernmentBEIS/ | 3k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 96 | A skill your agent uses when writing or modifying UE automated tests (Automation, CQTest, Functional, Gauntlet, LowLevel) with Rider MCP available. | JasonMa0012/ | 750 | — | ~2.1k | Automated safety check: Notes | Unknown | 23 days ago |