Search
Failing and flaky tests
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way. | openinterpreter/ | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations. | PowerShell/ | 56k | — | ~5.1k | Automated safety check: Pass | MIT | today |
| 3 | Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions. | langgenius/ | 158k | — | ~682 | Automated safety check: Pass | Unknown | today |
| 4 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes | ultralisp/ | 258 | 51 repos | ~2.4k | Automated safety check: Pass | No licence | 26 days ago |
| 5 | Evaluate ClickHouse performance test results from existing CI/dashboard data or local perf.py runs. | ClickHouse/ | 50k | — | ~3.9k | Automated safety check: Notes | Apache-2.0 | today |
| 6 | Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report. | appsmithorg/ | 41k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 7 | Log genuine, recurring repository friction to .agents/PAPERCUTS.md — confusing setup, a flaky repo command or script, a misleading in-repo error, stale generated files, or a non-obvious gotcha that… | every-app/ | 23k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 8 | Runs RustPython tests inside a Linux container built with Apple's container CLI, so macOS users can compare Linux results with their local ones. | RustPython/ | 22k | — | ~467 | Automated safety check: Pass | MIT | yesterday |
| 9 | Fixes failing Playwright specs by reading the error, classifying the cause in the test code and applying corrections that follow project conventions. | appsmithorg/ | 41k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | 10.Iterate PR Iterate on a PR until CI passes. An agent skill from meshery/meshery-operator. | meshery/ | 151 | 7 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 11 | Takes a GitHub or YouTrack issue for the Exposed project through reproduction, a failing test, a fix, validation and a pull request. | JetBrains/ | 9.3k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Downloads Azure Pipelines CI logs for an Ansible pull request or build so the agent can analyze test failures, after asking you first. | ansible/ | 71k | — | ~825 | Automated safety check: Pass | GPL-3.0 | yesterday |
| 13 | Inspect, analyze, troubleshoot, or review Codacy findings and local analyzer/API helpers. | netdata/ | 81k | — | ~2.2k | Automated safety check: Notes | GPL-3.0 | today |
| 14 | Handle a Perfherder performance regression bug end to end: read the alert bug, confirm whether the regression is real, find the cause, and iterate to a fix. | mozilla-firefox/ | 13k | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 15 | 15.Update V86 Build and install v86 (wasm + libv86.js + BIOS) into windows95. | felixrieseberg/ | 24k | — | ~1.7k | Automated safety check: Pass | Unknown | 28 days ago |
| 16 | Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real… | quay/ | 2.8k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent. | Chachamaru127/ | 3.2k | 1 repo | ~1.1k | Automated safety check: Notes | MIT | 4 days ago |
| 18 | Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated. | pingcap/ | 41k | — | ~498 | Automated safety check: Pass | Apache-2.0 | today |
| 19 | 19.Swig Test Run SWIG test suite for specific languages. An agent skill from swig/swig. | swig/ | 6.3k | — | ~2.3k | Automated safety check: Pass | Unknown | 2 days ago |
| 20 | A skill your agent uses when CI tests fail on main branch after PR merge, when investigating flaky test failures, or when user provides a PR URL/number to aggregate all failing tests | payloadcms/ | 45k | — | ~4.4k | Automated safety check: Pass | MIT | yesterday |
| 21 | Fixes a React Router bug reported in a GitHub issue end to end: fetching the issue, validating the reproduction, writing a failing test and implementing the fix on a new branch. | remix-run/ | 57k | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 22 | 22.Testing A skill your agent uses for every Kortix test task, behavior change, bug fix, refactor, API route change, CLI change, SDK change, browser journey, test failure, coverage question, local benchmark… | kortix-ai/ | 20k | — | ~3.6k | Automated safety check: Notes | Unknown | today |
| 23 | Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written. | ChrisWiles/ | 6.1k | 3 repos | ~1.2k | Automated safety check: Pass | No licence | 9 mo ago |
| 24 | Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses. | addyosmani/ | 103k | 1 repo | ~2.6k | Automated safety check: Pass | MIT | 6 days ago |
| 25 | Classify a failed CI as either caused by an active incident, flakiness, or a true code regression. | DataDog/ | 3.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code. | GreptimeTeam/ | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | Plans the smallest check that could disprove a code change in the OpenLogi project, then escalates through reproduction, focused tests and a final gate before a push. | AprilNEA/ | 23k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 28 | Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests. | comet-ml/ | 22k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 29 | Upgrades a Python standard library module from CPython into RustPython with update_lib, then triages and marks the tests that still fail. | RustPython/ | 22k | — | ~876 | Automated safety check: Pass | MIT | yesterday |
| 30 | 30.CI Fix Scan all CI builds and tests, find failures, fetch error logs, and fix the code. | FastLED/ | 7.5k | — | ~897 | Automated safety check: Pass | MIT | today |
| 31 | Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 32 | 32.Fix Issue Fix a reported issue in Remix from a GitHub issue. An agent skill from remix-run/remix. | remix-run/ | 33k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 33 | Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm. | guqiong96/ | 465 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | 17 days ago |
| 34 | Diagnose any Quay Prow job failure end to end: prowjob.json - top-level build log - JUnit - resolved failing step - Playwright results.json when the failing step is Playwright, continuing through… | quay/ | 2.8k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior during spec-superflow execution, before proposing fixes. | MageByte-Zero/ | 842 | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 8 days ago |
| 36 | Debug and verification workflow for runtime-bundle and module-resolution regressions. | vercel/ | 143k | 1 repo | ~618 | Automated safety check: Pass | MIT | today |
| 37 | Analyze Vortex GitHub Actions CI failures. An agent skill from vortex-data/vortex. | vortex-data/ | 3.2k | — | ~810 | Automated safety check: Pass | Apache-2.0 | today |
| 38 | Triggers, re-runs and unblocks the CI checks on an ONNX Runtime pull request, after diagnosing whether a failure is transient or needs a code change. | microsoft/ | 22k | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 39 | 39.CI Watchdog Continuously monitor GitHub PR CI checks and automatically fix failures until all checks pass. | latitude-dev/ | 4.7k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 40 | Runs a gated finish-line checklist before committing a PlotJuggler PJ4 change: build proof, red-test triage, hooks, docs freshness and a diff self-review. | PlotJuggler/ | 6.2k | — | ~1.3k | Automated safety check: Pass | MPL-2.0 | 8 days ago |
| 41 | Investigate and triage CI failures for dotnet/macios from Azure DevOps build URLs. | dotnet/ | 2.9k | — | ~2.3k | Automated safety check: Pass | Unknown | today |
| 42 | Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 43 | Reproduce a GitHub Actions Linux CI failure locally when it does not happen on your machine: a podman/docker image that mirrors the ubuntu-22.04 runner by reusing the real Tools/CI-linux-.sh install… | swig/ | 6.3k | — | ~1.2k | Automated safety check: Pass | Unknown | 2 days ago |
| 44 | Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests. | DynamoDS/ | 2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 45 | Run and debug Agmente iOS end-to-end tests against a real local Codex CLI app-server instance. | rebornix/ | 545 | — | ~625 | Automated safety check: Pass | MIT | 4 mo ago |
| 46 | 46.PR Review Address review comments and CI failures for the current branch's PR | wysaid/ | 1.9k | — | ~1.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 47 | A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes - four-phase framework (root cause investigation, pattern analysis, hypothesis… | ed3dai/ | 250 | 3 repos | ~2.4k | Automated safety check: Pass | No licence | 1 mo ago |
| 48 | Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala). | apache/ | 4.6k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |