Agent skill

Flaky Test Detector

by Donchitos in Donchitos/Claude-Code-Game-Studios

Finds flaky tests in CI logs by aggregating pass rates, explains likely causes and recommends whether to quarantine or fix each one.

MITAuto-check: notesTesting & QA

Install Flaky Test Detector

skills CLI
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Donchitos/Claude-Code-Game-Studios test-flakiness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-flakiness .claude/skills/test-flakiness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-flakiness
GitHub stars
26k
Token cost
~2.6k tokens
SKILL.md length
1,211 words
Files
1
Skills in repo
73
Repo updated
First seen
Licence
MIT

At a glance

Finds flaky tests in CI logs by aggregating pass rates, explains likely causes and recommends whether to quarantine or fix each one.

  • Works in 7 steps: Parse Arguments → Locate CI Log Data → Parse Test Results → …
  • Developers have started dismissing CI failures as probably flaky
  • SKILL.md covers 1. Parse Arguments, 2. Locate CI Log Data, 3. Parse Test Results and 4. Identify Flaky Tests, plus 4 more sections
  • Calls bash

What it does

A flaky test passes and fails without any code change, which teaches a team to ignore red CI runs. The skill reads CI results to spot intermittent failures, works out pass rates across runs, explains probable causes and recommends quarantine or a fix for each test. Modes are a path to one log file, `scan` for all logs under `.github/` and standard output folders, `registry` for guidance on already quarantined tests, or no argument, which picks scan when logs exist and registry otherwise.

It knows where common engines leave results: GitHub Actions artifacts and `test-results/`, GdUnit4 JUnit-style XML under `reports/` for Godot, NUnit XML from the game-ci runner for Unity, and automation logs in `Saved/Logs/` for Unreal. Findings update the quarantine section of `tests/regression-suite.md`, with an optional dated flakiness report. It works best late in a project, such as the polish phase, when enough runs exist for a reliable signal.

When your agent uses it

  • Developers have started dismissing CI failures as probably flaky
  • Diagnosing tests already quarantined in the regression suite
  • Aggregating pass rates over many CI runs during the polish phase

Example prompts

  • “Scan our CI logs and list the flaky tests with their pass rates.”
  • “Analyze test-results/run-latest.xml for intermittent failures.”
  • “Give me remediation advice for the tests quarantined in the regression suite.”

Requirements

  • CI logs or test result files from multiple runs
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, Write, Edit, Bash, Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helper.sh" resolve_config *)

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Parse Arguments
  2. Locate CI Log Data
  3. Parse Test Results
  4. Identify Flaky Tests
  5. Recommend Action
  6. Generate Reports
  7. Update Regression Suite + Optional Report File

What it can do on your machine

Read from SKILL.md and the folder at commit b21fa0f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • Write
    • Edit
    • Bash
    • Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helper.sh" resolve_config *)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flaky Test Detector loads about 2.6k tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 1,211 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Glob, Grep, Write, Edit, Bash, Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-hel

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Donchitos/Claude-Code-Game-Studios at commit b21fa0f, republished under its MIT licence (© Donchitos). 1,211 words, ~2,603 tokens.

Download SKILL.mdSave it as .claude/skills/test-flakiness/SKILL.md (or your agent's skills folder).
name
test-flakiness
description
Find flaky tests from CI logs — aggregates pass rates, spots intermittent failures, recommends quarantine. After multiple runs.
allowed-tools
Read, Glob, Grep, Write, Edit, Bash, Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helper.sh" resolve_config *)
argument-hint
[ci-log-path | scan | registry]
user-invocable
true
model
sonnet

!bash "${CLAUDE_SKILL_DIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation

Automation mode: Resolve modes.automation (project.local.yaml → project.yaml → default collaborative). Every AskUserQuestion call and every file write follows .claude/docs/automation-modes.md (collaborative asks always · guided major-only · autonomous logs and proceeds; automation_always_ask categories always prompt).

Test Flakiness Detection

A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.

Output: Updated tests/regression-suite.md quarantine section + optional production/qa/flakiness-report-[date].md

When to run:

  • Polish phase (tests have had many runs; statistical signal is reliable)
  • When developers start dismissing CI failures as "probably flaky"
  • After /regression-suite identifies quarantined tests that need diagnosis

1. Parse Arguments

Modes:

  • /test-flakiness [ci-log-path] — analyse a specific CI run log file
  • /test-flakiness scan — scan all available CI logs in .github/ or standard log output directories
  • /test-flakiness registry — read existing regression-suite.md quarantine section and provide remediation guidance for already-known flaky tests
  • No argument — auto-detect: run scan if CI logs are accessible, else registry

2. Locate CI Log Data

Option A — GitHub Actions (preferred)

Check for test result artifacts:

bash
ls -t .github/ 2>/dev/null
ls -t test-results/ 2>/dev/null

For Godot projects: GdUnit4 outputs XML results compatible with JUnit format, under reports/ (its default report folder, res://reports/). Check reports/, and any test-results/ you saved runs into, for .xml files.

For Unity projects: game-ci test runner outputs NUnit XML to test-results/ by default.

For Unreal projects: automation logs go to Saved/Logs/. Grep for Result={Success} and Result={Fail} — each test prints Test Completed. Result={<status>} (docs/engine-reference/unreal/current-best-practices.md, "Command Line").

Option B — Local log files

If a path argument is provided, read that file directly.

Option C — No log data available

If no logs found:

"No CI log data found. To detect flaky tests, this skill needs test result history from multiple runs. Options:

  1. Run the test suite at least 3 times and collect the output logs
  2. Check CI pipeline output and save a log to test-results/
  3. Run /test-flakiness registry to review tests already flagged as flaky in tests/regression-suite.md"

Stop and ask the user which option to pursue.


3. Parse Test Results

For each CI log or result file found, parse:

JUnit XML format (GdUnit4):

  • Grep for <testcase name= to get test names
  • Grep for <failure or <error to identify failures
  • Parse classname and name attributes for full test identifiers

NUnit XML format (Unity — the file whose <test-run> element /smoke-check reads):

  • Each test is a <test-case element; its fullname attribute is the identifier
  • Its result attribute is Passed, Failed, Inconclusive or Skipped (docs/engine-reference/unity/current-best-practices.md, "Command Line"); only Passed and Failed enter the history

Plain text logs:

  • Grep for pass/fail patterns:
    • Godot: PASSED / FAILED adjacent to test names
    • Unreal: Result={Success} / Result={Fail}
    • Unity: Test passed / Test failed

Build a table: test_id → [run1_result, run2_result, run3_result, ...]


4. Identify Flaky Tests

A test is flaky if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.

Flakiness thresholds:

  • High flakiness: Fails in >25% of runs — quarantine immediately
  • Moderate flakiness: Fails in 5–25% of runs — investigate and fix soon
  • Low/suspected flakiness: Fails in 1–5% of runs — monitor; may be genuinely rare failure

With fewer than 3 runs, every finding is suspected, whatever its fail rate. One failure in two runs reads as 50%, but it is one data point: do not quarantine it, label it suspected, and ask whether more run data is available. The tiers above, and quarantine, apply from 3 runs up.

For each flaky test, classify the likely cause:

Cause classification
CauseSymptomsFix direction
Timing / asyncFails after awaiting signals or timers; pass rate correlates with system loadAdd explicit await/synchronisation; avoid time-based delays
Order dependencyFails when run after specific other tests; passes in isolationAdd proper setup/teardown; ensure test isolation
Random seedFails intermittently with no pattern; involves RNGPass explicit seed; don't use randf() in tests
Resource leakFails more often later in a test runFix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity)
External stateFails when a file, scene, or global exists from a prior testIsolate test from file system; use in-memory mocks
Floating pointFails on comparisons like == 0.5Use epsilon comparison (is_equal_approx, Assert.AreApproximately)
Scene/prefab load raceFails when scenes are not yet readyAwait one frame after instantiation; use await get_tree().process_frame

Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.


Show full SKILL.md (461 more words)Show less

5. Recommend Action

For each flaky test:

Quarantine (High flakiness):

"Quarantine this test immediately. Skip it with the engine's own mechanism, log it in the tests/regression-suite.md quarantine section, and fix the root cause before removing quarantine."

Skip mechanisms, by engine — name only the one for the project's engine:

  • Godot (gdUnit4), GDScript: a skip parameter on the test function — func test_x(_do_skip := true, _skip_reason := "flaky: [cause]"). gdUnit4 reads the argument names do_skip and skip_reason (a leading _ is allowed) in addons/gdUnit4/src/core/GdUnitTestSuiteScanner.gd, as of gdUnit4 6.1.3; confirm them there for the installed version.
  • Godot (gdUnit4), C#: NOT SOURCEABLE — gdUnit4's C# test attributes are not in the addons/gdUnit4/ source, and docs/engine-reference/godot/ does not cover them. Log it in the quarantine section and ask the user how their C# tests are skipped; do not invent an attribute.
  • Unity (NUnit): [Ignore("flaky: [cause]")] — NUnit 3 requires the reason
  • Unreal: NOT SOURCEABLE — docs/engine-reference/unreal/ documents no way to skip an automation test. Log it in the quarantine section and ask the user how their CI excludes a test; do not invent a flag.

Investigate and fix soon (Moderate):

"This test is intermittently unreliable. Root cause appears to be [cause]. Suggested fix: [specific fix based on cause classification]. Do not quarantine yet — fix the test directly."

Monitor (Low/suspected):

"This test shows suspected flakiness. Collect more run data before quarantining. Note it as 'suspected' in the regression suite."


6. Generate Reports

In-conversation summary
## Flakiness Detection Results

**Runs analysed**: [N]
**Tests tracked**: [N]

### Flaky Tests Found

| Test | System | Fail Rate | Confidence | Likely Cause | Recommendation |
|------|--------|-----------|------------|--------------|----------------|
| [test_name] | [system] | [N]% | confirmed | Timing | Quarantine + fix async |
| [test_name] | [system] | [N]% | confirmed | Float comparison | Fix: use epsilon compare |
| [test_name] | [system] | [N]% | suspected (fewer than 3 runs) | Order dependency | Collect more runs before acting |

### Clean Tests (no flakiness detected)

[N] tests ran across [N] runs with consistent results — no flakiness detected.

### Data Limitations

[Note if fewer than 5 runs were available — fewer runs = less statistical confidence]

7. Update Regression Suite + Optional Report File

Ask: "May I update the quarantine section of tests/regression-suite.md with the flaky tests found?"

If yes: use Edit to append entries to the Quarantined Tests table. Never remove existing quarantine entries — only add new ones.

Ask (separately): "May I write a full flakiness report to production/qa/flakiness-report-[date].md?"

The full report includes per-test analysis with cause details and engine-specific fix snippets.

After writing:

  • For each quarantined test: "Add the engine-specific skip annotation to disable this test in CI. Re-enable after the root cause is fixed."
  • For fix-eligible tests: "The fix for [test] is straightforward — change the equality comparison on line [N] to use is_equal_approx."
  • Summary: "Once all quarantine annotations are applied, CI should run green. Schedule fix work for the [N] quarantined tests before the release gate."

Collaborative Protocol

  • Never delete test files — quarantine means annotate + list, not remove
  • Statistical confidence matters — with < 3 runs, flag findings as "suspected" not "confirmed"; ask if more run data is available
  • Fix is always the goal — quarantine is temporary; surface the fix direction even when recommending quarantine
  • Ask before writing — both the regression-suite update and the report file require explicit approval. On write: Verdict: COMPLETE — flakiness report written. On decline: Verdict: BLOCKED — user declined write.
  • Flakiness in CI is a team problem — surface the list and recommended actions clearly; do not just silently quarantine without the team knowing

© Donchitos, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/test-flakiness of Donchitos/Claude-Code-Game-Studios.

Open the folder on GitHubat commit b21fa0f

Compare with similar skills

Flaky Test Detector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flaky Test Detector compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flaky Test Detector this skillDonchitos/Claude-Code-Game-Studios26k—~2.6kAutomated safety check: NotesMIT
Debug Playwright Prowquay/quay2.8k—~2.2kAutomated safety check: PassApache-2.0
GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb6.7k—~4.4kAutomated safety check: PassApache-2.0
Debug Playwrightquay/quay2.8k—~1.2kAutomated safety check: PassApache-2.0
CI Triagecanton-network/splice118—~1.4kAutomated safety check: PassApache-2.0
Babysit PRZenUml/web-sequence150—~871Automated safety check: PassMIT

Similar skills

  • Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…

    2.8k GitHub stars~2.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Debug Playwright E2E test failures from GitHub Actions CI runs.

    2.8k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • CI Triage

    canton-network/splice

    Triage a failed splice GitHub Actions job (cn-test-failures ref) into a reproducible evidence packet - fetch job log and artifact, isolate the flagged lines, check the known flake families for…

    118 GitHub stars~1.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Babysit PR

    ZenUml/web-sequence

    Monitor and diagnose GitHub Actions checks on ZenUML web-sequence PRs, fixing code-caused CI failures when appropriate.

    150 GitHub stars~871 tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Instructor QA

    cognesy/instructor-php

    Run multi-dimensional quality assurance for InstructorPHP. An agent skill from cognesy/instructor-php.

    328 GitHub stars~1.3k tokensUpdated 2 days ago
    Testing & QAAuto-check passed

More from Donchitos/Claude-Code-Game-Studios

All 73 skills in this repo
  • Game Asset Audit

    Donchitos/Claude-Code-Game-Studios

    Audits game assets against naming conventions, file size budgets and format standards, and finds orphaned assets and missing references.

    26k GitHub stars~2k tokensUpdated 9 days ago
    Auto-check passed
  • Game Asset Spec Writer

    Donchitos/Claude-Code-Game-Studios

    Writes per-asset visual specs and AI image-generation prompts for a game's characters, enemies and screens, driven by the GDD, art bible and an entity inventory.

    26k GitHub stars~5k tokensUpdated 9 days ago
    Auto-check passed
  • Game Balance Check

    Donchitos/Claude-Code-Game-Studios

    Checks game data and formulas for balance outliers, broken progression, degenerate strategies and economy problems, and answers 'could not run' when the data is missing.

    26k GitHub stars~2.2k tokensUpdated 9 days ago
    Auto-check passed
  • Structured Bug Reports

    Donchitos/Claude-Code-Game-Studios

    Turns a description into a structured bug report, or scans code for likely bugs, then verifies and closes reports through four modes.

    26k GitHub stars~2.5k tokensUpdated 9 days ago
    Auto-check: notes
  • Bug Triage

    Donchitos/Claude-Code-Game-Studios

    Reviews the open bug backlog, separates severity from priority, assigns fixes to sprints and reports systemic trends, writing a dated triage file.

    26k GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed
  • Changelog Generator for Games

    Donchitos/Claude-Code-Game-Studios

    Generates an internal or player-facing changelog from git commits and sprint data, filtering out framework maintenance commits so that only work on the game itself reaches release copy.

    26k GitHub stars~2.5k tokensUpdated 9 days ago
    Auto-check: notes

Categories

Questions about Flaky Test Detector

What does Flaky Test Detector do?

Finds flaky tests in CI logs by aggregating pass rates, explains likely causes and recommends whether to quarantine or fix each one. A flaky test passes and fails without any code change, which teaches a team to ignore red CI runs. The skill reads CI results to spot intermittent failures, works out pass rates across runs, explains probable causes and recommends quarantine or a fix for each test.

When should I use Flaky Test Detector?

Flaky Test Detector fits situations like: developers have started dismissing CI failures as probably flaky; diagnosing tests already quarantined in the regression suite; aggregating pass rates over many CI runs during the polish phase.

How do I install Flaky Test Detector in Claude Code?

Run `npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a claude-code`. Or copy the skill folder (.claude/skills/test-flakiness in Donchitos/Claude-Code-Game-Studios) into .claude/skills/test-flakiness in your project. Claude Code loads it when a task matches its description.

How do I install Flaky Test Detector in Codex?

Run `npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a codex`. Or copy the skill folder (.claude/skills/test-flakiness in Donchitos/Claude-Code-Game-Studios) into .agents/skills/test-flakiness in your project. Codex loads it when a task matches its description.

Can I use Flaky Test Detector in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-flakiness, .gemini/skills/test-flakiness, .github/skills/test-flakiness and .opencode/skills/test-flakiness in your project.

What does Flaky Test Detector need to run?

Going by SKILL.md and its folder, Flaky Test Detector needs the command-line tools its instructions call (bash). Our summary lists: CI logs or test result files from multiple runs. Its frontmatter pre-approves these tools: Read, Glob, Grep, Write, Edit, Bash, Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helper.sh" resolve_config *).

Does Flaky Test Detector access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Flaky Test Detector safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Flaky Test Detector use?

Flaky Test Detector is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flaky Test Detector use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flaky Test Detector?

Skills that share tags, products or a category with Flaky Test Detector: Debug Playwright Prow (quay/quay, 2.8k stars), GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), Debug Playwright (quay/quay, 2.8k stars) and CI Triage (canton-network/splice, 118 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flaky Test Detector?

Donchitos (a GitHub user) maintains it in Donchitos/Claude-Code-Game-Studios, which has 25,871 GitHub stars. The repository holds 73 skills in this directory. The repository was last updated on September 29, 2026.

Source: Donchitos/Claude-Code-Game-Studios on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.