Debug Playwright Prow
quay/quay
Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…
Finds flaky tests in CI logs by aggregating pass rates, explains likely causes and recommends whether to quarantine or fix each one.
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Donchitos/Claude-Code-Game-Studios test-flakiness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/test-flakiness .claude/skills/test-flakiness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-flakiness" agent skill from https://github.com/Donchitos/Claude-Code-Game-Studios/tree/main/.claude/skills/test-flakiness into .claude/skills/test-flakiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-flakiness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Donchitos/Claude-Code-Game-Studios/tree/main/.claude/skills/test-flakinessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Donchitos/Claude-Code-Game-Studios test-flakiness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/test-flakiness .agents/skills/test-flakiness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-flakiness" agent skill from https://github.com/Donchitos/Claude-Code-Game-Studios/tree/main/.claude/skills/test-flakiness into .agents/skills/test-flakiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-flakiness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Donchitos/Claude-Code-Game-Studios test-flakiness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/test-flakiness .cursor/skills/test-flakiness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-flakiness" agent skill from https://github.com/Donchitos/Claude-Code-Game-Studios/tree/main/.claude/skills/test-flakiness into .cursor/skills/test-flakiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-flakiness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Donchitos/Claude-Code-Game-Studios.git --path .claude/skills/test-flakiness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Donchitos/Claude-Code-Game-Studios test-flakiness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/test-flakiness .gemini/skills/test-flakiness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-flakiness" agent skill from https://github.com/Donchitos/Claude-Code-Game-Studios/tree/main/.claude/skills/test-flakiness into .gemini/skills/test-flakiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-flakiness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Donchitos/Claude-Code-Game-Studios test-flakinessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/test-flakiness .github/skills/test-flakiness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-flakiness" agent skill from https://github.com/Donchitos/Claude-Code-Game-Studios/tree/main/.claude/skills/test-flakiness into .github/skills/test-flakiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-flakiness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Donchitos/Claude-Code-Game-Studios test-flakiness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/test-flakiness .opencode/skills/test-flakiness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-flakiness" agent skill from https://github.com/Donchitos/Claude-Code-Game-Studios/tree/main/.claude/skills/test-flakiness into .opencode/skills/test-flakiness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-flakiness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-flakinessFinds flaky tests in CI logs by aggregating pass rates, explains likely causes and recommends whether to quarantine or fix each one.
A flaky test passes and fails without any code change, which teaches a team to ignore red CI runs. The skill reads CI results to spot intermittent failures, works out pass rates across runs, explains probable causes and recommends quarantine or a fix for each test. Modes are a path to one log file, `scan` for all logs under `.github/` and standard output folders, `registry` for guidance on already quarantined tests, or no argument, which picks scan when logs exist and registry otherwise.
It knows where common engines leave results: GitHub Actions artifacts and `test-results/`, GdUnit4 JUnit-style XML under `reports/` for Godot, NUnit XML from the game-ci runner for Unity, and automation logs in `Saved/Logs/` for Unreal. Findings update the quarantine section of `tests/regression-suite.md`, with an optional dated flakiness report. It works best late in a project, such as the polish phase, when enough runs exist for a reliable signal.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b21fa0f. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGlobGrepWriteEditBashBash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helper.sh" resolve_config *)From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
bashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Flaky Test Detector loads about 2.6k tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 1,211 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Glob, Grep, Write, Edit, Bash, Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Donchitos/Claude-Code-Game-Studios at commit b21fa0f, republished under its MIT licence (© Donchitos). 1,211 words, ~2,603 tokens.
.claude/skills/test-flakiness/SKILL.md (or your agent's skills folder).!bash "${CLAUDE_SKILL_DIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation
Automation mode: Resolve modes.automation (project.local.yaml →
project.yaml → default collaborative). Every AskUserQuestion call and
every file write follows .claude/docs/automation-modes.md
(collaborative asks always · guided major-only · autonomous logs and proceeds;
automation_always_ask categories always prompt).
A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.
Output: Updated tests/regression-suite.md quarantine section + optional
production/qa/flakiness-report-[date].md
When to run:
/regression-suite identifies quarantined tests that need diagnosisModes:
/test-flakiness [ci-log-path] — analyse a specific CI run log file/test-flakiness scan — scan all available CI logs in .github/ or
standard log output directories/test-flakiness registry — read existing regression-suite.md quarantine
section and provide remediation guidance for already-known flaky testsscan if CI logs are accessible, else
registryCheck for test result artifacts:
ls -t .github/ 2>/dev/null
ls -t test-results/ 2>/dev/nullFor Godot projects: GdUnit4 outputs XML results compatible with JUnit format,
under reports/ (its default report folder, res://reports/). Check reports/,
and any test-results/ you saved runs into, for .xml files.
For Unity projects: game-ci test runner outputs NUnit XML to test-results/
by default.
For Unreal projects: automation logs go to Saved/Logs/. Grep for
Result={Success} and Result={Fail} — each test prints
Test Completed. Result={<status>}
(docs/engine-reference/unreal/current-best-practices.md, "Command Line").
If a path argument is provided, read that file directly.
If no logs found:
"No CI log data found. To detect flaky tests, this skill needs test result history from multiple runs. Options:
- Run the test suite at least 3 times and collect the output logs
- Check CI pipeline output and save a log to
test-results/- Run
/test-flakiness registryto review tests already flagged as flaky intests/regression-suite.md"
Stop and ask the user which option to pursue.
For each CI log or result file found, parse:
JUnit XML format (GdUnit4):
<testcase name= to get test names<failure or <error to identify failuresclassname and name attributes for full test identifiersNUnit XML format (Unity — the file whose <test-run> element /smoke-check
reads):
<test-case element; its fullname attribute is the identifierresult attribute is Passed, Failed, Inconclusive or Skipped
(docs/engine-reference/unity/current-best-practices.md, "Command Line");
only Passed and Failed enter the historyPlain text logs:
PASSED / FAILED adjacent to test namesResult={Success} / Result={Fail}Test passed / Test failedBuild a table: test_id → [run1_result, run2_result, run3_result, ...]
A test is flaky if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.
Flakiness thresholds:
With fewer than 3 runs, every finding is suspected, whatever its fail rate. One failure in two runs reads as 50%, but it is one data point: do not quarantine it, label it suspected, and ask whether more run data is available. The tiers above, and quarantine, apply from 3 runs up.
For each flaky test, classify the likely cause:
| Cause | Symptoms | Fix direction |
|---|---|---|
| Timing / async | Fails after awaiting signals or timers; pass rate correlates with system load | Add explicit await/synchronisation; avoid time-based delays |
| Order dependency | Fails when run after specific other tests; passes in isolation | Add proper setup/teardown; ensure test isolation |
| Random seed | Fails intermittently with no pattern; involves RNG | Pass explicit seed; don't use randf() in tests |
| Resource leak | Fails more often later in a test run | Fix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity) |
| External state | Fails when a file, scene, or global exists from a prior test | Isolate test from file system; use in-memory mocks |
| Floating point | Fails on comparisons like == 0.5 | Use epsilon comparison (is_equal_approx, Assert.AreApproximately) |
| Scene/prefab load race | Fails when scenes are not yet ready | Await one frame after instantiation; use await get_tree().process_frame |
Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.
For each flaky test:
Quarantine (High flakiness):
"Quarantine this test immediately. Skip it with the engine's own mechanism, log it in the
tests/regression-suite.mdquarantine section, and fix the root cause before removing quarantine."
Skip mechanisms, by engine — name only the one for the project's engine:
func test_x(_do_skip := true, _skip_reason := "flaky: [cause]"). gdUnit4 reads
the argument names do_skip and skip_reason (a leading _ is allowed) in
addons/gdUnit4/src/core/GdUnitTestSuiteScanner.gd, as of gdUnit4 6.1.3;
confirm them there for the installed version.addons/gdUnit4/ source, and docs/engine-reference/godot/ does not
cover them. Log it in the quarantine section and ask the user how their C#
tests are skipped; do not invent an attribute.[Ignore("flaky: [cause]")] — NUnit 3 requires the reasondocs/engine-reference/unreal/ documents no way
to skip an automation test. Log it in the quarantine section and ask the user
how their CI excludes a test; do not invent a flag.Investigate and fix soon (Moderate):
"This test is intermittently unreliable. Root cause appears to be [cause]. Suggested fix: [specific fix based on cause classification]. Do not quarantine yet — fix the test directly."
Monitor (Low/suspected):
"This test shows suspected flakiness. Collect more run data before quarantining. Note it as 'suspected' in the regression suite."
## Flakiness Detection Results
**Runs analysed**: [N]
**Tests tracked**: [N]
### Flaky Tests Found
| Test | System | Fail Rate | Confidence | Likely Cause | Recommendation |
|------|--------|-----------|------------|--------------|----------------|
| [test_name] | [system] | [N]% | confirmed | Timing | Quarantine + fix async |
| [test_name] | [system] | [N]% | confirmed | Float comparison | Fix: use epsilon compare |
| [test_name] | [system] | [N]% | suspected (fewer than 3 runs) | Order dependency | Collect more runs before acting |
### Clean Tests (no flakiness detected)
[N] tests ran across [N] runs with consistent results — no flakiness detected.
### Data Limitations
[Note if fewer than 5 runs were available — fewer runs = less statistical confidence]Ask: "May I update the quarantine section of tests/regression-suite.md
with the flaky tests found?"
If yes: use Edit to append entries to the Quarantined Tests table.
Never remove existing quarantine entries — only add new ones.
Ask (separately): "May I write a full flakiness report to
production/qa/flakiness-report-[date].md?"
The full report includes per-test analysis with cause details and engine-specific fix snippets.
After writing:
is_equal_approx."© Donchitos, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/test-flakiness of Donchitos/Claude-Code-Game-Studios.
Open the folder on GitHubat commit b21fa0f
Flaky Test Detector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Flaky Test Detector this skillDonchitos/Claude-Code-Game-Studios | 26k | — | ~2.6k | Automated safety check: Notes | MIT | |
| Debug Playwright Prowquay/quay | 2.8k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb | 6.7k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| Debug Playwrightquay/quay | 2.8k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| CI Triagecanton-network/splice | 118 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Babysit PRZenUml/web-sequence | 150 | — | ~871 | Automated safety check: Pass | MIT |
quay/quay
Deep-dive diagnosis of a Playwright test failure already isolated to one Quay Prow/OpenShift CI run: downloads its GCS artifacts (results.json, JUnit, build/pod logs, Jaeger traces), classifies real…
GreptimeTeam/greptimedb
Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.
quay/quay
Debug Playwright E2E test failures from GitHub Actions CI runs.
canton-network/splice
Triage a failed splice GitHub Actions job (cn-test-failures ref) into a reproducible evidence packet - fetch job log and artifact, isolate the flagged lines, check the known flake families for…
ZenUml/web-sequence
Monitor and diagnose GitHub Actions checks on ZenUML web-sequence PRs, fixing code-caused CI failures when appropriate.
cognesy/instructor-php
Run multi-dimensional quality assurance for InstructorPHP. An agent skill from cognesy/instructor-php.
Donchitos/Claude-Code-Game-Studios
Audits game assets against naming conventions, file size budgets and format standards, and finds orphaned assets and missing references.
Donchitos/Claude-Code-Game-Studios
Writes per-asset visual specs and AI image-generation prompts for a game's characters, enemies and screens, driven by the GDD, art bible and an entity inventory.
Donchitos/Claude-Code-Game-Studios
Checks game data and formulas for balance outliers, broken progression, degenerate strategies and economy problems, and answers 'could not run' when the data is missing.
Donchitos/Claude-Code-Game-Studios
Turns a description into a structured bug report, or scans code for likely bugs, then verifies and closes reports through four modes.
Donchitos/Claude-Code-Game-Studios
Reviews the open bug backlog, separates severity from priority, assigns fixes to sprints and reports systemic trends, writing a dated triage file.
Donchitos/Claude-Code-Game-Studios
Generates an internal or player-facing changelog from git commits and sprint data, filtering out framework maintenance commits so that only work on the game itself reaches release copy.
Works with
Categories
Finds flaky tests in CI logs by aggregating pass rates, explains likely causes and recommends whether to quarantine or fix each one. A flaky test passes and fails without any code change, which teaches a team to ignore red CI runs. The skill reads CI results to spot intermittent failures, works out pass rates across runs, explains probable causes and recommends quarantine or a fix for each test.
Flaky Test Detector fits situations like: developers have started dismissing CI failures as probably flaky; diagnosing tests already quarantined in the regression suite; aggregating pass rates over many CI runs during the polish phase.
Run `npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a claude-code`. Or copy the skill folder (.claude/skills/test-flakiness in Donchitos/Claude-Code-Game-Studios) into .claude/skills/test-flakiness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a codex`. Or copy the skill folder (.claude/skills/test-flakiness in Donchitos/Claude-Code-Game-Studios) into .agents/skills/test-flakiness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-flakiness, .gemini/skills/test-flakiness, .github/skills/test-flakiness and .opencode/skills/test-flakiness in your project.
Going by SKILL.md and its folder, Flaky Test Detector needs the command-line tools its instructions call (bash). Our summary lists: CI logs or test result files from multiple runs. Its frontmatter pre-approves these tools: Read, Glob, Grep, Write, Edit, Bash, Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helper.sh" resolve_config *).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Flaky Test Detector is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Flaky Test Detector: Debug Playwright Prow (quay/quay, 2.8k stars), GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), Debug Playwright (quay/quay, 2.8k stars) and CI Triage (canton-network/splice, 118 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Donchitos (a GitHub user) maintains it in Donchitos/Claude-Code-Game-Studios, which has 25,871 GitHub stars. The repository holds 73 skills in this directory. The repository was last updated on September 29, 2026.
Source: Donchitos/Claude-Code-Game-Studios on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.