Agent skill

Detect Flaky Tests

by agent-substrate in agent-substrate/substrate

Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM…

Apache-2.0Auto-check passedTesting & QA

Install Detect Flaky Tests

skills CLI
$ npx skills add agent-substrate/substrate --skill detect-flaky-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agent-substrate/substrate detect-flaky-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agent-substrate/substrate.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/detect-flaky-tests .claude/skills/detect-flaky-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
detect-flaky-tests
GitHub stars
4.5k
Token cost
~3k tokens
SKILL.md length
1,181 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM…

  • Works in 6 steps: Collect workflow run IDs (last 7 days) → Download and parse logs: two jobs, three… → Infra issue triage (do this BEFORE… → …
  • Tasks that involve Failing and flaky tests
  • SKILL.md covers Flakiness threshold (keeps…, Step 1 — Collect workflow run…, Step 2 — Download and parse… and Step 3 — Infra issue triage…, plus 6 more sections
  • Calls gh, go and kind

What it does

Detect Flaky Tests is an agent skill from agent-substrate/substrate. Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM lanes). For each newly-detected flaky test or infra issue, opens a GitHub issue with full evidence and a draft fix PR. Does not touch BigQuery, dashboards, or any external storage — those are cron-job concerns layered on top.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests and End-to-end testing. It works with Google BigQuery, GitHub and GitHub Actions. The repository describes itself as: Agent Substrate: the core system. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Failing and flaky tests
  • Tasks that involve End-to-end testing

Example prompts

  • “Use the detect-flaky-tests skill to detect flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering…”
  • “/detect-flaky-tests”

Requirements

  • Docker

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Collect workflow run IDs (last 7 days)
  2. Download and parse logs: two jobs, three lanes per run
  3. Infra issue triage (do this BEFORE aggregating flakiness)
  4. Aggregate test flakiness across runs
  5. Open a draft fix PR for each new flaky test
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 0b91488. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • go
    • kind

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Detect Flaky Tests loads about 3k tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 1,181 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agent-substrate/substrate at commit 0b91488, republished under its Apache-2.0 licence (© agent-substrate). 1,181 words, ~3,003 tokens.

Download SKILL.mdSave it as .claude/skills/detect-flaky-tests/SKILL.md (or your agent's skills folder).
name
detect-flaky-tests
description
Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM lanes). For each newly-detected flaky test or infra issue, opens a GitHub issue with full evidence and a draft fix PR. Does not touch BigQuery, dashboards, or any external storage — those are cron-job concerns layered on top.

Detect Flaky Tests

A test is flaky when it produces both PASS and FAIL outcomes across multiple independent CI runs in the last 7 days, with no code change to that test's package explaining the inconsistency. Cross-PR analysis provides the strongest signal: if the same test fails on PR-A but passes on PR-B, that inconsistency is almost certainly non-determinism, not a legitimate regression.

Flakiness threshold (keeps false-positive rate low)

A test is flagged only when all three conditions hold in the 7-day window:

ConditionRationale
fail_count >= 2One failure could be infra noise
pass_count >= 2One pass could be a pre-fix lucky run
0.05 < fail_rate < 0.95Outside this band it is either reliably broken or reliably passing

Step 1 — Collect workflow run IDs (last 7 days)

bash
SINCE=$(date -u -v-7d +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || date -u -d '7 days ago' +%Y-%m-%dT%H:%M:%SZ)
gh api --paginate \
  "repos/agent-substrate/substrate/actions/workflows/pr-workflow.yaml/runs?status=completed&per_page=100&created=>=$SINCE" \
  --jq '.workflow_runs[] | {id: .id, conclusion: .conclusion, head_sha: .head_sha, created_at: .created_at}'

--paginate is required: a typical week has several hundred completed runs (600+ as of August 2026), far more than one page of 100.

Collect all run IDs. Process both successful and failed runs — both contain test output.


Step 2 — Download and parse logs: two jobs, three lanes per run

For each run, you need logs from two jobs:

Job nameCoverage
run-testsUnit + integration tests (go test -race -v ./...)
e2e-testE2E suite, both sandbox classes — two sequential steps in the one job
bash
# List all jobs for a run
gh api "repos/agent-substrate/substrate/actions/runs/<RUN_ID>/jobs" \
  --jq '.jobs[] | {id: .id, name: .name, conclusion: .conclusion}'

# Download log for a specific job
gh api "repos/agent-substrate/substrate/actions/jobs/<JOB_ID>/logs" > /tmp/job_<JOB_ID>.log

The two e2e lanes are NOT a matrix — e2e-test is a single job that runs the step "Run E2E tests (gVisor)" followed by "Run E2E tests (micro-VM)" (same hack/run-e2e-kind.sh command, the second with E2E_SANDBOX_CLASS: microvm). Split the one e2e-test job log at the "Run E2E tests (micro-VM)" step boundary: PASS/FAIL lines before it belong to the gVisor lane, lines after it to the microVM lane. Because the steps are sequential, a gVisor-lane failure means the micro-VM step never ran — record no microVM results for that run rather than counting them as failures.

Parse go test -v output from each log:

bash
grep -E '^--- (PASS|FAIL): ' /tmp/job_<JOB_ID>.log \
  | awk '{print $2, $3}' | sed 's/://'

Track results per (test_name, job_type) where job_type is one of: unit, e2e-gvisor, e2e-microvm.


Step 3 — Infra issue triage (do this BEFORE aggregating flakiness)

An infra failure is when the job itself breaks before or during setup — not when a test produces FAIL output. Infra failures must be identified and reported separately; they do NOT count toward a test's fail_count.

How to detect an infra failure

A job log is an infra failure (not a test failure) when it contains ANY of:

SignalExample log pattern
Go module proxy errorINTERNAL_ERROR, proxy.golang.org: dial, go mod download: ...500
kind cluster creation failureERROR: failed to create cluster, node(s) not ready, timed out waiting for the condition
Image pull failurefailed to pull image, ErrImagePull, ImagePullBackOff
Docker/containerd failurefailed to start containerd, Error response from daemon
OOM / out of diskOOMKilled, No space left on device
Network/DNS failure in setupdial tcp: lookup, connection refused during setup steps (not inside a test)
No test output at allLog ends before any --- PASS or --- FAIL line appears
Setup step non-zero exitA step before the go test or hack/run-e2e-kind.sh command fails

Critical rule: If a job log contains --- FAIL: TestFoo AND infra error patterns, you must determine which came first chronologically. If the infra error appears before the first test ran, treat the whole job as an infra failure (zero test results). If tests started running and then an infra error interrupted them mid-run, count only the completed test results and note the truncation.

Triple-check protocol

For any run where you suspect an infra issue, verify it three ways:

  1. Pattern match: does the log contain one of the signals above?
  2. Timeline check: does the error appear before === RUN Test lines, or between test completions?
  3. Cross-run consistency: did the same infra error appear in ≥2 other runs around the same time? If yes, it is definitely infrastructure. If it appeared only once, it may be a transient fluke — still classify as infra but flag lower confidence.
Infra issue aggregation

Collect infra failures separately:

infra_patternjob_typefail_countexample_run_ids
proxy.golang.org INTERNAL_ERRORunit4[run_1, run_2, ...]
kind cluster: node(s) not readye2e-gvisor2[run_5, ...]

An infra pattern that appears in ≥2 runs warrants a GitHub issue (Step 5b).


Show full SKILL.md (500 more words)Show less

Step 4 — Aggregate test flakiness across runs

Using only the runs NOT classified as infra failures, build a per-test table:

test_namejob_typefail_countpass_counttotal_runs

Treat each lane independently. A test that is flaky only in e2e-gvisor is still flagged — it does not need to be flaky in e2e-microvm too.

Apply the threshold: fail_count >= 2 AND pass_count >= 2 AND 0.05 < fail_rate < 0.95.


Step 5a — Create issues for flaky tests

Before creating, check for an existing open issue:

bash
gh issue list \
  --repo agent-substrate/substrate \
  --state open \
  --label "kind/bug,area/tests" \
  --search "flaky: <TEST_NAME>" \
  --json number,title

If none exists, create:

bash
gh issue create \
  --repo agent-substrate/substrate \
  --title "flaky: <TEST_NAME>" \
  --label "kind/bug,area/tests" \
  --body "$(cat <<'BODY'
## Flaky test detected

**Test:** `<TEST_NAME>`
**Job:** `<e2e-gvisor | e2e-microvm | unit>`

### Evidence (last 7 days)

| Metric | Value |
|---|---|
| Runs analysed | <TOTAL_RUNS> |
| Failures | <FAIL_COUNT> |
| Passes | <PASS_COUNT> |
| Flake rate | <FLAKE_RATE>% |
| Infra-failure runs excluded | <INFRA_EXCLUDED> |

### Failing run examples
<links to 2-3 failing runs>

### Passing run examples
<links to 1-2 passing runs>

### Infra triage
Infra failures were excluded before computing this flake rate. The remaining
failures cannot be explained by cluster setup, image pull, or proxy errors.

A draft fix PR will be opened by the detect-flaky-tests agent.
BODY
)"

Step 5b — Create issues for recurring infra failures

For each infra pattern appearing in ≥2 runs:

bash
gh issue list \
  --repo agent-substrate/substrate \
  --state open \
  --search "infra: <PATTERN_SUMMARY>" \
  --json number,title

If none exists, create:

bash
gh issue create \
  --repo agent-substrate/substrate \
  --title "infra: <PATTERN_SUMMARY>" \
  --label "kind/bug,area/dev-infra" \
  --body "$(cat <<'BODY'
## Recurring infrastructure failure in CI

**Pattern:** `<infra error pattern>`
**Job type:** `<unit | e2e-gvisor | e2e-microvm>`

### Evidence

| Metric | Value |
|---|---|
| Occurrences in last 7 days | <COUNT> |
| Example runs | <links> |

### Impact

This failure causes entire CI jobs to abort before tests run. It inflates
apparent failure rates and masks real test flakiness. Fixing it will improve
flakiness signal quality.

### Log excerpt
\`\`\`
<paste 3-5 lines of the actual error from the log>
\`\`\`
BODY
)"

Step 6 — Open a draft fix PR for each new flaky test

Read the test source file. Diagnose the likely cause using the patterns below (unit tests and e2e tests share most root causes, but e2e has additional patterns):

Common Go flakiness patterns and fixes
PatternSymptomsFix
Timing / sleeptime.Sleep before an assertionReplace with require.Eventually or testutil.WaitFor
Shared global statePackage-level var mutated without cleanupMove to test-local; t.Cleanup to restore
Port conflictsHardcoded port or race on ephemeral portUse ln.Addr() from the actual listener
Goroutine leakGoroutines from one test race the nextt.Cleanup(cancel) + wait for goroutines to exit
File system racesShared temp path across parallel testsUse t.TempDir()
Context not cancelledLong operation outlives testt.Context() (Go 1.21+) or t.Cleanup(cancel)
Order dependencyTest relies on prior test's side effectsMake each test self-contained
E2E: actor/pod not readyTest proceeds before actor reaches Running statePoll with require.Eventually on status, increase timeout with justification
E2E: resource cleanup racePrior test's namespace/actor not fully deleted before next testAdd explicit WaitForDeletion in t.Cleanup
E2E: network policy timingPolicy applied but not yet enforced at assertion timeRetry the connectivity check, not just the policy application

Steps for the fix PR:

  1. Create a branch: fix/flaky-<test-name-kebab> from main
  2. Apply the fix
  3. Commit: fix(tests): resolve flakiness in <TestName>\n\nFixes #<issue_number>
  4. Open a draft PR:
bash
gh pr create \
  --repo agent-substrate/substrate \
  --title "fix(tests): resolve flakiness in <TestName>" \
  --draft \
  --body "$(cat <<'BODY'
## Summary

Fixes the flaky test `<TestName>` in `<package>` (`<job_type>` lane).

**Root cause:** <one sentence>
**Fix:** <one sentence>

Closes #<issue_number>

## Evidence

Flake rate over last 7 days: <FLAKE_RATE>% (<FAIL_COUNT> fail / <PASS_COUNT> pass)
Infra-failure runs excluded from count: <INFRA_EXCLUDED>

Failing runs: <links>
Passing runs: <links>

## Test plan

- [ ] Run `go test -race -count=10 ./path/to/package/...` locally for unit tests
- [ ] For e2e: re-run the affected suite 3+ times against a kind cluster
BODY
)"

Do not open a fix PR for infra issues — those require infra investigation, not test code changes.


Step 7 — Report

Output two tables:

Flaky tests
TestJobFlake rateFail/PassInfra excludedIssueFix PRAction
TestFooe2e-gvisor40%4/62 runs#NNN#MMMcreated
Infra issues
PatternJobOccurrencesIssueAction
proxy.golang.org INTERNAL_ERRORunit4#OOOcreated

If nothing found in either category: No new flaky tests or infra issues detected.


Notes

  • This skill does NOT write to BigQuery, update a dashboard, or perform any storage operations. Those are handled by the cron job that invokes this skill.
  • If log download fails for a run (e.g. logs expired after 90 days), skip that run and note it in the report.
  • The --add-label flag may fail if a label does not exist on the repo; fall back to omitting labels and add a comment instead.
  • E2E flakiness in one lane (gVisor or microVM) is flagged independently — a test does not need to be flaky in both lanes to warrant an issue.

© agent-substrate, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/detect-flaky-tests of agent-substrate/substrate.

Open the folder on GitHubat commit 0b91488

Compare with similar skills

Detect Flaky Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Detect Flaky Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Detect Flaky Tests this skillagent-substrate/substrate4.5k—~3kAutomated safety check: PassApache-2.0
GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb6.7k—~4.4kAutomated safety check: PassApache-2.0
Debugging Opik E2E Testscomet-ml/opik22k—~1.8kAutomated safety check: PassApache-2.0
Debug Playwrightquay/quay2.8k—~1.2kAutomated safety check: PassApache-2.0
Babysit PRZenUml/web-sequence150—~871Automated safety check: PassMIT
Analysing CI Failuresgolemcloud/golem1.5k—~862Automated safety check: PassCustom licence

Similar skills

  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Debug Playwright E2E test failures from GitHub Actions CI runs.

    2.8k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Babysit PR

    ZenUml/web-sequence

    Monitor and diagnose GitHub Actions checks on ZenUML web-sequence PRs, fixing code-caused CI failures when appropriate.

    150 GitHub stars~871 tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Analysing CI Failures

    golemcloud/golem

    Analysing GitHub Actions CI failures from a run URL. An agent skill from golemcloud/golem.

    1.5k GitHub stars~862 tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug.

    23k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed

More from agent-substrate/substrate

  • Post Draft Review

    agent-substrate/substrate

    Posts pull request review findings as GitHub draft (pending) inline comments for a human to edit and submit, instead of publishing them straight to the PR author.

    4.5k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Triage Issues

    agent-substrate/substrate

    Triages open GitHub issues by applying the correct labels. An agent skill from agent-substrate/substrate.

    4.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Review Crds

    agent-substrate/substrate

    Reviews CRDs for compliance with Kubernetes API conventions.

    4.5k GitHub stars~228 tokensUpdated today
    Auto-check passed
  • Security Status Report

    agent-substrate/substrate

    Generates a security status report based on docs/threats.json by spinning up sub-agents for each threat to compute a quality score.

    4.5k GitHub stars~524 tokensUpdated today
    Auto-check passed
  • Agents Md

    agent-substrate/substrate

    Generates or updates an AGENTS.md file

    4.5k GitHub stars~635 tokensUpdated today
    Auto-check passed

Categories

Questions about Detect Flaky Tests

What does Detect Flaky Tests do?

Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM…. Detect Flaky Tests is an agent skill from agent-substrate/substrate. Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM lanes).

When should I use Detect Flaky Tests?

Detect Flaky Tests fits situations like: tasks that involve Failing and flaky tests; tasks that involve End-to-end testing.

How do I install Detect Flaky Tests in Claude Code?

Run `npx skills add agent-substrate/substrate --skill detect-flaky-tests -a claude-code`. Or copy the skill folder (.agents/skills/detect-flaky-tests in agent-substrate/substrate) into .claude/skills/detect-flaky-tests in your project. Claude Code loads it when a task matches its description.

How do I install Detect Flaky Tests in Codex?

Run `npx skills add agent-substrate/substrate --skill detect-flaky-tests -a codex`. Or copy the skill folder (.agents/skills/detect-flaky-tests in agent-substrate/substrate) into .agents/skills/detect-flaky-tests in your project. Codex loads it when a task matches its description.

Can I use Detect Flaky Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agent-substrate/substrate --skill detect-flaky-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/detect-flaky-tests, .gemini/skills/detect-flaky-tests, .github/skills/detect-flaky-tests and .opencode/skills/detect-flaky-tests in your project.

What does Detect Flaky Tests need to run?

Going by SKILL.md and its folder, Detect Flaky Tests needs the command-line tools its instructions call (gh, go and kind). Our summary lists: Docker.

Does Detect Flaky Tests access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Detect Flaky Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Detect Flaky Tests use?

Detect Flaky Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Detect Flaky Tests use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Detect Flaky Tests?

Skills that share tags, products or a category with Detect Flaky Tests: GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), Debugging Opik E2E Tests (comet-ml/opik, 22k stars), Debug Playwright (quay/quay, 2.8k stars) and Babysit PR (ZenUml/web-sequence, 150 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Detect Flaky Tests?

agent-substrate (a GitHub organization) maintains it in agent-substrate/substrate, which has 4,459 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 7, 2026.

Source: agent-substrate/substrate on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.