Agent skill

Flaky Test Detector

by ArabelaTso in ArabelaTso/Skills-4-SE

Identifies non-deterministic or unreliable tests through static code analysis and test result analysis.

Apache-2.0Auto-check passedTesting & QA

Install Flaky Test Detector

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill flaky-test-detector -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE flaky-test-detector --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/flaky-test-detector .claude/skills/flaky-test-detector && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flaky-test-detector
GitHub stars
253
Token cost
~2k tokens
SKILL.md length
825 words
Files
4 (incl. scripts, references)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Identifies non-deterministic or unreliable tests through static code analysis and test result analysis.

  • Works in 5 steps: Understand the Context → Choose Detection Method → Analyze for Flakiness → …
  • Claude needs to find flaky tests
  • SKILL.md covers Quick Start, What Makes Tests Flaky, Detection Methods and Framework-Specific Guidance, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Flaky Test Detector is an agent skill from ArabelaTso/Skills-4-SE. Identifies non-deterministic or unreliable tests through static code analysis and test result analysis. Use when Claude needs to find flaky tests, analyze test reliability, or investigate intermittent test failures. Supports Python (pytest, unittest) and Java (JUnit, TestNG) test frameworks. Trigger when users mention "flaky tests", "intermittent failures", "non-deterministic tests", "unreliable tests", or ask to "find flaky tests", "analyze test stability", or "why tests fail randomly".

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/flaky-patterns.md`, `references/remediation-strategies.md` and `scripts/analyze_test_results.py`).

It sits in Testing & QA, covering Failing and flaky tests and Unit testing. It works with JUnit, pytest, Java and Python. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • Claude needs to find flaky tests
  • Analyze test reliability
  • Investigate intermittent test failures
  • Users mention flaky tests

Example prompts

  • “flaky tests”
  • “intermittent failures”
  • “non-deterministic tests”
  • “/flaky-test-detector”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Understand the Context
  2. Choose Detection Method
  3. Analyze for Flakiness
  4. Report Findings
  5. Suggest Remediation

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flaky Test Detector loads about 2k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 825 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 825 words, ~2,011 tokens.

Download SKILL.mdSave it as .claude/skills/flaky-test-detector/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
flaky-test-detector
description
Identifies non-deterministic or unreliable tests through static code analysis and test result analysis. Use when Claude needs to find flaky tests, analyze test reliability, or investigate intermittent test failures. Supports Python (pytest, unittest) and Java (JUnit, TestNG) test frameworks. Trigger when users mention "flaky tests", "intermittent failures", "non-deterministic tests", "unreliable tests", or ask to "find flaky tests", "analyze test stability", or "why tests fail randomly".

Flaky Test Detector

Identify and fix non-deterministic tests that intermittently fail without code changes.

Quick Start

When a user reports flaky tests or asks for test reliability analysis:

  1. Identify the approach: Determine if analyzing code patterns or test execution results
  2. Analyze for flakiness: Look for common flaky patterns in test code or execution history
  3. Report findings: List identified flaky tests with specific issues
  4. Suggest fixes: Provide concrete remediation strategies

What Makes Tests Flaky

Flaky tests fail intermittently without code changes due to:

  • Timing issues: Race conditions, fixed sleeps, async/await problems
  • State management: Shared state between tests, improper cleanup
  • External dependencies: Network calls, database connections, file system
  • Randomness: Unseeded random data, UUID generation
  • Time dependencies: Current time/date, timezone assumptions
  • Resource issues: Leaks, insufficient cleanup
  • Test order: Dependencies between tests
  • Environment: Hardcoded paths, missing env vars

Detection Methods

Static Code Analysis

Analyze test code for common flaky patterns.

When to use:

  • Reviewing test code for potential issues
  • Proactive flakiness prevention
  • Code review of new tests
  • Refactoring existing tests

Process:

  1. Read test files
  2. Search for flaky patterns (see flaky-patterns.md)
  3. Identify specific issues with line numbers
  4. Suggest fixes (see remediation-strategies.md)

Common patterns to detect:

Timing issues:

  • time.sleep(), Thread.sleep() - Fixed waits
  • Missing await in async functions
  • Race conditions with threading

State issues:

  • Class or global variables in test classes
  • Missing setUp/tearDown or fixtures
  • Database operations without cleanup

External dependencies:

  • requests.get(), http.client - Real network calls
  • Database connections to production/external DBs
  • File operations without temp directories

Randomness:

  • random. without seed
  • UUID.randomUUID() without mocking
  • Non-deterministic data generation

Time dependencies:

  • datetime.now(), System.currentTimeMillis()
  • Timezone-dependent assertions
  • Date comparisons without mocking
Test Result Analysis

Analyze test execution history to find inconsistent results.

When to use:

  • Tests are failing intermittently in CI/CD
  • Investigating specific test reliability
  • Analyzing test suite health
  • Tracking flakiness over time

Process:

  1. Collect test results from multiple runs
  2. Use scripts/analyze_test_results.py to analyze patterns
  3. Review flakiness scores and patterns
  4. Investigate high-scoring tests

Script usage:

bash
python scripts/analyze_test_results.py test_results.json

Input format (JSON):

json
[
  {
    "test_name": "test_user_login",
    "status": "passed",
    "timestamp": "2024-01-01T10:00:00",
    "duration": 1.23
  },
  {
    "test_name": "test_user_login",
    "status": "failed",
    "timestamp": "2024-01-01T11:00:00",
    "duration": 1.45
  }
]

Metrics:

  • Flakiness score: 0-1, higher = more flaky (based on pass rate variance)
  • Pass rate: Percentage of successful runs
  • Pattern: Recent pass/fail sequence (P = pass, F = fail)
  • Alternating: Whether test alternates between pass/fail
  • Duration variance: Inconsistent execution time indicates issues

Framework-Specific Guidance

Python (pytest, unittest)

Common issues:

  • Missing fixtures or improper fixture scope
  • Shared class variables
  • Not using tmp_path for file operations
  • Missing @pytest.mark.django_db for database tests
  • Unseeded random module usage

Best practices:

  • Use fixtures for test data and cleanup
  • Use tmp_path fixture for file operations
  • Mock external calls with pytest-mock or unittest.mock
  • Use freezegun for time mocking
  • Seed random with random.seed()
Java (JUnit, TestNG)

Common issues:

  • Static variables in test classes
  • Missing @Before/@After cleanup
  • Not using @Transactional for database tests
  • Fixed Thread.sleep() calls
  • Hardcoded file paths

Best practices:

  • Use @Before/@After for setup/cleanup
  • Use @Transactional for automatic rollback
  • Mock with Mockito
  • Use Clock for time mocking
  • Use try-with-resources for resource management

Workflow

1. Understand the Context

Ask clarifying questions:

  • What tests are flaky?
  • How often do they fail?
  • What's the failure pattern?
  • Any recent changes?
  • CI/CD or local environment?
Show full SKILL.md (319 more words)Show less
2. Choose Detection Method

Static analysis if:

  • Reviewing code proactively
  • No test execution history available
  • Want to prevent flakiness

Result analysis if:

  • Have test execution history
  • Tests failing intermittently
  • Need to quantify flakiness
3. Analyze for Flakiness

For static analysis:

  • Read test files
  • Search for patterns from flaky-patterns.md
  • Note specific issues with line numbers
  • Categorize by issue type

For result analysis:

  • Run analyze_test_results.py script
  • Review flakiness scores
  • Identify high-risk tests
  • Examine failure patterns
4. Report Findings

Structure the report:

  • Summary: Number of flaky tests found
  • High priority: Tests with highest flakiness scores
  • By category: Group by issue type
  • Specific issues: File paths and line numbers

Example format:

Found 5 potentially flaky tests:

HIGH PRIORITY:
- test_user_login (flakiness: 0.85)
  - Line 45: time.sleep(2) - fixed wait
  - Line 52: Shared class variable 'user_data'

MEDIUM PRIORITY:
- test_api_call (flakiness: 0.62)
  - Line 23: requests.get() - unmocked network call
5. Suggest Remediation

For each issue, provide:

  • What's wrong: Explain the flaky pattern
  • Why it's flaky: Describe the non-determinism
  • How to fix: Concrete code example

Reference remediation-strategies.md for detailed fixes.

Example Usage Patterns

User: "Our test_checkout test keeps failing randomly" → Analyze test code for flaky patterns, report findings with fixes

User: "Find all flaky tests in the test suite" → Scan all test files for common flaky patterns

User: "This test has a 60% pass rate, why?" → Analyze test code and suggest specific fixes

User: "Analyze these test results for flakiness" → Use analyze_test_results.py script on provided data

User: "How do I fix this race condition in my test?" → Provide remediation strategy with code examples

User: "Review this test for potential flakiness" → Static analysis of specific test with recommendations

Best Practices

Detection
  • Look for multiple flaky patterns, not just one
  • Consider the test framework's idioms
  • Check both test code and test fixtures/setup
  • Review recent changes that might introduce flakiness
Reporting
  • Prioritize by severity and frequency
  • Provide specific line numbers
  • Group related issues together
  • Include confidence level in assessment
Remediation
  • Suggest framework-appropriate fixes
  • Provide complete code examples
  • Explain why the fix works
  • Consider test maintainability
Prevention
  • Recommend test design patterns
  • Suggest CI/CD improvements (retry policies, test isolation)
  • Encourage test independence
  • Promote proper mocking and fixtures

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/flaky-test-detector of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • references/flaky-patterns.md
  • references/remediation-strategies.md
  • scripts/analyze_test_results.py

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Flaky Test Detector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flaky Test Detector compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flaky Test Detector this skillArabelaTso/Skills-4-SE253—~2kAutomated safety check: PassApache-2.0
Test Analysis Extensionsmicrosoft/testfx1k2 repos~1.1kAutomated safety check: PassMIT
TDD GuideLeoYeAI/openclaw-master-skills2.2k—~1.4kAutomated safety check: PassMIT
ONNX Runtime Test Runnermicrosoft/onnxruntime22k—~1.8kAutomated safety check: PassMIT
ONNX Runtime GPU Transformers Testsmicrosoft/onnxruntime22k—~2.9kAutomated safety check: PassMIT
Temporal Python Testingwshobson/agents40k12 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Test Analysis Extensions

    microsoft/testfx

    Official

    Provides file paths to language-specific reference files for the test ANALYSIS skills (assertion-quality, test-anti-patterns, test-gap-analysis, test-smell-detection, test-tagging).

    1k GitHub starsUsed in 2 repos~1.1k tokens
    Testing & QAAuto-check passed
  • TDD Guide

    LeoYeAI/openclaw-master-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    2.2k GitHub stars~1.4k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • ONNX Runtime Test Runner

    microsoft/onnxruntime

    Official

    Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

    22k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Temporal workflows with pytest, time-skipping, and mocking strategies.

    40k GitHub starsUsed in 12 repos~1.2k tokens
    Testing & QAAuto-check passed
  • Squid Testing Python

    iusztinpaul/squid

    Write and evaluate effective Python tests using pytest. An agent skill from iusztinpaul/squid.

    203 GitHub stars~1.3k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from ArabelaTso/Skills-4-SE

All 150 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Flaky Test Detector

What does Flaky Test Detector do?

Identifies non-deterministic or unreliable tests through static code analysis and test result analysis. Flaky Test Detector is an agent skill from ArabelaTso/Skills-4-SE. Identifies non-deterministic or unreliable tests through static code analysis and test result analysis.

When should I use Flaky Test Detector?

Flaky Test Detector fits situations like: Claude needs to find flaky tests; analyze test reliability; investigate intermittent test failures; users mention flaky tests.

How do I install Flaky Test Detector in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill flaky-test-detector -a claude-code`. Or copy the skill folder (skills/flaky-test-detector in ArabelaTso/Skills-4-SE) into .claude/skills/flaky-test-detector in your project. Claude Code loads it when a task matches its description.

How do I install Flaky Test Detector in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill flaky-test-detector -a codex`. Or copy the skill folder (skills/flaky-test-detector in ArabelaTso/Skills-4-SE) into .agents/skills/flaky-test-detector in your project. Codex loads it when a task matches its description.

Can I use Flaky Test Detector in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill flaky-test-detector -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flaky-test-detector, .gemini/skills/flaky-test-detector, .github/skills/flaky-test-detector and .opencode/skills/flaky-test-detector in your project.

What does Flaky Test Detector need to run?

Going by SKILL.md and its folder, Flaky Test Detector needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Flaky Test Detector access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Flaky Test Detector safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Flaky Test Detector use?

Flaky Test Detector is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flaky Test Detector use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Flaky Test Detector?

Skills that share tags, products or a category with Flaky Test Detector: Test Analysis Extensions (microsoft/testfx, 1k stars), TDD Guide (LeoYeAI/openclaw-master-skills, 2.2k stars), ONNX Runtime Test Runner (microsoft/onnxruntime, 22k stars) and ONNX Runtime GPU Transformers Tests (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flaky Test Detector?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.