Agent skill

Fixing Flaky Tests

by rileyhilliard in rileyhilliard/claude-essentials

Diagnose and fix tests that pass in isolation but fail when run concurrently.

MITAuto-check passedTesting & QA

Install Fixing Flaky Tests

skills CLI
$ npx skills add rileyhilliard/claude-essentials --skill fixing-flaky-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rileyhilliard/claude-essentials fixing-flaky-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rileyhilliard/claude-essentials.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ce/skills/fixing-flaky-tests .claude/skills/fixing-flaky-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fixing-flaky-tests
GitHub stars
130
Token cost
~1.1k tokens
SKILL.md length
233 words
Files
4 (incl. references)
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Diagnose and fix tests that pass in isolation but fail when run concurrently.

  • Works in 3 steps: Run failing test 10x alone - does it… → Run failing test 10x with the suite -… → Check error message - mentions…
  • Tasks that involve Failing and flaky tests
  • SKILL.md covers Diagnose first, Shared state (deterministic…, Race conditions (random… and Resource conflicts (port/file…, plus 2 more sections
  • Calls pytest and jest

What it does

Fixing Flaky Tests is an agent skill from rileyhilliard/claude-essentials. Diagnose and fix tests that pass in isolation but fail when run concurrently. Covers shared state isolation, resource conflicts, and timing-based flakiness.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/jest.md`, `references/playwright.md` and `references/python.md`).

It sits in Testing & QA, covering Failing and flaky tests. It works with Python, Jest, Playwright and TypeScript. The licence is MIT.

When your agent uses it

  • Tasks that involve Failing and flaky tests

Example prompts

  • “/fixing-flaky-tests”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Run failing test 10x alone - does it always pass?
  2. Run failing test 10x with the suite - same error or different?
  3. Check error message - mentions port/file/connection?

What it can do on your machine

Read from SKILL.md and the folder at commit 3a67a01. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest
    • jest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fixing Flaky Tests loads about 1.1k tokens when it runs, and up to ~4.6k if it reads all its reference files. Until then it costs about 44 tokens; SKILL.md has 233 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rileyhilliard/claude-essentials at commit 3a67a01, republished under its MIT licence (© rileyhilliard). 233 words, ~1,064 tokens.

Download SKILL.mdSave it as .claude/skills/fixing-flaky-tests/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
fixing-flaky-tests
description
Diagnose and fix tests that pass in isolation but fail when run concurrently. Covers shared state isolation, resource conflicts, and timing-based flakiness.

If the current repo has its own rules/skills covering this topic (check .claude/rules/ and repo CLAUDE.md), those take precedence — apply this skill only where they're silent.

Fixing Flaky Tests

Target symptom: Tests pass when run alone, fail when run with other tests.

Diagnose first

Test passes alone, fails with others?
    │
    ├─ Same error every time → Shared state
    │   └─ Database, globals, files, singletons
    │
    ├─ Random/timing failures → Race condition
    │   └─ See async waiting patterns in `writing-tests` skill
    │
    └─ Resource errors (port, file lock) → Resource conflict
        └─ Need unique resources per test/worker

Quick diagnosis:

  1. Run failing test 10x alone - does it always pass?
  2. Run failing test 10x with the suite - same error or different?
  3. Check error message - mentions port/file/connection?

Shared state (deterministic failures)

Tests pollute state that other tests depend on. Fix by isolating state per test.

State TypeIsolation Pattern
DatabaseTransaction rollback, savepoints, worker-specific DBs
Global variablesReset in beforeEach/afterEach
SingletonsProvide fresh instance per test
Module statejest.resetModules() or equivalent
FilesUnique paths per test, temp directories
Environment varsSave/restore in setup/teardown

Database isolation (most common):

python
# Python: Savepoint rollback - each test gets rolled back
@pytest.fixture
async def db_session(db_engine):
    async with db_engine.connect() as conn:
        await conn.begin()
        await conn.begin_nested()  # Savepoint
        # ... yield session ...
        await conn.rollback()  # All changes vanish
typescript
// Jest: Reset mocks between tests
beforeEach(() => {
  jest.clearAllMocks()
  jest.resetModules()  // Clear module cache before test
})

afterEach(() => {
  jest.restoreAllMocks()  // Restore spied functions
})

See language-specific references for complete patterns.

Race conditions (random failures)

Tests don't wait for async operations to complete.

See the writing-tests skill for async waiting patterns:

  • Framework-specific waiting (Testing Library findBy, Playwright auto-wait)
  • Custom polling helpers
  • When arbitrary timeouts are acceptable

Quick summary: Wait for conditions, not time:

typescript
// Bad
await sleep(500)

// Good
await waitFor(() => expect(result).toBe('done'))

Resource conflicts (port/file errors)

Multiple tests or workers compete for same resource.

Worker-specific resources:

python
# Python pytest-xdist: unique DB per worker
@pytest.fixture(scope="session")
def database_url(worker_id):
    if worker_id == "master":
        return "postgresql://localhost/test"
    return f"postgresql://localhost/test_{worker_id}"
typescript
// Jest/Node: dynamic port allocation
const server = app.listen(0)  // OS assigns available port
const port = server.address().port

File conflicts:

python
import tempfile

@pytest.fixture
def temp_dir():
    with tempfile.TemporaryDirectory() as d:
        yield d

Language-specific isolation patterns

StackReference
Python (pytest, SQLAlchemy)references/python.md
Jest / Testing Libraryreferences/jest.md
Playwright E2Ereferences/playwright.md
Async waiting patterns (TypeScript)writing-tests waiting-typescript
Async waiting patterns (Python)writing-tests waiting-python

Verification

After fixing, verify the fix worked:

bash
# Run the specific test many times
pytest tests/test_flaky.py -x --count=20

# Run with parallelism
pytest -n auto

# Jest equivalent
jest --runInBand  # First verify serial works
jest              # Then verify parallel works

© rileyhilliard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in plugins/ce/skills/fixing-flaky-tests of rileyhilliard/claude-essentials.

  • SKILL.md
  • references/jest.md
  • references/playwright.md
  • references/python.md

Open the folder on GitHubat commit 3a67a01

Compare with similar skills

Fixing Flaky Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fixing Flaky Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fixing Flaky Tests this skillrileyhilliard/claude-essentials130—~1.1kAutomated safety check: PassMIT
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0
RStudio Selenium to Playwright Migrationrstudio/rstudio5.1k—~3.6kAutomated safety check: PassCustom licence
Playwright Testingchongdashu/vibejam-starter-pack149—~2.1kAutomated safety check: PassNone
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone
Testingradix-ng/primitives274—~3.3kAutomated safety check: PassMIT

Similar skills

  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Converts RStudio Python Selenium electron tests into TypeScript Playwright tests, checking each against a live RStudio before counting it as migrated.

    5.1k GitHub stars~3.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.1k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Testing

    radix-ng/primitives

    Test Radix NG primitives across every layer and pick the RIGHT one for a change: Vitest unit (zoneless), jest-axe a11y, Playwright browser regression (apps/visual-regression), SSR…

    274 GitHub stars~3.3k tokensUpdated 11 days ago
    Testing & QAAuto-check passed
  • Test Specialist

    travisjneuman/.claude

    Test-writing patterns for JS/TS, Python, Go, and Rust (unit, integration, E2E, visual regression).

    101 GitHub stars~3.7k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from rileyhilliard/claude-essentials

All 17 skills in this repo
  • Handling Errors

    rileyhilliard/claude-essentials

    Prevents silent failures and context loss in error handling.

    130 GitHub stars~632 tokensUpdated 1 mo ago
    Auto-check passed
  • Planning Products

    rileyhilliard/claude-essentials

    Defines product features from a PM perspective (JTBD, competitive research, scope negotiation) before technical planning.

    130 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Writing Tests

    rileyhilliard/claude-essentials

    Writes behavior-focused tests using Testing Trophy model with real dependencies.

    130 GitHub stars~999 tokensUpdated 1 mo ago
    Auto-check passed
  • Optimizing Performance

    rileyhilliard/claude-essentials

    Measure-first performance optimization that balances gains against complexity.

    130 GitHub stars~470 tokensUpdated 1 mo ago
    Auto-check passed
  • Architecting Systems

    rileyhilliard/claude-essentials

    Guides clean, scalable system architecture during the build phase.

    130 GitHub stars~713 tokensUpdated 1 mo ago
    Auto-check passed
  • Design

    rileyhilliard/claude-essentials

    Enforces precise, minimal design for dashboards and admin interfaces.

    130 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Fixing Flaky Tests

What does Fixing Flaky Tests do?

Diagnose and fix tests that pass in isolation but fail when run concurrently. Fixing Flaky Tests is an agent skill from rileyhilliard/claude-essentials. Diagnose and fix tests that pass in isolation but fail when run concurrently.

When should I use Fixing Flaky Tests?

Fixing Flaky Tests fits situations like: tasks that involve Failing and flaky tests.

How do I install Fixing Flaky Tests in Claude Code?

Run `npx skills add rileyhilliard/claude-essentials --skill fixing-flaky-tests -a claude-code`. Or copy the skill folder (plugins/ce/skills/fixing-flaky-tests in rileyhilliard/claude-essentials) into .claude/skills/fixing-flaky-tests in your project. Claude Code loads it when a task matches its description.

How do I install Fixing Flaky Tests in Codex?

Run `npx skills add rileyhilliard/claude-essentials --skill fixing-flaky-tests -a codex`. Or copy the skill folder (plugins/ce/skills/fixing-flaky-tests in rileyhilliard/claude-essentials) into .agents/skills/fixing-flaky-tests in your project. Codex loads it when a task matches its description.

Can I use Fixing Flaky Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rileyhilliard/claude-essentials --skill fixing-flaky-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fixing-flaky-tests, .gemini/skills/fixing-flaky-tests, .github/skills/fixing-flaky-tests and .opencode/skills/fixing-flaky-tests in your project.

What does Fixing Flaky Tests need to run?

Going by SKILL.md and its folder, Fixing Flaky Tests needs the command-line tools its instructions call (pytest and jest). Our summary lists: Python 3.

Does Fixing Flaky Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Fixing Flaky Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fixing Flaky Tests use?

Fixing Flaky Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fixing Flaky Tests use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.

What are the alternatives to Fixing Flaky Tests?

Skills that share tags, products or a category with Fixing Flaky Tests: Testing Patterns (softspark/ai-toolkit, 179 stars), RStudio Selenium to Playwright Migration (rstudio/rstudio, 5.1k stars), Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fixing Flaky Tests?

rileyhilliard (a GitHub user) maintains it in rileyhilliard/claude-essentials, which has 130 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on August 18, 2026.

Source: rileyhilliard/claude-essentials on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.