Agent skill

Flaky Test Deflaker

by QwenLM in QwenLM/qwen-code

Stabilizes one flaky test with the smallest possible fix, such as a longer timeout or a real wait, without weakening, skipping or deleting any assertion.

Apache-2.0Auto-check passedTesting & QA

Install Flaky Test Deflaker

skills CLI
$ npx skills add QwenLM/qwen-code --skill deflake -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install QwenLM/qwen-code deflake --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/deflake .claude/skills/deflake && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deflake
GitHub stars
28k
Token cost
~856 tokens
SKILL.md length
492 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Stabilizes one flaky test with the smallest possible fix, such as a longer timeout or a real wait, without weakening, skipping or deleting any assertion.

  • Works in 4 steps: Raise a timeout / poll budget. A test… → Stabilize timing / waiting. Replace a… → Make randomness / time deterministic.… → …
  • A CI test passes when rerun on the same commit
  • SKILL.md covers The only allowed fixes, Hard rules and Verify
  • Calls npx

What it does

The input is a deflake issue naming one test that failed and then passed on a rerun of the same commit, with its file, name and failure signature. The agent reads the test, works out the source of the nondeterminism and applies the smallest fix from a fixed list of four: raise a timeout or poll budget, stabilize timing by awaiting the real condition, make randomness and time deterministic, or isolate tests that collide on shared resources.

Hard rules forbid deleting, skipping or loosening the test, widening expected ranges, adding blanket try/catch or retry wrappers, or refactoring unrelated code. If none of the four fixes applies, or the failure looks like a real intermittent product bug, the agent writes a failure.md describing what it found and stops so a person can take over.

When your agent uses it

  • A CI test passes when rerun on the same commit
  • Fixing a test that times out under CI load
  • Replacing fixed sleeps with waits on the real condition
  • Making a test that depends on the clock or random values deterministic

Example prompts

  • “Deflake the retry test in api-client.test.ts; it timed out once in CI.”
  • “This test passes on rerun, so fix the race without touching its assertions.”
  • “Replace the fixed sleep in the file watcher test with an explicit wait.”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Raise a timeout / poll budget. A test that blows vitest's default under
  2. Stabilize timing / waiting. Replace a bare setTimeout/fixed sleep
  3. Make randomness / time deterministic. Seed the RNG, vi.useFakeTimers()
  4. Isolate / serialize interference. Give tests that collide on a shared

What it can do on your machine

Read from SKILL.md and the folder at commit d9c6f8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flaky Test Deflaker loads about 856 tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 492 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~856

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from QwenLM/qwen-code at commit d9c6f8c, republished under its Apache-2.0 licence (© QwenLM). 492 words, ~856 tokens.

Download SKILL.mdSave it as .claude/skills/deflake/SKILL.md (or your agent's skills folder).
name
deflake
description
Stabilize a flaky test with a minimal, assertion-preserving fix — never by weakening or deleting the check.

Deflake a flaky test

A deflake: issue names ONE test that has been observed failing and then passing on a rerun of the same commit — the definitive flaky signature. Your job is to make that test deterministic without changing what it verifies.

The issue body carries the test identity (file + name) and the observed failure signature (e.g. Test timed out in 5000ms, Timed out waiting for …, an order-dependent assertion, a wall-clock/random-dependent value). Read the test, reproduce the mechanism in your head, and apply the SMALLEST fix from the allowed set below that removes the nondeterminism.

The only allowed fixes

  1. Raise a timeout / poll budget. A test that blows vitest's default under CI contention (fully-mocked or I/O-bound, not a perf test) gets a generous per-test testTimeout (3rd arg to it), or its internal poll loop is given a real wall-clock budget instead of a fixed iteration count (a fixed count of setImmediate turns elapses in milliseconds and races real I/O).
  2. Stabilize timing / waiting. Replace a bare setTimeout/fixed sleep with an explicit await of the real condition (vi.waitFor, a resolved promise, an event). Pre-warm a lazy load (e.g. a WASM runtime) in beforeAll so per-test time doesn't include first-load cost.
  3. Make randomness / time deterministic. Seed the RNG, vi.useFakeTimers() / mock Date.now, or pin the input so a value that depends on the real clock or Math.random can't drift.
  4. Isolate / serialize interference. Give tests that collide on a shared resource (a same-named tempdir, a fixed port, a global singleton) unique per-test resources, or serialize them.
Show full SKILL.md (237 more words)Show less

Hard rules

  • Never delete the test, skip/todo it, loosen an assertion, widen an expected range, add a blanket try/catch, or add a retry wrapper around the assertion. Those hide the flake instead of fixing it — and could hide a real bug. If none of the four fixes applies, or the failure looks like a REAL intermittent product bug (not test nondeterminism), write <workdir>/failure.md explaining what you found and stop. A human deflakes it.
  • Keep the diff minimal and local to the named test (and its file's helpers). Do not refactor unrelated code.
  • Preserve every assertion and every input exactly. A timeout bump changes only the ceiling; a determinism fix changes only the source of nondeterminism.
  • Prefer a per-test or per-file change over a global config change unless the same class demonstrably spans the whole package (then a testTimeout in that package's vitest.config.ts is acceptable, as it only raises the ceiling and weakens no assertion).

Verify

Run the named test's file several times (npx vitest run <file> in the right package, repeated) — it must pass every time. Then run the standard verify gate (build / typecheck / lint / the changed test). If you cannot make it pass deterministically, write <workdir>/failure.md and stop.

Then follow .qwen/skills/prepare-pr/SKILL.md for the PR body and write the bilingual <workdir>/e2e-report.md (per the Shared Rules) stating: the flaky mechanism, which of the four fixes you applied and why, and the repeated-run evidence that it is now deterministic.

© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .qwen/skills/deflake of QwenLM/qwen-code.

Open the folder on GitHubat commit d9c6f8c

Compare with similar skills

Flaky Test Deflaker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flaky Test Deflaker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flaky Test Deflaker this skillQwenLM/qwen-code28k—~856Automated safety check: PassApache-2.0
Ckeditor5 TestingTriliumNext/Trilium38k—~3.3kAutomated safety check: PassAGPL-3.0
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT
Caliber Testingcaliber-ai-org/ai-setup1.3k—~3.2kAutomated safety check: PassMIT
Playwright Testingchongdashu/vibejam-starter-pack149—~2.1kAutomated safety check: PassNone
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone

Similar skills

  • Ckeditor5 Testing

    TriliumNext/Trilium

    Testing CKEditor 5 plugins in the Trilium monorepo. An agent skill from TriliumNext/Trilium.

    38k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Caliber Testing

    caliber-ai-org/ai-setup

    Writes Vitest tests following project patterns: tests/ directories, vi.mock() for module mocking with vi.hoisted() for test-time factories, global LLM mock from src/test/setup.ts, environment…

    1.3k GitHub stars~3.2k tokensUpdated 17 days ago
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.1k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Testing

    radix-ng/primitives

    Test Radix NG primitives across every layer and pick the RIGHT one for a change: Vitest unit (zoneless), jest-axe a11y, Playwright browser regression (apps/visual-regression), SSR…

    274 GitHub stars~3.3k tokensUpdated 12 days ago
    Testing & QAAuto-check passed

More from QwenLM/qwen-code

All 41 skills in this repo
  • Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.

    28k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.

    28k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.

    28k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Flaky Test Deflaker

What does Flaky Test Deflaker do?

Stabilizes one flaky test with the smallest possible fix, such as a longer timeout or a real wait, without weakening, skipping or deleting any assertion. The input is a deflake issue naming one test that failed and then passed on a rerun of the same commit, with its file, name and failure signature. The agent reads the test, works out the source of the nondeterminism and applies the smallest fix from a fixed list of four: raise a timeout or poll budget, stabilize timing by awaiting the real condition, make randomness and time deterministic, or isolate tests that collide on shared resources.

When should I use Flaky Test Deflaker?

Flaky Test Deflaker fits situations like: A CI test passes when rerun on the same commit; fixing a test that times out under CI load; replacing fixed sleeps with waits on the real condition; making a test that depends on the clock or random values deterministic.

How do I install Flaky Test Deflaker in Claude Code?

Run `npx skills add QwenLM/qwen-code --skill deflake -a claude-code`. Or copy the skill folder (.qwen/skills/deflake in QwenLM/qwen-code) into .claude/skills/deflake in your project. Claude Code loads it when a task matches its description.

How do I install Flaky Test Deflaker in Codex?

Run `npx skills add QwenLM/qwen-code --skill deflake -a codex`. Or copy the skill folder (.qwen/skills/deflake in QwenLM/qwen-code) into .agents/skills/deflake in your project. Codex loads it when a task matches its description.

Can I use Flaky Test Deflaker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill deflake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deflake, .gemini/skills/deflake, .github/skills/deflake and .opencode/skills/deflake in your project.

What does Flaky Test Deflaker need to run?

Going by SKILL.md and its folder, Flaky Test Deflaker needs the command-line tools its instructions call (npx).

Does Flaky Test Deflaker access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Flaky Test Deflaker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Flaky Test Deflaker use?

Flaky Test Deflaker is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flaky Test Deflaker use?

About 856 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flaky Test Deflaker?

Skills that share tags, products or a category with Flaky Test Deflaker: Ckeditor5 Testing (TriliumNext/Trilium, 38k stars), Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars), Caliber Testing (caliber-ai-org/ai-setup, 1.3k stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flaky Test Deflaker?

QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,410 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 11, 2026.

Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.