Agent skill

Routine Flaky Test Fixer

by bex-co in bex-co/beancount-io

Autonomous maintenance routine that root-causes intermittently failing CI tests — mining GitHub Actions history for same-commit red-to-green reruns, reproducing locally with repeated runs, fixing…

MITAuto-check passedTesting & QA

Install Routine Flaky Test Fixer

skills CLI
$ npx skills add bex-co/beancount-io --skill routine-flaky-test-fixer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bex-co/beancount-io routine-flaky-test-fixer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bex-co/beancount-io.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/routine-flaky-test-fixer .claude/skills/routine-flaky-test-fixer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
routine-flaky-test-fixer
GitHub stars
296
Token cost
~949 tokens
SKILL.md length
387 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Autonomous maintenance routine that root-causes intermittently failing CI tests — mining GitHub Actions history for same-commit red-to-green reruns, reproducing locally with repeated runs, fixing…

  • Works in 6 steps: Scope → Discover → Prove it → …
  • Explicitly invokes /routine-flaky-test-fixer
  • SKILL.md covers Contract, Workflow, What NOT to do and Arguments
  • Calls gh and yarn

What it does

Routine Flaky Test Fixer is an agent skill from bex-co/beancount-io. Autonomous maintenance routine that root-causes intermittently failing CI tests — mining GitHub Actions history for same-commit red-to-green reruns, reproducing locally with repeated runs, fixing the real nondeterminism (timers, teardown, ordering, unawaited promises), and shipping via /ship after 20 consecutive green runs. Use only when the user explicitly invokes /routine-flaky-test-fixer or asks to hunt down flaky tests. SKIP for tests that can never fail (routine-useless-test-pruner) and for logic bugs that…

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests and Root cause analysis. It works with GitHub Actions and Jest. The repository describes itself as: 💰 Double-entry bookkeeping made easy — plain-text accounting for humans and AI agents. Polished iOS & Android app built with React Native + Expo. The licence is MIT.

When your agent uses it

  • Explicitly invokes /routine-flaky-test-fixer
  • Asks to hunt down flaky tests
  • Is this test sometimes fails
  • Make it deterministic

Example prompts

  • “this test sometimes fails; make it deterministic”
  • “/routine-flaky-test-fixer”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Scope
  2. Discover
  3. Prove it
  4. Fix
  5. Verify
  6. Ship

What it can do on your machine

Read from SKILL.md and the folder at commit 5614bc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • yarn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and yarn, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Routine Flaky Test Fixer loads about 949 tokens when it runs. Until then it costs about 166 tokens; SKILL.md has 387 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~166
When it runs · the whole SKILL.md, loaded when a task matches
~949

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from bex-co/beancount-io at commit 5614bc7, republished under its MIT licence (© bex-co). 387 words, ~949 tokens.

Download SKILL.mdSave it as .claude/skills/routine-flaky-test-fixer/SKILL.md (or your agent's skills folder).
name
routine-flaky-test-fixer
description
Autonomous maintenance routine that root-causes intermittently failing CI tests — mining GitHub Actions history for same-commit red-to-green reruns, reproducing locally with repeated runs, fixing the real nondeterminism (timers, teardown, ordering, unawaited promises), and shipping via /ship after 20 consecutive green runs. Use only when the user explicitly invokes /routine-flaky-test-fixer or asks to hunt down flaky tests. SKIP for tests that can never fail (routine-useless-test-pruner) and for logic bugs that fail deterministically (routine-logic-bugfixer). The core trigger is "this test sometimes fails; make it deterministic".

Task: Root-cause flaky CI tests and make them deterministic

Usage: /routine-flaky-test-fixer [package-or-path]

Find a test that sometimes fails, reproduce or root-cause the nondeterminism, fix the real cause, prove determinism with repeated runs, and ship via /ship.

Contract

Read .agents/skills/routine-shared/contract.md first — preconditions, scope resolution, verify gates, ship protocol, budget, universal STOPs. Floor even without it: never ship red; one finding per ship; anything in a DO_NOT_DO.md is a hard STOP.

Workflow

1. Scope

Resolve scope per the contract. A test that is consistently red on main is the degenerate case and jumps the queue — fix or escalate it first.

2. Discover

Mine CI history on the path-filtered workflows:

zsh
gh run list --workflow=ci.yml --branch main --limit 50            # mobile
gh run list --workflow=ci-dashboard.yml --branch main --limit 50
gh run list --workflow=ci-cli.yml --branch main --limit 50
gh run list --workflow=ci-skills.yml --branch main --limit 50
gh run list --workflow=ci-backend-parity.yml --branch main --limit 50  # backend-v2

Flake signals: the same commit red then green on rerun; failures whose job logs (gh run view <id> --log-failed) show different tests failing across runs. Locally: rerun the owning suite — or the suspect file — 20–30 times (mobile via yarn test:unit, the bespoke mobile/scripts/jest-lite-runner.js; dashboard/ backend via their jest commands; cli via pytest), with shuffled order where the runner supports it.

3. Prove it

Either a local reproduction, or an unambiguous root cause read from the logs and code. The usual suspects: real timers and wall-clock time, missing teardown (handles, listeners, temp files — the dashboard OTP-teardown fix is the house example), shared state and test-order dependence, port/tmp-path collisions, unawaited promises, network reliance, unseeded randomness. Name the mechanism before touching the test.

Show full SKILL.md (160 more words)Show less
4. Fix

Fix the nondeterminism itself: fake timers, complete teardown, per-test unique ports/paths, seeded randomness, awaited async, order independence. Sometimes the bug is in the production code's lifecycle (leaked handle, race) — that is still this routine's finding; fix it at the root.

5. Verify

20 consecutive green runs of the fixed test (in-suite, not isolated, when order was implicated), then the owning package's full gate.

6. Ship

Compose /ship for this one flake. Loop within budget.

What NOT to do

  • Never add retry wrappers, jest.retryTimes, or blind timeout bumps — masking a flake is worse than leaving it visibly flaky.
  • Never delete a flaky test to make it stop flaking. If it is genuinely unsalvageable, that must independently pass routine-useless-test-pruner's proof standard — and the summary says so honestly.
  • Never mark a fix done on fewer than 20 consecutive greens; three passes is luck, not determinism.
  • A test failing the same way every run is not flaky — defer to routine-logic-bugfixer.

Arguments

$ARGUMENTS

© bex-co, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/routine-flaky-test-fixer of bex-co/beancount-io.

Open the folder on GitHubat commit 5614bc7

Compare with similar skills

Routine Flaky Test Fixer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Routine Flaky Test Fixer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Routine Flaky Test Fixer this skillbex-co/beancount-io296—~949Automated safety check: PassMIT
Pester Failure AnalysisPowerShell/PowerShell56k—~5.1kAutomated safety check: PassMIT
GreptimeDB Fuzz CI Failure InvestigationGreptimeTeam/greptimedb6.7k—~4.4kAutomated safety check: PassApache-2.0
Debugging Opik E2E Testscomet-ml/opik22k—~1.8kAutomated safety check: PassApache-2.0
CI Triagecanton-network/splice118—~1.4kAutomated safety check: PassApache-2.0
Fixing Flaky TestsPostHog/posthog40k—~5.9kAutomated safety check: PassCustom licence

Similar skills

  • Pester Failure Analysis

    PowerShell/PowerShell

    Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

    56k GitHub stars~5.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Investigates a failed Opik end-to-end test from CI, TestOps or a local run, decides regression versus flake, and proposes a fix without editing tests.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • CI Triage

    canton-network/splice

    Triage a failed splice GitHub Actions job (cn-test-failures ref) into a reproducible evidence packet - fetch job log and artifact, isolate the flagged lines, check the known flake families for…

    118 GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Fixing Flaky Tests

    PostHog/posthog

    Official

    Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo.

    40k GitHub stars~5.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Analyze Test Run

    microsoft/GitHub-Copilot-for-Azure

    Official

    Analyze a GitHub Actions integration test run and produce a skill invocation report with failure root-cause issues.

    255 GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check passed

More from bex-co/beancount-io

All 27 skills in this repo
  • Beancount Close

    bex-co/beancount-io

    Close an accounting period in a Beancount ledger by reconciling each active account through beancount-reconcile, checking assertions and recurring gaps, reviewing flags, then proposing a commit with…

    296 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Beancount Import

    bex-co/beancount-io

    Import a bank or card CSV, OFX/QFX, or QIF export into an existing Beancount ledger.

    296 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Beancount Importer Author

    bex-co/beancount-io

    Write or repair a reusable Beangulp importer from a sample bank export, with reviewed golden files and a passing test harness.

    296 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Beancount Init

    bex-co/beancount-io

    Scaffold a new Beancount ledger with bea init and validation, with optional Fava browser setup when requested.

    296 GitHub stars~1.2k tokensUpdated today
    Auto-check: notes
  • Beancount Options

    bex-co/beancount-io

    Record a described options trade or lifecycle event as validated Beancount transactions through bea, after review and confirmation.

    296 GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Beancount Reconcile

    bex-co/beancount-io

    Reconcile one Beancount account against a CSV statement or pasted PDF text.

    296 GitHub stars~4.1k tokensUpdated today
    Auto-check passed

Categories

Questions about Routine Flaky Test Fixer

What does Routine Flaky Test Fixer do?

Autonomous maintenance routine that root-causes intermittently failing CI tests — mining GitHub Actions history for same-commit red-to-green reruns, reproducing locally with repeated runs, fixing…. Routine Flaky Test Fixer is an agent skill from bex-co/beancount-io. Autonomous maintenance routine that root-causes intermittently failing CI tests — mining GitHub Actions history for same-commit red-to-green reruns, reproducing locally with repeated runs, fixing the real nondeterminism (timers, teardown, ordering, unawaited promises), and shipping via /ship after 20 consecutive green runs.

When should I use Routine Flaky Test Fixer?

Routine Flaky Test Fixer fits situations like: explicitly invokes /routine-flaky-test-fixer; asks to hunt down flaky tests; is this test sometimes fails; make it deterministic.

How do I install Routine Flaky Test Fixer in Claude Code?

Run `npx skills add bex-co/beancount-io --skill routine-flaky-test-fixer -a claude-code`. Or copy the skill folder (.agents/skills/routine-flaky-test-fixer in bex-co/beancount-io) into .claude/skills/routine-flaky-test-fixer in your project. Claude Code loads it when a task matches its description.

How do I install Routine Flaky Test Fixer in Codex?

Run `npx skills add bex-co/beancount-io --skill routine-flaky-test-fixer -a codex`. Or copy the skill folder (.agents/skills/routine-flaky-test-fixer in bex-co/beancount-io) into .agents/skills/routine-flaky-test-fixer in your project. Codex loads it when a task matches its description.

Can I use Routine Flaky Test Fixer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bex-co/beancount-io --skill routine-flaky-test-fixer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/routine-flaky-test-fixer, .gemini/skills/routine-flaky-test-fixer, .github/skills/routine-flaky-test-fixer and .opencode/skills/routine-flaky-test-fixer in your project.

What does Routine Flaky Test Fixer need to run?

Going by SKILL.md and its folder, Routine Flaky Test Fixer needs the command-line tools its instructions call (gh and yarn).

Does Routine Flaky Test Fixer access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Routine Flaky Test Fixer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Routine Flaky Test Fixer use?

Routine Flaky Test Fixer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Routine Flaky Test Fixer use?

About 949 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Routine Flaky Test Fixer?

Skills that share tags, products or a category with Routine Flaky Test Fixer: Pester Failure Analysis (PowerShell/PowerShell, 56k stars), GreptimeDB Fuzz CI Failure Investigation (GreptimeTeam/greptimedb, 6.7k stars), Debugging Opik E2E Tests (comet-ml/opik, 22k stars) and CI Triage (canton-network/splice, 118 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Routine Flaky Test Fixer?

bex-co (a GitHub organization) maintains it in bex-co/beancount-io, which has 296 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 9, 2026.

Source: bex-co/beancount-io on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.