Agent skill

Fixing Flaky Failures

by bikeindex in bikeindex/bike_index

How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky.

AGPL-3.0Auto-check passedTesting & QA

Install Fixing Flaky Failures

skills CLI
$ npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bikeindex/bike_index fixing-flaky-failures --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bikeindex/bike_index.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fixing-flaky-failures .claude/skills/fixing-flaky-failures && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fixing-flaky-failures
GitHub stars
308
Token cost
~5k tokens
SKILL.md length
2,672 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
AGPL-3.0

At a glance

How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky.

  • Works in 5 steps: Get the real failure, not the summary → Read the message literally → Instrument rather than theorise → …
  • A CI failure the user calls flaky
  • SKILL.md covers The rule that overrides…, Retries and longer waits:…, Diagnose before you touch… and Known causes in this repo, plus 2 more sections
  • Calls gh, rails and bundle

What it does

Fixing Flaky Failures is an agent skill from bikeindex/bike_index. How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky. Read this before touching any spec that is failing intermittently, and before adding, raising, or relying on a flaky: tag or a wait: bump. Trigger on a CI failure the user calls flaky, unreliable, intermittent, "green locally", "passes on retry", "fix CI", or a pasted gh run/Actions URL whose failure isn't reproducible; also whenever you catch…

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests. The repository describes itself as: All the code for Bike Index, because we love you. The licence is AGPL-3.0.

When your agent uses it

  • A CI failure the user calls flaky
  • Passes on retry
  • A pasted gh run/Actions URL whose failure isnt reproducible
  • Also whenever you catch yourself about to delete

Example prompts

  • “green locally”
  • “passes on retry”
  • “fix CI”
  • “/fixing-flaky-failures”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Get the real failure, not the summary
  2. Read the message literally
  3. Instrument rather than theorise
  4. Try to reproduce, and treat failure-to-reproduce as information
  5. Blaming your own change needs both arms measured together

What it can do on your machine

Read from SKILL.md and the folder at commit b9be85f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • rails
    • bundle

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fixing Flaky Failures loads about 5k tokens when it runs. Until then it costs about 227 tokens; SKILL.md has 2,672 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~227
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from bikeindex/bike_index at commit b9be85f, republished under its AGPL-3.0 licence (© bikeindex). 2,672 words, ~5,042 tokens.

Download SKILL.mdSave it as .claude/skills/fixing-flaky-failures/SKILL.md (or your agent's skills folder).
name
fixing-flaky-failures
description
How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged `:flaky`. **Read this before touching any spec that is failing intermittently, and before adding, raising, or relying on a `flaky:` tag or a `wait:` bump.** Trigger on a CI failure the user calls flaky, unreliable, intermittent, "green locally", "passes on retry", "fix CI", or a pasted `gh run`/Actions URL whose failure isn't reproducible; also whenever you catch yourself about to delete, skip, loosen, or retry an assertion to make a red build go green. **Equally: any spec that fails intermittently while you verify your own work** — deciding whether your change caused a flake, or whether to ship past one, is this skill's problem too, and it arrives with no CI run, no `:flaky` tag, and nobody but you calling it flaky.

Fixing flaky failures

Covers diagnosing the real mechanism, attributing a flake to a change, this repo's flaky: retry harness, and the known false-flake causes — a missing Tailwind build, the shared Redis autocomplete cache, Turbo frame timing, probe-run interference.

The rule that overrides everything else here

You may not reduce test coverage to make a flaky test pass. Not as a last resort, not "temporarily", not when the coverage looks redundant, not when you have already decided the failing step is a harness artifact rather than a real bug.

All of these are reducing coverage, and none of them is a fix:

  • Deleting an assertion, a step, a helper call, or a whole example.
  • Loosening a matcher so it can't fail (have_css(count: 2) → have_css, an exact string → a regex, have_no_content → nothing).
  • skip/pending/xit, or excluding the spec from a CI shard.

A flaky test is a reporting problem — the suite is telling you something real and telling you unreliably. Every item above changes the reporting and leaves the underlying behaviour untested.

If the only fix you can see requires giving up coverage, that is a decision for the user — describe what you'd have to give up and ask. Do not make that trade yourself and mention it in the summary afterwards; announcing a scope reduction is not the same as getting agreement for it.

The one exception that isn't an exception: if the assertion is wrong — it asserts behaviour the app never promised — then fixing it is correcting a bad test, not reducing coverage. Say explicitly why it was wrong.

Retries and longer waits: allowed, but earn them first

:flaky (and raising an existing flaky: N) and a bigger wait: are a different category from the list above, because they keep every assertion and still require it to pass. Nothing goes untested. They're a legitimate last resort — reach for them when the diagnosis below has genuinely run out, not as the first thing you try.

What "earn them" means in practice:

  • Diagnose first. Instrument, read the failure literally, check the known causes below. Most flakes here have a mechanism you can find in one focused pass, and a retry over a findable cause just makes it intermittent for longer.
  • Leave a comment saying what you found and why the retry stands in for a fix — the harness artifact, the contention, the thing you ruled out. The existing flaky: 4 on search_registrations_spec is the pattern: it names the click landing on about:blank, lists what it ruled out, and says it went unreproduced under local CPU throttling. A bare flaky: true with no comment tells the next person nothing and will outlive the problem.
  • Say plainly in your summary that you papered over it rather than fixed it, so the user can decide whether that's good enough.

Two signals that a retry is the wrong answer even as a last resort: an example that fails through all its retries (the cause is structural, and a bigger N won't help), and a wait you can't say what it's waiting for. A bump is honest when the assertion starts before the work does — a held route released, a job enqueued, an expensive response — and the comment names that.

Diagnose before you touch anything

1. Get the real failure, not the summary
bash
gh run view <run-id> --repo bikeindex/bike_index --json jobs \
  --jq '.jobs[] | "\(.databaseId) \(.name) \(.conclusion)"'
gh run view --repo bikeindex/bike_index --job <job-id> --log-failed

The log is large and ANSI-coloured; pipe through sed 's/\x1b\[[0-9;]*m//g' and grep for Failure/Error, expected, and rspec ./spec/... to get the failing example ids and the actual message.

Then download the Capybara screenshot. Every :js failure writes one, and ci.yml uploads tmp/capybara/ to the shard's test-results-<node-index> artifact for exactly this — but the log names the file without saying it's fetchable, so it usually goes unread. It shows the page the failure saw, which routinely settles a mechanism the log can only hint at: what a click actually landed on, a frame still loading, a control in a state no step in the example set.

bash
gh api repos/bikeindex/bike_index/actions/runs/<run-id>/artifacts \
  --jq '.artifacts[] | "\(.id) \(.name)"'
gh api repos/bikeindex/bike_index/actions/artifacts/<id>/zip > tmp/a.zip && unzip -o tmp/a.zip -d tmp/ci_artifact
2. Read the message literally

The exact failure text usually names the mechanism, and it is easy to skim past into a wrong assumption. Worked example from this repo: expected nil to match /\/bikes\/\d+/ was long assumed to mean "the click was lost". It doesn't — a nil current_path means the URL had no path for Capybara to return, which is three different browser states, not one (capybara/session.rb:207: nil for an about: scheme, then path unless path&.empty?):

URLhow it got there
about:blanktraversed to entry 0, or the page was replaced
chrome-error://chromewebdataa cross-document navigation failed outright
""no document has committed yet

All three screenshot blank, so the picture can't tell them apart — tmp/capybara/browser_events.log (written by spec/support/capybara.rb, uploaded with the screenshots) can. Check the matcher's source when a message is surprising, and don't read one of these three as another: the fix differs, and "about:blank" has been the standing wrong guess.

3. Instrument rather than theorise

When you can't reproduce, you can still make the browser tell you what happened. Copy the spec to a scratch file, record the events the app actually emits, and run it — a log beats an argument, and this routinely overturns the theory you were about to ship:

ruby
page.execute_script(<<~JS)
  window.__events = []
  const stamp = (name, extra) => window.__events.push(Object.assign({name, t: Math.round(performance.now())}, extra))
  document.addEventListener("turbo:before-fetch-response", (e) => {
    stamp("response", {target: e.target?.id, url: e.detail?.fetchResponse?.response?.url, prevented: e.defaultPrevented})
  })
JS
# ...drive the spec...
# A file, not puts - rtk's rspec wrapper reports a summary and drops the run's stdout
File.open("tmp/probe.log", "a") { |f| page.evaluate_script("window.__events").each { |e| f.puts e.inspect } }

Parameterise the scratch spec over the variable you suspect ([0, 8].each do |delay| around an injected sleep) so one run compares the fast and slow paths. Delete the scratch file when you're done. Two things this buys that reasoning doesn't: it distinguishes "the event never arrived" from "the event arrived and the assertion misread it", and it tells you when things happened, which is usually the answer.

4. Try to reproduce, and treat failure-to-reproduce as information
bash
for i in 1 2 3; do bundle exec rspec <the spec file> 2>&1 | grep -E "examples, " | tail -1; done

Keep that loop on the one spec file — escalating it to bin/ci costs minutes of parallel workers and browsers per iteration, and answers the same question no better.

Green locally three times doesn't mean "not reproducible, add a retry". It narrows the cause to something CI has and you don't: contention (browser, Rails and Postgres sharing one runner) or ordering (a different seed, knapsack handing this shard a different set of files, or state left by another example). Reason about which, then look for the mechanism.

For contention, slow the renderer rather than the machine — CPU hogs slow the Ruby side too, so a loop of runs takes minutes and the extra load is spent where the race isn't, and one backgrounded from a non-interactive shell survives kill $(jobs -p) — pgrep -f it. CDP throttles the browser alone, and the driver hands you a session:

ruby
page.driver.with_playwright_page do |playwright_page|
  session = playwright_page.context.new_cdp_session(playwright_page)
  session.send_message("Emulation.setCPUThrottlingRate", params: {rate: 6})
end

A rate that leaves the spec green over ~20 runs is evidence, not proof: it stretches main-thread work, not the network or a parallel shard's I/O.

Which is why it can't reach a race about when a response arrives. The common one here is a lazily loaded Stimulus controller, since the module is a fetch — hold it on the route and the late connect is deterministic, no loop:

ruby
playwright_page.route(%r{serial_controller}, ->(route, request) { held << request.url; sleep 1; route.continue })

Registering a route disables the http cache, so the reload asks again. Assert on what the handler held: a route that stops matching (a moved asset path) otherwise leaves the example green on a page that held nothing back. Measured against register--serial's connect-time reconcile, throttling at 6, 12 and 25 left it green over 30+ runs; the route hold failed it every time.

A sleep in the handler is still a race — it has to outlast whatever else the page is doing, and the duration that wins locally is not the one that wins on a loaded CI runner. When the spec can observe the state the module must arrive after, block the handler on a Queue and release it from the example instead:

ruby
playwright_page.context.route(%r{parking_notification_form_controller}, proc { |route, request|
  held << request.url
  release.pop
  route.continue
})
# ...only the accordion can reveal the panel, so this is it having already opened
expect(page).to have_content("Set on map", wait: 10)
release << :continue

Blocking the handler doesn't stall the driver, so Capybara still polls while it waits.

Caveat when measuring locally: after a heavy record-creating run (seeding, probe scripts, a big suite), :js specs fail spuriously for a while. Re-measure in a quiet environment before concluding a spec is flaky.

5. Blaming your own change needs both arms measured together

"Is this spec flaky?" and "did my change make it flaky?" are different questions. The second one is where sequential sampling lies to you: local load drifts — a browser left open, another suite, the machine waking up — so a sample taken now and one taken an hour ago aren't comparable, and whichever arm ran while things were busy looks guilty.

Revert only the suspect change and run both arms the same number of times, back to back, then compare. Measured that way here, a patch blamed for :js failures on sequential samples (0 failures in 9 clean runs against 5 in 13 patched ones) came out at 2/6 versus the baseline's 1/6 — indistinguishable, and the spec was flaky on its own. Sample sizes this small can't separate a 17% failure rate from a 7% one, so treat a handful of green runs as weak evidence in either direction.

Ruling ordering out is cheap and worth doing first: RSpec prints Randomized with seed N, and --seed N replays that order. A failing seed that passes on replay leaves timing, not ordering or leaked state.

Measure the "without the fix" arm against the base ref by name — git checkout origin/main -- <paths>. Once the fix is committed, git checkout -- <paths> restores it, so the arm you think is unpatched is the patched one, and a regression test that does fail without the fix reads as passing.

Show full SKILL.md (1,117 more words)Show less

Known causes in this repo

Work through these before inventing a new theory — most flakes here are one of them, and several look like timing but aren't.

Not actually flaky — the environment is wrong. A missing app/assets/builds/tailwind.css makes tw:hidden silently not apply, so visibility assertions fail in ways that read as flakes — bin/rails tailwindcss:build (see sandbox-test-setup). Same class of thing: an unmigrated test DB, a stale VCR cassette.

A build that's present but predates a merge fails the same way and reads worse, because the class the failing spec needs is in the source and the whole suite is otherwise green — Tailwind only generates what the content scan saw, so a class arriving with the merge (tw:h-64, used by one preview) isn't in a build from before it. bin/dev down means no watcher, so its app/assets/builds/*.css mtime against the merge commit's is the check; bin/rails tailwindcss:build is the fix. Deterministic, not intermittent: three identical failures with no ordering component is this rather than a flake.

Shared state across examples. The autocomplete cache (autc:test:*) lives in a Redis DB shared across :js examples and survives 600s, and load_all never invalidates it — so a stale entry from an earlier spec changes what a combobox returns. The fix is Autocomplete::Loader.clear_redis in before, not a retry. Browser history looks like the same shape and isn't: the driver closes the browser context between examples, so no entry outlives one. A spec that resets history is treating its own earlier steps as contamination — walk them the way a reader would instead, and note that entry 0 of every example is about:blank, which a traversal lands on as a nil current_path.

Interacting with a page whose controllers haven't connected. application.js lazy loads every Stimulus controller, so a freshly rendered page answers to none of them until each module lands: a combobox filters nothing, a one-shot event (like form-persist's restore) reaches no listener, and a fill_in's text can end up in whatever autofocus left focused. Waiting on any one controller proves nothing about the rest — wait_for_stimulus (spec/support/integration_spec_helpers.rb) waits for every identifier the page names, and pass it the one you're about to interact with (wait_for_stimulus("shared-blocks--navbar")): bare, it is vacuously true on a document that has parsed none yet, so it returns before that element even exists. A reload of a form with a saved draft runs the other way: form-persist's restore can land after the example has checked or typed, and puts the draft back over it — a tick after connect, so wait_for_stimulus doesn't cover it. Wait for a value the draft restores (have_field(..., with:)) first; register_organized_spec's single-page example is the pattern.

Interacting before the legacy page script has bound. The same shape, one era back: init.coffee's loadPageScript constructs the per-page class in $(document).ready, while click_link returns with the new document still parsing — so an interaction landing between the two is swallowed with nothing on the page to say so. wait_for_page_script (spec/support/integration_spec_helpers.rb) waits on window.pageScript; reach for it after any navigation into a jQuery-driven control.

Filling a field while a turbo-stream renders. Every stream render runs through Turbo's withPreservedFocus, which puts focus back a frame later on whatever held it when the render began — so a fill_in landing in that frame types into the field before it (the failure reads as one field empty and its neighbour holding both values). A multiselect pick's chip is one such stream: wait for it, as combobox_select in spec/components/pages/search/form/component_system_spec.rb does. An async combobox pick sends a filter request whose response can land several steps later; click_combobox_option waits that one out.

Clicking something that is being re-rendered. The dominant :js flake. A Turbo frame that reloads (an eager frame, reloadFrameIfUrlStale on turbo:load, a broadcast morph) detaches the element mid-click, and the click lands nowhere. Fix it by waiting for the settled state the user would wait for, then clicking:

ruby
expect(page).to have_css("turbo-frame#results_frame[complete]:not([busy])", wait: 10)
retry_on_detach { first(".bike-box-item .title-link a").click }

retry_on_detach (spec/support/integration_spec_helpers.rb) rescues the raw Playwright::Error for a detached node, which Capybara's own retry does not. This is not a coverage reduction: the assertions are untouched, the click just happens on a DOM that has stopped moving.

A probe that measures itself instead of the app. When a spec observes behaviour by listening for an event, check whether what it reads is a property of the app or of the listener's position. event.defaultPrevented read inside a document listener only reflects preventDefault calls from listeners that already ran, so it reports registration order as much as the app's verdict. Measured in this repo, same event, same run: read in the listener false, read on the next tick true. Defer the read so every listener has had its turn:

js
document.addEventListener("turbo:before-fetch-response", (event) => {
  if (event.target?.id !== "results_frame") return
  // Next tick: the verdict no longer depends on where this sits in the order
  setTimeout(() => { document.body.dataset.testRejected = event.defaultPrevented ? "true" : "false" })
})

Order is easy to invert without noticing, because the two sides have different lifetimes: document listeners survive a Turbo Drive body swap, while Stimulus controllers disconnect and re-add theirs at the end of the list on reconnect. So a probe registered before a Turbo visit ends up ahead of the controller it's watching. Two habits that keep this class of bug visible: record the verdict for both outcomes ("true"/"false") rather than only writing the marker on success, so a wrong verdict fails loudly instead of timing out as if the event never arrived; and prefer asserting the user-visible consequence alongside the internal verdict.

Programmatic back/forward. Playwright drives these specs and there is no BFCache, so go_back/go_forward onto a turbo-action: advance entry is genuinely unreliable. Prefer not to chain a real navigation onto the tail of a back/forward sequence. If a spec must, expect to need the settle-then-click pattern above.

What a finished fix looks like

  • The change names a mechanism ("the frame reloads on turbo:load and detaches the link"), not a symptom ("this is flaky on CI").
  • You can explain why the fix addresses that mechanism, even though you probably can't reproduce the original failure locally.
  • Or — where you fell back to a retry or a longer wait — you say so as the headline rather than the footnote, and the comment in the spec records what you ruled out, so the next person starts where you stopped.
  • Comments describing the flake are corrected if your diagnosis contradicts them — a wrong comment sends the next person down the same wrong path.
  • CI is the only real verification; local green doesn't prove it.

Working on an already-tagged spec

flaky: retries only run on CI (RETRY_FLAKY, see spec/rails_helper.rb). flaky: true retries twice; flaky: <n> overrides the count. So a plain bundle exec rspec doesn't retry (bin/ci sets RETRY_FLAKY), and a flaky:-tagged spec failing once locally is not automatically "the known flake" — it may be a plain reproducible failure that the tag has been hiding on CI.

Removing a now-unnecessary flaky: tag after you've fixed the cause is good housekeeping — that direction adds signal rather than removing it.

© bikeindex, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fixing-flaky-failures of bikeindex/bike_index.

Open the folder on GitHubat commit b9be85f

Compare with similar skills

Fixing Flaky Failures next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fixing Flaky Failures compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fixing Flaky Failures this skillbikeindex/bike_index308—~5kAutomated safety check: PassAGPL-3.0
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Triage CI FailureDataDog/datadog-agent3.8k—~2.3kAutomated safety check: PassApache-2.0
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Fix Ready PRsfastrepl/anarlog9.5k—~1.4kAutomated safety check: PassMIT
Trx Analysismicrosoft/vstest969—~1.8kAutomated safety check: PassMIT

Similar skills

  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Triage CI Failure

    DataDog/datadog-agent

    Official

    Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

    3.8k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Fix Ready PRs

    fastrepl/anarlog

    Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.

    9.5k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Trx Analysis

    microsoft/vstest

    Official

    Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.

    969 GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    204 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from bikeindex/bike_index

All 11 skills in this repo
  • Admin Data API

    bikeindex/bike_index

    Read live production data from Bike Index through the admin OAuth token — Sidekiq and PgHero status, and the user-submitted bug reports — the same data as the cookie-gated dashboards, but…

    308 GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • GitHub PR Images

    bikeindex/bike_index

    Embed a local image file into an existing GitHub PR — either in the PR body or as a comment.

    308 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Manufacturers

    bikeindex/bike_index

    Add a manufacturer to Bike Index in production through the admin OAuth token (POST /admin/manufacturers).

    308 GitHub stars~452 tokensUpdated today
    Auto-check passed
  • PR

    bikeindex/bike_index

    Create or update a pull request for the current branch. An agent skill from bikeindex/bike_index.

    308 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Honeybadger Debugging

    bikeindex/bike_index

    Investigate and fix a specific Honeybadger exception in the Bike Index app — pull the fault, read its backtrace, find the offending code, write the fix.

    308 GitHub stars~1k tokensUpdated today
    Auto-check: notes
  • Merge Conflicts

    bikeindex/bike_index

    How to bring a branch up to date and resolve git merge conflicts the way this repo expects — merge (never rebase or force-push), understand each side's intent before choosing, ask when a resolution…

    308 GitHub stars~3.2k tokensUpdated today
    Auto-check: notes

Categories

Questions about Fixing Flaky Failures

What does Fixing Flaky Failures do?

How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky. Fixing Flaky Failures is an agent skill from bikeindex/bike_index. How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky.

When should I use Fixing Flaky Failures?

Fixing Flaky Failures fits situations like: A CI failure the user calls flaky; passes on retry; A pasted gh run/Actions URL whose failure isnt reproducible; also whenever you catch yourself about to delete.

How do I install Fixing Flaky Failures in Claude Code?

Run `npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a claude-code`. Or copy the skill folder (.claude/skills/fixing-flaky-failures in bikeindex/bike_index) into .claude/skills/fixing-flaky-failures in your project. Claude Code loads it when a task matches its description.

How do I install Fixing Flaky Failures in Codex?

Run `npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a codex`. Or copy the skill folder (.claude/skills/fixing-flaky-failures in bikeindex/bike_index) into .agents/skills/fixing-flaky-failures in your project. Codex loads it when a task matches its description.

Can I use Fixing Flaky Failures in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fixing-flaky-failures, .gemini/skills/fixing-flaky-failures, .github/skills/fixing-flaky-failures and .opencode/skills/fixing-flaky-failures in your project.

What does Fixing Flaky Failures need to run?

Going by SKILL.md and its folder, Fixing Flaky Failures needs the command-line tools its instructions call (gh, rails and bundle).

Does Fixing Flaky Failures access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fixing Flaky Failures safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fixing Flaky Failures use?

Fixing Flaky Failures is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fixing Flaky Failures use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fixing Flaky Failures?

Skills that share tags, products or a category with Fixing Flaky Failures: Swig Test (swig/swig, 6.3k stars), Triage CI Failure (DataDog/datadog-agent, 3.8k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars) and Fix Ready PRs (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fixing Flaky Failures?

bikeindex (a GitHub organization) maintains it in bikeindex/bike_index, which has 308 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 10, 2026.

Source: bikeindex/bike_index on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.