Swig Test
swig/swig
Run SWIG test suite for specific languages. An agent skill from swig/swig.
How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky.
$ npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bikeindex/bike_index fixing-flaky-failures --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/bikeindex/bike_index.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fixing-flaky-failures .claude/skills/fixing-flaky-failures && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "fixing-flaky-failures" agent skill from https://github.com/bikeindex/bike_index/tree/main/.claude/skills/fixing-flaky-failures into .claude/skills/fixing-flaky-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-failures", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/bikeindex/bike_index/tree/main/.claude/skills/fixing-flaky-failuresType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bikeindex/bike_index fixing-flaky-failures --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bikeindex/bike_index.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/fixing-flaky-failures .agents/skills/fixing-flaky-failures && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "fixing-flaky-failures" agent skill from https://github.com/bikeindex/bike_index/tree/main/.claude/skills/fixing-flaky-failures into .agents/skills/fixing-flaky-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-failures", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bikeindex/bike_index fixing-flaky-failures --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bikeindex/bike_index.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/fixing-flaky-failures .cursor/skills/fixing-flaky-failures && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "fixing-flaky-failures" agent skill from https://github.com/bikeindex/bike_index/tree/main/.claude/skills/fixing-flaky-failures into .cursor/skills/fixing-flaky-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-failures", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/bikeindex/bike_index.git --path .claude/skills/fixing-flaky-failures--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bikeindex/bike_index fixing-flaky-failures --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bikeindex/bike_index.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/fixing-flaky-failures .gemini/skills/fixing-flaky-failures && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "fixing-flaky-failures" agent skill from https://github.com/bikeindex/bike_index/tree/main/.claude/skills/fixing-flaky-failures into .gemini/skills/fixing-flaky-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-failures", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bikeindex/bike_index fixing-flaky-failuresInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/bikeindex/bike_index.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/fixing-flaky-failures .github/skills/fixing-flaky-failures && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "fixing-flaky-failures" agent skill from https://github.com/bikeindex/bike_index/tree/main/.claude/skills/fixing-flaky-failures into .github/skills/fixing-flaky-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-failures", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bikeindex/bike_index fixing-flaky-failures --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bikeindex/bike_index.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/fixing-flaky-failures .opencode/skills/fixing-flaky-failures && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "fixing-flaky-failures" agent skill from https://github.com/bikeindex/bike_index/tree/main/.claude/skills/fixing-flaky-failures into .opencode/skills/fixing-flaky-failures/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "fixing-flaky-failures", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
fixing-flaky-failuresHow to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky.
Fixing Flaky Failures is an agent skill from bikeindex/bike_index. How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky. Read this before touching any spec that is failing intermittently, and before adding, raising, or relying on a flaky: tag or a wait: bump. Trigger on a CI failure the user calls flaky, unreliable, intermittent, "green locally", "passes on retry", "fix CI", or a pasted gh run/Actions URL whose failure isn't reproducible; also whenever you catch…
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Failing and flaky tests. The repository describes itself as: All the code for Bike Index, because we love you. The licence is AGPL-3.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b9be85f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghrailsbundleFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Fixing Flaky Failures loads about 5k tokens when it runs. Until then it costs about 227 tokens; SKILL.md has 2,672 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from bikeindex/bike_index at commit b9be85f, republished under its AGPL-3.0 licence (© bikeindex). 2,672 words, ~5,042 tokens.
.claude/skills/fixing-flaky-failures/SKILL.md (or your agent's skills folder).Covers diagnosing the real mechanism, attributing a flake to a change, this
repo's flaky: retry harness, and the known false-flake causes — a missing
Tailwind build, the shared Redis autocomplete cache, Turbo frame timing,
probe-run interference.
You may not reduce test coverage to make a flaky test pass. Not as a last resort, not "temporarily", not when the coverage looks redundant, not when you have already decided the failing step is a harness artifact rather than a real bug.
All of these are reducing coverage, and none of them is a fix:
have_css(count: 2) → have_css,
an exact string → a regex, have_no_content → nothing).skip/pending/xit, or excluding the spec from a CI shard.A flaky test is a reporting problem — the suite is telling you something real and telling you unreliably. Every item above changes the reporting and leaves the underlying behaviour untested.
If the only fix you can see requires giving up coverage, that is a decision for the user — describe what you'd have to give up and ask. Do not make that trade yourself and mention it in the summary afterwards; announcing a scope reduction is not the same as getting agreement for it.
The one exception that isn't an exception: if the assertion is wrong — it asserts behaviour the app never promised — then fixing it is correcting a bad test, not reducing coverage. Say explicitly why it was wrong.
:flaky (and raising an existing flaky: N) and a bigger wait: are a
different category from the list above, because they keep every assertion and
still require it to pass. Nothing goes untested. They're a legitimate last
resort — reach for them when the diagnosis below has genuinely run out, not as
the first thing you try.
What "earn them" means in practice:
flaky: 4 on search_registrations_spec is the pattern: it names the click
landing on about:blank, lists what it ruled out, and says it went
unreproduced under local CPU throttling. A bare flaky: true with no comment tells the
next person nothing and will outlive the problem.Two signals that a retry is the wrong answer even as a last resort: an example that fails through all its retries (the cause is structural, and a bigger N won't help), and a wait you can't say what it's waiting for. A bump is honest when the assertion starts before the work does — a held route released, a job enqueued, an expensive response — and the comment names that.
gh run view <run-id> --repo bikeindex/bike_index --json jobs \
--jq '.jobs[] | "\(.databaseId) \(.name) \(.conclusion)"'
gh run view --repo bikeindex/bike_index --job <job-id> --log-failedThe log is large and ANSI-coloured; pipe through sed 's/\x1b\[[0-9;]*m//g'
and grep for Failure/Error, expected, and rspec ./spec/... to get the
failing example ids and the actual message.
Then download the Capybara screenshot. Every :js failure writes one, and
ci.yml uploads tmp/capybara/ to the shard's test-results-<node-index>
artifact for exactly this — but the log names the file without saying it's
fetchable, so it usually goes unread. It shows the page the failure saw, which
routinely settles a mechanism the log can only hint at: what a click actually
landed on, a frame still loading, a control in a state no step in the example set.
gh api repos/bikeindex/bike_index/actions/runs/<run-id>/artifacts \
--jq '.artifacts[] | "\(.id) \(.name)"'
gh api repos/bikeindex/bike_index/actions/artifacts/<id>/zip > tmp/a.zip && unzip -o tmp/a.zip -d tmp/ci_artifactThe exact failure text usually names the mechanism, and it is easy to skim past
into a wrong assumption. Worked example from this repo: expected nil to match /\/bikes\/\d+/ was long assumed to mean "the click was lost". It doesn't — a nil
current_path means the URL had no path for Capybara to return, which is three
different browser states, not one (capybara/session.rb:207: nil for an about:
scheme, then path unless path&.empty?):
| URL | how it got there |
|---|---|
about:blank | traversed to entry 0, or the page was replaced |
chrome-error://chromewebdata | a cross-document navigation failed outright |
"" | no document has committed yet |
All three screenshot blank, so the picture can't tell them apart — tmp/capybara/browser_events.log
(written by spec/support/capybara.rb, uploaded with the screenshots) can. Check the
matcher's source when a message is surprising, and don't read one of these three as
another: the fix differs, and "about:blank" has been the standing wrong guess.
When you can't reproduce, you can still make the browser tell you what happened. Copy the spec to a scratch file, record the events the app actually emits, and run it — a log beats an argument, and this routinely overturns the theory you were about to ship:
page.execute_script(<<~JS)
window.__events = []
const stamp = (name, extra) => window.__events.push(Object.assign({name, t: Math.round(performance.now())}, extra))
document.addEventListener("turbo:before-fetch-response", (e) => {
stamp("response", {target: e.target?.id, url: e.detail?.fetchResponse?.response?.url, prevented: e.defaultPrevented})
})
JS
# ...drive the spec...
# A file, not puts - rtk's rspec wrapper reports a summary and drops the run's stdout
File.open("tmp/probe.log", "a") { |f| page.evaluate_script("window.__events").each { |e| f.puts e.inspect } }Parameterise the scratch spec over the variable you suspect ([0, 8].each do |delay|
around an injected sleep) so one run compares the fast and slow paths. Delete the
scratch file when you're done. Two things this buys that reasoning doesn't: it
distinguishes "the event never arrived" from "the event arrived and the assertion
misread it", and it tells you when things happened, which is usually the answer.
for i in 1 2 3; do bundle exec rspec <the spec file> 2>&1 | grep -E "examples, " | tail -1; doneKeep that loop on the one spec file — escalating it to bin/ci costs minutes of
parallel workers and browsers per iteration, and answers the same question no better.
Green locally three times doesn't mean "not reproducible, add a retry". It narrows the cause to something CI has and you don't: contention (browser, Rails and Postgres sharing one runner) or ordering (a different seed, knapsack handing this shard a different set of files, or state left by another example). Reason about which, then look for the mechanism.
For contention, slow the renderer rather than the machine — CPU hogs slow the Ruby
side too, so a loop of runs takes minutes and the extra load is spent where the race
isn't, and one backgrounded from a non-interactive shell survives kill $(jobs -p) —
pgrep -f it. CDP throttles the browser alone, and the driver hands you a session:
page.driver.with_playwright_page do |playwright_page|
session = playwright_page.context.new_cdp_session(playwright_page)
session.send_message("Emulation.setCPUThrottlingRate", params: {rate: 6})
endA rate that leaves the spec green over ~20 runs is evidence, not proof: it stretches main-thread work, not the network or a parallel shard's I/O.
Which is why it can't reach a race about when a response arrives. The common one here is a lazily loaded Stimulus controller, since the module is a fetch — hold it on the route and the late connect is deterministic, no loop:
playwright_page.route(%r{serial_controller}, ->(route, request) { held << request.url; sleep 1; route.continue })Registering a route disables the http cache, so the reload asks again. Assert on what
the handler held: a route that stops matching (a moved asset path) otherwise leaves the
example green on a page that held nothing back. Measured against register--serial's
connect-time reconcile, throttling at 6, 12 and 25 left it green over 30+ runs; the
route hold failed it every time.
A sleep in the handler is still a race — it has to outlast whatever else the page is
doing, and the duration that wins locally is not the one that wins on a loaded CI
runner. When the spec can observe the state the module must arrive after, block the
handler on a Queue and release it from the example instead:
playwright_page.context.route(%r{parking_notification_form_controller}, proc { |route, request|
held << request.url
release.pop
route.continue
})
# ...only the accordion can reveal the panel, so this is it having already opened
expect(page).to have_content("Set on map", wait: 10)
release << :continueBlocking the handler doesn't stall the driver, so Capybara still polls while it waits.
Caveat when measuring locally: after a heavy record-creating run (seeding,
probe scripts, a big suite), :js specs fail spuriously for a while. Re-measure
in a quiet environment before concluding a spec is flaky.
"Is this spec flaky?" and "did my change make it flaky?" are different questions. The second one is where sequential sampling lies to you: local load drifts — a browser left open, another suite, the machine waking up — so a sample taken now and one taken an hour ago aren't comparable, and whichever arm ran while things were busy looks guilty.
Revert only the suspect change and run both arms the same number of times, back
to back, then compare. Measured that way here, a patch blamed for :js failures
on sequential samples (0 failures in 9 clean runs against 5 in 13 patched ones)
came out at 2/6 versus the baseline's 1/6 — indistinguishable, and the spec was
flaky on its own. Sample sizes this small can't separate a 17% failure rate from
a 7% one, so treat a handful of green runs as weak evidence in either direction.
Ruling ordering out is cheap and worth doing first: RSpec prints Randomized with seed N, and --seed N replays that order. A failing seed that passes on replay
leaves timing, not ordering or leaked state.
Measure the "without the fix" arm against the base ref by name — git checkout origin/main -- <paths>. Once the fix is committed, git checkout -- <paths> restores
it, so the arm you think is unpatched is the patched one, and a regression test that
does fail without the fix reads as passing.
Work through these before inventing a new theory — most flakes here are one of them, and several look like timing but aren't.
Not actually flaky — the environment is wrong. A missing
app/assets/builds/tailwind.css makes tw:hidden silently not apply, so
visibility assertions fail in ways that read as flakes —
bin/rails tailwindcss:build (see sandbox-test-setup).
Same class of thing: an unmigrated test DB, a stale VCR cassette.
A build that's present but predates a merge fails the same way and reads worse, because
the class the failing spec needs is in the source and the whole suite is otherwise green —
Tailwind only generates what the content scan saw, so a class arriving with the merge
(tw:h-64, used by one preview) isn't in a build from before it. bin/dev down means no
watcher, so its app/assets/builds/*.css mtime against the merge commit's is the check;
bin/rails tailwindcss:build is the fix. Deterministic, not intermittent: three identical
failures with no ordering component is this rather than a flake.
Shared state across examples. The autocomplete cache (autc:test:*) lives
in a Redis DB shared across :js examples and survives 600s, and load_all
never invalidates it — so a stale entry from an earlier spec changes what a
combobox returns. The fix is Autocomplete::Loader.clear_redis in before,
not a retry. Browser history looks like the same shape and isn't: the driver closes
the browser context between examples, so no entry outlives one. A spec that resets
history is treating its own earlier steps as contamination — walk them the way a
reader would instead, and note that entry 0 of every example is about:blank, which
a traversal lands on as a nil current_path.
Interacting with a page whose controllers haven't connected. application.js
lazy loads every Stimulus controller, so a freshly rendered page answers to none of
them until each module lands: a combobox filters nothing, a one-shot event (like
form-persist's restore) reaches no listener, and a fill_in's text can end up in
whatever autofocus left focused. Waiting on any one controller proves nothing about
the rest — wait_for_stimulus (spec/support/integration_spec_helpers.rb) waits for
every identifier the page names, and pass it the one you're about to interact with
(wait_for_stimulus("shared-blocks--navbar")): bare, it is vacuously true on a document
that has parsed none yet, so it returns before that element even exists.
A reload of a form with a saved draft runs the other way: form-persist's restore can land
after the example has checked or typed, and puts the draft back over it — a tick after
connect, so wait_for_stimulus doesn't cover it. Wait for a value the draft restores
(have_field(..., with:)) first; register_organized_spec's single-page example is the pattern.
Interacting before the legacy page script has bound. The same shape, one era
back: init.coffee's loadPageScript constructs the per-page class in
$(document).ready, while click_link returns with the new document still
parsing — so an interaction landing between the two is swallowed with nothing on
the page to say so. wait_for_page_script
(spec/support/integration_spec_helpers.rb) waits on window.pageScript; reach for
it after any navigation into a jQuery-driven control.
Filling a field while a turbo-stream renders. Every stream render runs through Turbo's
withPreservedFocus, which puts focus back a frame later on whatever held it when the render
began — so a fill_in landing in that frame types into the field before it (the failure reads as
one field empty and its neighbour holding both values). A multiselect pick's chip is one such
stream: wait for it, as combobox_select in spec/components/pages/search/form/component_system_spec.rb
does. An async combobox pick sends a filter request whose response can land several steps later;
click_combobox_option waits that one out.
Clicking something that is being re-rendered. The dominant :js flake.
A Turbo frame that reloads (an eager frame, reloadFrameIfUrlStale on
turbo:load, a broadcast morph) detaches the element mid-click, and the click
lands nowhere. Fix it by waiting for the settled state the user would wait for,
then clicking:
expect(page).to have_css("turbo-frame#results_frame[complete]:not([busy])", wait: 10)
retry_on_detach { first(".bike-box-item .title-link a").click }retry_on_detach (spec/support/integration_spec_helpers.rb) rescues the raw
Playwright::Error for a detached node, which Capybara's own retry does not.
This is not a coverage reduction: the assertions are untouched, the click just
happens on a DOM that has stopped moving.
A probe that measures itself instead of the app. When a spec observes
behaviour by listening for an event, check whether what it reads is a property of
the app or of the listener's position. event.defaultPrevented read inside a
document listener only reflects preventDefault calls from listeners that
already ran, so it reports registration order as much as the app's verdict.
Measured in this repo, same event, same run: read in the listener false, read on
the next tick true. Defer the read so every listener has had its turn:
document.addEventListener("turbo:before-fetch-response", (event) => {
if (event.target?.id !== "results_frame") return
// Next tick: the verdict no longer depends on where this sits in the order
setTimeout(() => { document.body.dataset.testRejected = event.defaultPrevented ? "true" : "false" })
})Order is easy to invert without noticing, because the two sides have different
lifetimes: document listeners survive a Turbo Drive body swap, while Stimulus
controllers disconnect and re-add theirs at the end of the list on reconnect. So
a probe registered before a Turbo visit ends up ahead of the controller it's
watching. Two habits that keep this class of bug visible: record the verdict for
both outcomes ("true"/"false") rather than only writing the marker on success,
so a wrong verdict fails loudly instead of timing out as if the event never
arrived; and prefer asserting the user-visible consequence alongside the internal
verdict.
Programmatic back/forward. Playwright drives these specs and there is no
BFCache, so go_back/go_forward onto a turbo-action: advance entry is
genuinely unreliable. Prefer not to chain a real navigation onto the tail of a
back/forward sequence. If a spec must, expect to need the settle-then-click
pattern above.
turbo:load and detaches
the link"), not a symptom ("this is flaky on CI").flaky: retries only run on CI (RETRY_FLAKY, see spec/rails_helper.rb).
flaky: true retries twice; flaky: <n> overrides the count. So a plain
bundle exec rspec doesn't retry (bin/ci sets RETRY_FLAKY), and a flaky:-tagged spec failing once locally is not
automatically "the known flake" — it may be a plain reproducible failure that
the tag has been hiding on CI.
Removing a now-unnecessary flaky: tag after you've fixed the cause is good
housekeeping — that direction adds signal rather than removing it.
© bikeindex, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/fixing-flaky-failures of bikeindex/bike_index.
Open the folder on GitHubat commit b9be85f
Fixing Flaky Failures next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Fixing Flaky Failures this skillbikeindex/bike_index | 308 | — | ~5k | Automated safety check: Pass | AGPL-3.0 | |
| Swig Testswig/swig | 6.3k | — | ~2.3k | Automated safety check: Pass | Custom licence | |
| Triage CI FailureDataDog/datadog-agent | 3.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Dynamo Jira TicketDynamoDS/Dynamo | 2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Fix Ready PRsfastrepl/anarlog | 9.5k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Trx Analysismicrosoft/vstest | 969 | — | ~1.8k | Automated safety check: Pass | MIT |
swig/swig
Run SWIG test suite for specific languages. An agent skill from swig/swig.
DataDog/datadog-agent
Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.
DynamoDS/Dynamo
Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.
fastrepl/anarlog
Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.
microsoft/vstest
Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.
workersio/skills
Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.
bikeindex/bike_index
Read live production data from Bike Index through the admin OAuth token — Sidekiq and PgHero status, and the user-submitted bug reports — the same data as the cookie-gated dashboards, but…
bikeindex/bike_index
Embed a local image file into an existing GitHub PR — either in the PR body or as a comment.
bikeindex/bike_index
Add a manufacturer to Bike Index in production through the admin OAuth token (POST /admin/manufacturers).
bikeindex/bike_index
Create or update a pull request for the current branch. An agent skill from bikeindex/bike_index.
bikeindex/bike_index
Investigate and fix a specific Honeybadger exception in the Bike Index app — pull the fault, read its backtrace, find the offending code, write the fix.
bikeindex/bike_index
How to bring a branch up to date and resolve git merge conflicts the way this repo expects — merge (never rebase or force-push), understand each side's intent before choosing, ask when a resolution…
Categories
How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky. Fixing Flaky Failures is an agent skill from bikeindex/bike_index. How to fix a test that fails intermittently in Bike Index — one that passes locally but fails on CI, fails on one shard, passes on re-run, or is already tagged :flaky.
Fixing Flaky Failures fits situations like: A CI failure the user calls flaky; passes on retry; A pasted gh run/Actions URL whose failure isnt reproducible; also whenever you catch yourself about to delete.
Run `npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a claude-code`. Or copy the skill folder (.claude/skills/fixing-flaky-failures in bikeindex/bike_index) into .claude/skills/fixing-flaky-failures in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a codex`. Or copy the skill folder (.claude/skills/fixing-flaky-failures in bikeindex/bike_index) into .agents/skills/fixing-flaky-failures in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bikeindex/bike_index --skill fixing-flaky-failures -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fixing-flaky-failures, .gemini/skills/fixing-flaky-failures, .github/skills/fixing-flaky-failures and .opencode/skills/fixing-flaky-failures in your project.
Going by SKILL.md and its folder, Fixing Flaky Failures needs the command-line tools its instructions call (gh, rails and bundle).
SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Fixing Flaky Failures is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Fixing Flaky Failures: Swig Test (swig/swig, 6.3k stars), Triage CI Failure (DataDog/datadog-agent, 3.8k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars) and Fix Ready PRs (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bikeindex (a GitHub organization) maintains it in bikeindex/bike_index, which has 308 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 10, 2026.
Source: bikeindex/bike_index on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.