Agent skill

Flake Triage

by yschimke in yschimke/compose-ai-tools

Decide whether a preview the visual-diff bot flagged actually regressed or is simply nondeterministic, using a repeat-render oracle at a single commit.

Apache-2.0Auto-check passedMobile

Install Flake Triage

skills CLI
$ npx skills add yschimke/compose-ai-tools --skill flake-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yschimke/compose-ai-tools flake-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yschimke/compose-ai-tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flake-triage .claude/skills/flake-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flake-triage
GitHub stars
117
Token cost
~1.5k tokens
SKILL.md length
706 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Decide whether a preview the visual-diff bot flagged actually regressed or is simply nondeterministic, using a repeat-render oracle at a single commit.

  • Works in 3 steps: Re-run the oracle. Consecutive --rerun… → Add the pixel test that keeps it pinned… → Commit the evidence — the two differing…
  • A PRs preview diff reports a changed preview whose source the PR does not touch
  • SKILL.md covers The oracle, Localising an unstable render and Fixing it, and proving the fix
  • Calls node and npx

What it does

Flake Triage is an agent skill from yschimke/compose-ai-tools. Decide whether a preview the visual-diff bot flagged actually regressed or is simply nondeterministic, using a repeat-render oracle at a single commit. Use when a PR's preview diff reports a changed preview whose source the PR does not touch, or when a render, GIF or filmstrip is suspected of being unstable.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Mobile, covering Visual regression testing and Android development. It works with Jetpack Compose, Gradle, Visual Studio Code and Kotlin. The repository describes itself as: Helping the Agents Compose the Things. The licence is Apache-2.0.

When your agent uses it

  • A PRs preview diff reports a changed preview whose source the PR does not touch
  • Filmstrip is suspected of being unstable

Example prompts

  • “/flake-triage”

Requirements

  • Node.js

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Re-run the oracle. Consecutive --rerun renders should be byte-identical;
  2. Add the pixel test that keeps it pinned (SharedElementFilmstripPixelTest is the
  3. Commit the evidence — the two differing before-renders and the stable after —

What it can do on your machine

Read from SKILL.md and the folder at commit dd6fb61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flake Triage loads about 1.5k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 706 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yschimke/compose-ai-tools at commit dd6fb61, republished under its Apache-2.0 licence (© yschimke). 706 words, ~1,505 tokens.

Download SKILL.mdSave it as .claude/skills/flake-triage/SKILL.md (or your agent's skills folder).
name
flake-triage
description
Decide whether a preview the visual-diff bot flagged actually regressed or is simply nondeterministic, using a repeat-render oracle at a single commit. Use when a PR's preview diff reports a changed preview whose source the PR does not touch, or when a render, GIF or filmstrip is suspected of being unstable.

Is it a regression, or is the preview unstable?

A changed preview whose source the PR does not touch — clocks, timestamps, randomness, animation frames, network-loaded images — is instability to triage, not a regression to rubber-stamp and not a fix to write blind. The distinction is decidable in a few minutes, at one commit, with no reference to the base branch.

The oracle

Render the same preview N times at the same commit and compare the bytes. If two renders of one commit differ, nothing about the PR's diff is evidence of anything.

for i in 1 2 3; do
  ./gradlew :samples:android:composePreviewRender --rerun \
    -PcomposePreview.filter=<PreviewFunctionName>
  cp <module>/build/compose-previews/renders/<Preview>.png run-$i.png
done
md5sum run-*.png

--rerun is load-bearing: a render task goes UP-TO-DATE off its own outputs (including .error.json sidecars from a failed run), so a plain re-invocation compares a file against itself and always "passes".

A harness capture has the same oracle, and it is cheaper. The captures the visual-diff bot compares are Playwright shots of committed static fixtures, so N runs cost seconds and need no JVM.

The serve captures (serve-, viewer-) are not triaged here any more. That harness moved with the server to yschimke/compose-preview-server, where it runs as the visual-harness CI job; triage them in that repository, against its preview-harness/. What stays here is the fixture tree those captures read (preview-server/preview-harness/fixtures/, still asserted by ServeWebFixtureTest and friends) — a fixture edit here changes what that repository captures.

# compose-preview-vscode/preview-harness — the VS Code panel's own fixtures.
# Run from `compose-preview-vscode/`, not from inside the harness directory, and build
# the webview bundle first: the fixtures load it, so a stale one moves pixels for
# reasons no commit explains.
cd ../compose-preview-vscode
node esbuild.webview.mjs
HARNESS_CHROMIUM=/path/to/chromium HARNESS_FIXTURE=<fixture> HARNESS_THEME=light \
  npx playwright test -c preview-harness/playwright.config.mjs snapshot.spec.mjs
md5sum preview-harness/out/<capture>.light.png

HARNESS_THEME narrows to one theme; HARNESS_FIXTURE is the extension harness's own selector and is cleaner than -g there.

There is no --rerun equivalent to remember: each invocation rewrites out/.

Read the hashes:

  • All identical → the preview is deterministic at this commit. A diff against the base is real; go find it in the source.
  • They differ → the preview is unstable. The bot flagging it on unrelated PRs is a symptom, not a coincidence. Continue below.

Three runs is usually enough to catch it; five identical runs is a reasonable bar for calling a fix proven.

Localising an unstable render

Once you know it moves, find what moves:

  • Still PNGs — diff the pair as images rather than eyeballing them; the percentage of changed pixels and where they sit names the culprit (a clock face, one panel of a filmstrip, a gradient band).
  • GIFs and filmstrips — decode to frames and compare frame by frame. A filmstrip whose panels re-freeze somewhere new on each run differs in one to three panels while the rest is byte-identical, which reads as "small diff" until you split it. docs/design/evidence/filmstrip-determinism/ is the worked example: two same-commit renders differing across 2.99% / 7.97% / 10.96% of the image depending on the pair, root-caused to panels freezing at unpinned points rather than at their labelled transition fractions.
  • Animated images in a harness capture — an APNG or GIF a spec swaps in plays on its own clock, and img.complete says decoded, not finished. A shot held only for the decode lands on an arbitrary frame. docs/design/evidence/motion-index-playing-determinism/ is the worked example: eight same-commit runs, three distinct hashes, 0.11% of the image moving in one 19-row band per card. Decode the stub's acTL/fcTL chunks to learn its frame count and delays, then hold for rest — poll the container's pixels until two reads spaced wider than one frame agree, rather than hard-coding the duration.
Show full SKILL.md (180 more words)Show less

The usual sources, in rough order of frequency: an unpinned animation clock, a system-time variable, unseeded randomness, a network- or disk-loaded image, and layout that depends on measurement order.

Fixing it, and proving the fix

Pin the nondeterminism at its source — hold the clock, seed the randomness, pin the panel to its labelled fraction — rather than loosening a comparison threshold. Then:

  1. Re-run the oracle. Consecutive --rerun renders should be byte-identical; quote the shared hash.
  2. Add the pixel test that keeps it pinned (SharedElementFilmstripPixelTest is the pattern) so the next regression fails a test rather than a review.
  3. Commit the evidence — the two differing before-renders and the stable after — under docs/design/evidence/<slug>/ with a README stating the commit, the command, and the hashes. That is what makes the next occurrence a lookup instead of a rediscovery.

Reporting an unstable preview without the hashes is an assertion, not a triage. The hashes are cheap; include them.

Related: render-evidence for the capture mechanics, and the stability reference in the published compose-preview-review skill for how to flag instability on someone else's PR.

© yschimke, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/flake-triage of yschimke/compose-ai-tools.

Open the folder on GitHubat commit dd6fb61

Compare with similar skills

Flake Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flake Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flake Triage this skillyschimke/compose-ai-tools117—~1.5kAutomated safety check: PassApache-2.0
Android Developmentdpconde/claude-android-skill337—~1.7kAutomated safety check: PassMIT
Claude Android NinjaDrjacky/claude-android-ninja124—~5.2kAutomated safety check: PassApache-2.0
Composewebview Developmentparkwoocheol/compose-webview103—~1.3kAutomated safety check: PassMIT
Update Libs VersionsCosminMihuMDC/KtorMonitor254—~800Automated safety check: PassApache-2.0
Android DevelopmentCoWork-OS/CoWork-OS474—~459Automated safety check: PassMIT

Similar skills

  • Android Development

    dpconde/claude-android-skill

    Create production-quality Android applications following Google's official architecture guidance and NowInAndroid best practices.

    337 GitHub stars~1.7k tokensUpdated 10 mo ago
    MobileAuto-check passed
  • Claude Android Ninja

    Drjacky/claude-android-ninja

    Build and migrate Android apps with Kotlin, Jetpack Compose, MVVM, Hilt, Room 3 (KSP, SQLiteDriver, Flow/suspend DAOs), Navigation3, and multi-module Gradle.

    124 GitHub stars~5.2k tokensUpdated 9 days ago
    MobileAuto-check passed
  • Composewebview Development

    parkwoocheol/compose-webview

    Builds, tests, and formats ComposeWebView multiplatform library.

    103 GitHub stars~1.3k tokensUpdated 1 mo ago
    MobileAuto-check passed
  • Update Libs Versions

    CosminMihuMDC/KtorMonitor

    Update dependency versions in gradle/libs.versions.toml by querying Maven repositories.

    254 GitHub stars~800 tokensUpdated 26 days ago
    MobileAuto-check passed
  • Android Development

    CoWork-OS/CoWork-OS

    Android/Kotlin development: Jetpack Compose, Room database, Gradle builds, emulator management, and Play Store submission.

    474 GitHub stars~459 tokensUpdated today
    MobileAuto-check passed
  • Kotlin Android

    ericrisco/rsc-harness

    A skill your agent uses when building or fixing a native Android app in Kotlin and Jetpack Compose on the UDF layered architecture — ViewModel/StateFlow, Hilt, Room, Retrofit, coroutines, type-safe…

    174 GitHub stars~2.9k tokensUpdated yesterday
    MobileAuto-check passed

More from yschimke/compose-ai-tools

  • Render Evidence

    yschimke/compose-ai-tools

    Capture the before/after rendered PNGs a UI-affecting PR in this repository needs as visual evidence, including bringing an Android SDK up in a fresh sandbox.

    117 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Steward

    yschimke/compose-ai-tools

    Drive a pull request on this repository to green — which fast checks to run before pushing, how to read a red check, and what to do about a review comment.

    117 GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Flake Triage

What does Flake Triage do?

Decide whether a preview the visual-diff bot flagged actually regressed or is simply nondeterministic, using a repeat-render oracle at a single commit. Flake Triage is an agent skill from yschimke/compose-ai-tools. Decide whether a preview the visual-diff bot flagged actually regressed or is simply nondeterministic, using a repeat-render oracle at a single commit.

When should I use Flake Triage?

Flake Triage fits situations like: A PRs preview diff reports a changed preview whose source the PR does not touch; filmstrip is suspected of being unstable.

How do I install Flake Triage in Claude Code?

Run `npx skills add yschimke/compose-ai-tools --skill flake-triage -a claude-code`. Or copy the skill folder (.claude/skills/flake-triage in yschimke/compose-ai-tools) into .claude/skills/flake-triage in your project. Claude Code loads it when a task matches its description.

How do I install Flake Triage in Codex?

Run `npx skills add yschimke/compose-ai-tools --skill flake-triage -a codex`. Or copy the skill folder (.claude/skills/flake-triage in yschimke/compose-ai-tools) into .agents/skills/flake-triage in your project. Codex loads it when a task matches its description.

Can I use Flake Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yschimke/compose-ai-tools --skill flake-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flake-triage, .gemini/skills/flake-triage, .github/skills/flake-triage and .opencode/skills/flake-triage in your project.

What does Flake Triage need to run?

Going by SKILL.md and its folder, Flake Triage needs the command-line tools its instructions call (node and npx). Our summary lists: Node.js.

Does Flake Triage access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Flake Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Flake Triage use?

Flake Triage is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Flake Triage use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flake Triage?

Skills that share tags, products or a category with Flake Triage: Android Development (dpconde/claude-android-skill, 337 stars), Claude Android Ninja (Drjacky/claude-android-ninja, 124 stars), Composewebview Development (parkwoocheol/compose-webview, 103 stars) and Update Libs Versions (CosminMihuMDC/KtorMonitor, 254 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flake Triage?

yschimke (a GitHub user) maintains it in yschimke/compose-ai-tools, which has 117 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 9, 2026.

Source: yschimke/compose-ai-tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.