Agent skill

Triage Visual Changes

by igrlk in igrlk/storybook-addon-test-codegen

Triage a UI Verify build from your coding agent via the UI Verify MCP — bucket a build's changed stories into real regressions vs cosmetic reflow vs rendering noise, summarize the real regressions…

MITAuto-check passedTesting & QA

Install Triage Visual Changes

skills CLI
$ npx skills add igrlk/storybook-addon-test-codegen --skill triage-visual-changes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install igrlk/storybook-addon-test-codegen triage-visual-changes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/igrlk/storybook-addon-test-codegen.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triage-visual-changes .claude/skills/triage-visual-changes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
triage-visual-changes
GitHub stars
154
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,570 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Triage a UI Verify build from your coding agent via the UI Verify MCP — bucket a build's changed stories into real regressions vs cosmetic reflow vs rendering noise, summarize the real regressions…

  • Works in 4 steps: Regression — a real change that's wrong… → Intended change — content changed and… → Cosmetic reflow — real diff pixels, but… → …
  • - triage this build
  • SKILL.md covers Connect the MCP (once), Recipe 1 — "Triage this build", Recipe 2 — "Summarize the… and Recipe 3 — "Accept all the new…, plus 3 more sections
  • Calls claude and curl; reaches uiverify.ai; needs UIVERIFY_API_KEY

What it does

Triage Visual Changes is an agent skill from igrlk/storybook-addon-test-codegen. Triage a UI Verify build from your coding agent via the UI Verify MCP — bucket a build's changed stories into real regressions vs cosmetic reflow vs rendering noise, summarize the real regressions for a PR comment, and accept baselines in bulk. Use after a UI Verify build reports changes and you want the agent to review, summarize, or accept from the terminal instead of clicking through the dashboard. Triggers - "triage this build", "summarize the visual regressions", "accept all the new baselines".

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Visual regression testing and MCP servers. It works with Model Context Protocol. The repository describes itself as: Addon for Storybook that generates test code for your stories. The licence is MIT.

When your agent uses it

  • - triage this build
  • Summarize the visual regressions
  • Accept all the new baselines

Example prompts

  • “triage this build”
  • “summarize the visual regressions”
  • “accept all the new baselines”
  • “/triage-visual-changes”

Requirements

  • A credential in UIVERIFY_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Regression — a real change that's wrong — content changed and the after is broken or worse
  2. Intended change — content changed and the after is correct and matches the PR's stated intent
  3. Cosmetic reflow — real diff pixels, but the diff image shows the same content shifted a few px
  4. Rendering noise — safe to accept — sub-0.4% jitter on chart edges / anti-aliased lines. Pure

What it can do on your machine

Read from SKILL.md and the folder at commit cdfbf4e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • uiverify.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • UIVERIFY_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triage Visual Changes loads about 2.9k tokens when it runs. Until then it costs about 132 tokens; SKILL.md has 1,570 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from igrlk/storybook-addon-test-codegen at commit cdfbf4e, republished under its MIT licence (© igrlk). 1,570 words, ~2,916 tokens.

Download SKILL.mdSave it as .claude/skills/triage-visual-changes/SKILL.md (or your agent's skills folder).
name
triage-visual-changes
description
Triage a UI Verify build from your coding agent via the UI Verify MCP — bucket a build's changed stories into real regressions vs cosmetic reflow vs rendering noise, summarize the real regressions for a PR comment, and accept baselines in bulk. Use after a UI Verify build reports changes and you want the agent to review, summarize, or accept from the terminal instead of clicking through the dashboard. Triggers - "triage this build", "summarize the visual regressions", "accept all the new baselines".

Triage a UI Verify build from your agent

When a build comes back changed, the dashboard shows a triptych per story (baseline / candidate / diff). This skill does that review from your agent over the UI Verify MCP — read what changed, look at the pixels when a number is ambiguous, then summarize or accept.

Connect the MCP (once)

The server is a remote Streamable-HTTP MCP at https://uiverify.ai/api/mcp, authed with your project's uv_proj_… API key as a Bearer token. Add it as a real MCP server in your agent — don't hand-roll curl against the endpoint. Over raw curl the image tools come back as MCP image blocks a shell can't render, so the pixels are useless to a script; a native client shows them to the model directly.

Claude Code:

sh
claude mcp add --transport http uiverify https://uiverify.ai/api/mcp \
  --header "Authorization: Bearer $UIVERIFY_API_KEY"

Cursor — add to .cursor/mcp.json (project) or ~/.cursor/mcp.json (global):

json
{
  "mcpServers": {
    "uiverify": {
      "url": "https://uiverify.ai/api/mcp",
      "headers": { "Authorization": "Bearer YOUR_uv_proj_KEY" }
    }
  }
}

Any other MCP client (Codex, Copilot, Gemini): point it at the same Streamable-HTTP URL with the Authorization: Bearer header. If a client can only reach it over HTTP by hand, prefer the get_diff tool (it returns downloadable image URLs) over render_diff_image (inline pixels) — see the tool table below.

Tools it exposes:

ToolDoesKey args
list_buildsRecent builds + gate status, to find the one to inspectbranch?, status?, limit?
get_buildLean triage of one build — the gate + counts {total, changed, failed, unchanged}, the AI tally, and the first page (25) of changed/failed stories (each with a diffResultId, diff %, and when AI review is on the judge's aiVerdict, aiConfidence, and for a regression aiFlagReason = its one-line "what looks unintended"), the page carrying a nextCursor. It does not dump the unchanged list — counts.unchanged is the signpost; page the rest with list_build_storiesone of commitSha / prNumber / buildId
list_build_storiesPage through one build's stories by status — the only way to browse the full changed / failed / unchanged set beyond get_build's first page. counts.unchanged from get_build is the exact count it pagesselector, status: changed | unchanged | failed, cursor?, limit? (default 25, max 100)
get_diffPer-story diff metrics + presigned image URLs (baseline / candidate / diff) you can download to a file or link straight into a PR; when AI review is on, the judge's full call: aiVerdict, aiConfidence, aiSummary (what changed), aiReasoning (why), aiFlagReason. Pass storyId to pull one story even if it didn't change — you get its current baseline URL, so you can show "identical to baseline" on a passed buildsame selector, optional storyId
render_diff_imageThe actual pixels of one image, inline for the vision model to look at (not a URL): baseline / candidate / diff, or before_after for a before-and-after crop zoomed to the changed region — one crop per region, stacked, when the story moved in several far-apart places (a header and a footer)diffResultId (or storyId for a story that didn't change), which
review_diffRecord a decision on one storydiffResultId, decision: accept | deny | ignore
accept_buildAccept every changed story in a build at oncesame selector

Recipe 1 — "Triage this build"

Bucket the changes so the human sees signal, not 60 rows. Call get_build for the PR — it carries the AI judge's call per story (aiVerdict, and for a regression the aiFlagReason) for the first page of changes; if counts.changed is larger than that page, keep paging list_build_stories { status: "changed" } on the nextCursor until it runs out, so no regression slips past the page boundary. Read the judge before you form your own view, and for anything that actually changed content, adjudicate it against the before image, not in isolation:

  • Render BOTH images, not just the diff. render_diff_image which: "baseline" and which: "candidate" and compare them. The green diff mask tells you where pixels moved, never whether the result is good — a code block that went dark-on-dark, text that lost contrast, a control that's now clipped all look like "some pixels changed" in the mask and only reveal themselves as regressions in the before/after. Judging the candidate alone is how "it has colours now → improvement" passes an illegible block. Pull get_diff for the judge's aiReasoning / aiFlagReason on the ones it flagged.

Report in four buckets, with real story names and % deltas:

  1. Regression — a real change that's wrong — content changed and the after is broken or worse: illegible/low-contrast text, clipping, overlap, missing content, a control that stopped reading as itself. A regression the current PR caused is still a regression — "my diff explains it" says nothing about whether it's correct. Never accept these; they're the reason to triage.
  2. Intended change — content changed and the after is correct and matches the PR's stated intent (a deliberate restyle, new copy, a genuinely new story). These still want a human eye, but they're accept candidates.
  3. Cosmetic reflow — real diff pixels, but the diff image shows the same content shifted a few px (a row moved down ~5px, text ghosted lower). Not a content change.
  4. Rendering noise — safe to accept — sub-0.4% jitter on chart edges / anti-aliased lines. Pure render jitter.

The judge's regression verdict is a reason to look harder, never to dismiss. If you're about to call a story the judge flagged "intended anyway", you need a concrete reason from the before/after that refutes its aiFlagReason — not "it's a colour change, probably fine." When you can't refute it, trust it.

Example output shape:

67 changed stories → buckets. ⚠️ Regressions (2) — real changes that look wrong, do NOT accept: SetupWizard/Playwright (8.3%, code block now dark-on-dark / illegible — judge flagged regression), Card/Compact (5.1%, title clipped) … 1. Intended changes (11) — content changed, looks correct, accept candidates: Hero/Default (14%, new headline copy), … 2. Cosmetic reflow (9) — real pixels, vertical shift only: EventStatusCard/All Variants (4.1%, rows shifted ~5px), TestingInit/Playground (3.7%), … 3. Chart anti-aliasing noise (6) — sub-0.4% edge jitter, safe to accept: FunnelChart/Smaller Screen (3.2%), UpliftChart (single 1px mark), … New stories (12) — first baseline, nothing to compare: Button/AllVariants, …

Never accept from inside the triage step — surface the buckets and let the human decide, or do it explicitly in recipe 3.

Show full SKILL.md (596 more words)Show less

Recipe 2 — "Summarize the visual regressions"

Same read, but output the changes that need a human — the ⚠️ regressions first, then the intended content changes — in a PR-comment shape, one line each, no reflow/noise. This is the comment to drop on the PR:

🔎 Visual review — 1 regression, 2 intended (of 67 flagged; 64 reflow/noise, listed below the fold)

  • ⚠️ SetupWizard/Playwright — code block dark-on-dark, illegible (8.3%, judge: regression)
  • Pricing/Card — CTA button color changed blue→green (12.4%)
  • Nav/Header — logo 8px larger (3.1%)

Recipe 3 — "Accept all the new baselines"

When the changes are all intended (a deliberate restyle) or all noise and you've confirmed no regression is hiding in the list (bucket ⚠️ empty), advance every changed story's baseline in one call:

accept_build { prNumber: 123 }

Don't let "the headline change is intended" halo onto the whole build — a global CSS or theme change sweeps into stories the PR never meant to touch, and "the PR caused it" is not "the PR intended it." Each changed story earns its bucket on its own before-and-after.

For a targeted accept (some real, some not), loop review_diff per diffResultId with accept / ignore / deny instead — accept the intended ones, ignore the noise, leave real regressions for a human.

Recipe 4 — "Put the before/after in the PR"

Keep the human in the PR instead of sending them to the dashboard. After you've triaged, post the changed stories inline:

  • The fastest single image: render_diff_image which: "before_after" gives you one PNG, before and after side by side, cropped to the changed region (one crop per region, stacked, if the story moved in several far-apart places) — the least to eyeball. Save it and attach it.
  • Downloadable files / durable links: get_diff returns presigned baselineUrl / candidateUrl / diffUrl. Download them (curl -L "$url" -o before.png) to attach to the PR, or drop the build link and the URLs into a comment. Always include the build link (https://uiverify.ai/builds/<id>) so the reviewer can open the full triptych.

A good PR comment: the one-line-per-regression summary from Recipe 2, the before/after image for the regressions, and the build link. Post it, don't make them go looking.

Passed build? You can still show the pixels

A build with no changes shows no triptych — but the components still rendered, and "nothing changed" is only trustworthy if you can see it. Enumerate the unchanged set with list_build_stories { status: "unchanged" } (page it on the returned nextCursor), then pull any story's current image by id: get_diff { buildId, storyId } for its baseline URL, or render_diff_image { storyId, which: "baseline" } for the pixels. Use it to confirm a component looks right on a green build, or to hand the reviewer "here it is, identical to baseline" without opening the dashboard.

Guardrails

  • Read before you write. Always get_build (and render_diff_image for anything ambiguous) before review_diff / accept_build. Don't accept a build you haven't looked at.
  • Compare before-and-after, never the after alone. For any story that changed content, render which: "baseline" AND which: "candidate" and look at both — the diff mask shows where, the pair shows whether it got worse (contrast/legibility, clipping, layout). "It has colours / it moved" is not "it's better."
  • The AI judge's regression is a signal to trust by default. Overriding it needs a concrete reason from the before/after that refutes its aiFlagReason, not a hand-wave. Attributable-to-this-PR is not a reason to dismiss it.
  • Bulk-accept is a baseline change. accept_build makes the candidate the new truth for every changed story — only reach for it when you've confirmed there's no real regression hiding in the list.
  • Distinguish "new story" from "changed story." A first-baseline story has nothing to diff; it's not a regression, just needs a baseline. Bucket it separately.

© igrlk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/triage-visual-changes of igrlk/storybook-addon-test-codegen.

Open the folder on GitHubat commit cdfbf4e

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in igrlk/storybook-addon-test-codegen, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Triage Visual Changes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triage Visual Changes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triage Visual Changes this skilligrlk/storybook-addon-test-codegen1541 repos~2.9kAutomated safety check: PassMIT
Glance TestDebugBase/glance156—~827Automated safety check: PassMIT
Qwen Code E2E TestingQwenLM/qwen-code28k—~2.1kAutomated safety check: PassApache-2.0
Adding LLM MCP ToolsTriliumNext/Trilium38k—~2.5kAutomated safety check: PassAGPL-3.0
Agentacct Workflowmikehasa/agentacct765—~1.6kAutomated safety check: PassMIT
Chatgpt App Submissionnteract/semiotic2.7k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • Glance Test

    DebugBase/glance

    Run E2E browser tests on any web application using Glance MCP.

    156 GitHub stars~827 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Adding LLM MCP Tools

    TriliumNext/Trilium

    A skill your agent uses when adding, changing, or reviewing an LLM/MCP tool in Trilium (the defineTools definitions under packages/trilium-core/src/services/llm/tools/ —…

    38k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check passed
  • Agentacct Workflow

    mikehasa/agentacct

    A skill your agent uses when working in a repo with agentacct MCP configured, or when asked to track coding-agent work, smoke-test agentacct integrations, or report objective AI-agent task evidence.

    765 GitHub stars~1.6k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Chatgpt App Submission

    nteract/semiotic

    Inspect a ChatGPT Apps MCP server codebase and generate chatgpt-app-submission.json with app info suggestions, tool hint justifications, test cases, and negative test cases, then report review-check…

    2.7k GitHub stars~2.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Ue Test Authoring

    JasonMa0012/MooaToon

    A skill your agent uses when writing or modifying UE automated tests (Automation, CQTest, Functional, Gauntlet, LowLevel) with Rider MCP available.

    749 GitHub stars~2.1k tokensUpdated 19 days ago
    Testing & QAAuto-check: notes

More from igrlk/storybook-addon-test-codegen

  • Check Visual Changes

    igrlk/storybook-addon-test-codegen

    Check the visual impact of an edit mid-task with uiverify check — render just the components you're touching on the UI Verify fleet and diff them against the real CI baseline, without opening a PR…

    154 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Economical Visual Tests

    igrlk/storybook-addon-test-codegen

    Author economical visual tests — full visual coverage in the fewest billable snapshots.

    154 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Storybook Visual Testing

    igrlk/storybook-addon-test-codegen

    Make Storybook stories deterministic for UI Verify so captures stop coming back "changed" without a real change (flaky diffs).

    154 GitHub stars~1.9k tokensUpdated 24 days ago
    Auto-check passed
  • Making UI Changes

    igrlk/storybook-addon-test-codegen

    Read before changing any component, page, or styles. An agent skill from igrlk/storybook-addon-test-codegen.

    154 GitHub starsUsed in 1 repo~493 tokens
    Auto-check passed

Categories

Questions about Triage Visual Changes

What does Triage Visual Changes do?

Triage a UI Verify build from your coding agent via the UI Verify MCP — bucket a build's changed stories into real regressions vs cosmetic reflow vs rendering noise, summarize the real regressions…. Triage Visual Changes is an agent skill from igrlk/storybook-addon-test-codegen. Triage a UI Verify build from your coding agent via the UI Verify MCP — bucket a build's changed stories into real regressions vs cosmetic reflow vs rendering noise, summarize the real regressions for a PR comment, and accept baselines in bulk.

When should I use Triage Visual Changes?

Triage Visual Changes fits situations like: - triage this build; summarize the visual regressions; accept all the new baselines.

How do I install Triage Visual Changes in Claude Code?

Run `npx skills add igrlk/storybook-addon-test-codegen --skill triage-visual-changes -a claude-code`. Or copy the skill folder (.agents/skills/triage-visual-changes in igrlk/storybook-addon-test-codegen) into .claude/skills/triage-visual-changes in your project. Claude Code loads it when a task matches its description.

How do I install Triage Visual Changes in Codex?

Run `npx skills add igrlk/storybook-addon-test-codegen --skill triage-visual-changes -a codex`. Or copy the skill folder (.agents/skills/triage-visual-changes in igrlk/storybook-addon-test-codegen) into .agents/skills/triage-visual-changes in your project. Codex loads it when a task matches its description.

Can I use Triage Visual Changes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add igrlk/storybook-addon-test-codegen --skill triage-visual-changes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triage-visual-changes, .gemini/skills/triage-visual-changes, .github/skills/triage-visual-changes and .opencode/skills/triage-visual-changes in your project.

What does Triage Visual Changes need to run?

Going by SKILL.md and its folder, Triage Visual Changes needs the command-line tools its instructions call (claude and curl) and credentials named UIVERIFY_API_KEY. Our summary lists: A credential in UIVERIFY_API_KEY.

Does Triage Visual Changes access the network?

SKILL.md names 1 domain. In commands or code: uiverify.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Triage Visual Changes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triage Visual Changes use?

Triage Visual Changes is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triage Visual Changes use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triage Visual Changes?

Skills that share tags, products or a category with Triage Visual Changes: Glance Test (DebugBase/glance, 156 stars), Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars), Adding LLM MCP Tools (TriliumNext/Trilium, 38k stars) and Agentacct Workflow (mikehasa/agentacct, 765 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triage Visual Changes?

igrlk (a GitHub user) maintains it in igrlk/storybook-addon-test-codegen, which has 154 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 13, 2026.

Source: igrlk/storybook-addon-test-codegen on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.