Agent skill

UI Verification

by mblode in mblode/agent-skills

Runs scoped browser probes for focus, hit targets, overflow, themes, request failures, and performance attribution, with evidence linked to UI rule IDs.

MITAuto-check passedMobile

Install UI Verification

skills CLI
$ npx skills add mblode/agent-skills --skill ui-verification -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mblode/agent-skills ui-verification --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mblode/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ui-verification .claude/skills/ui-verification && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ui-verification
GitHub stars
143
Token cost
~3.5k tokens
SKILL.md length
1,844 words
Files
14 (incl. references)
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Runs scoped browser probes for focus, hit targets, overflow, themes, request failures, and performance attribution, with evidence linked to UI rule IDs.

  • Works in 6 steps: Establish the session → Select the probe set → Run the probes → …
  • Asked to verify this in the browser
  • SKILL.md covers Contents, When to run, Probe catalogue and Progress checklist, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

UI Verification is an agent skill from mblode/agent-skills. Runs scoped browser probes for focus, hit targets, overflow, themes, request failures, and performance attribution, with evidence linked to UI rule IDs. Use when asked to "verify this in the browser", "reproduce this UI finding", or "re-measure after the UI fix".

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including reference files (for example `evals/evals.json`, `probes/axe-scan.md` and `probes/console-network.md`). Compatibility notes: Requires access to the target app and browser automation. Bundled JavaScript recipes use the Playwright page API.

It sits in Mobile, covering Mobile testing and debugging. The repository describes itself as: Nobody ships AI slop on purpose. These skills make sure you don’t. The licence is MIT.

When your agent uses it

  • Asked to verify this in the browser
  • Reproduce this UI finding
  • Re-measure after the UI fix

Example prompts

  • “verify this in the browser”
  • “reproduce this UI finding”
  • “re-measure after the UI fix”
  • “/ui-verification”

Requirements

  • Compatibility (from SKILL.md): Requires access to the target app and browser automation. Bundled JavaScript recipes use the Playwright page API.

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Establish the session
  2. Select the probe set
  3. Run the probes
  4. Decide each finding
  5. Clearing re-run
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit cef4cfa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires access to the target app and browser automation. Bundled JavaScript recipes use the Playwright page API.

    From compatibility in the SKILL.md frontmatter.

Context cost

UI Verification loads about 3.5k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 1,844 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mblode/agent-skills at commit cef4cfa, republished under its MIT licence (© mblode). 1,844 words, ~3,528 tokens.

Download SKILL.mdSave it as .claude/skills/ui-verification/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
ui-verification
description
Runs scoped browser probes for focus, hit targets, overflow, themes, request failures, and performance attribution, with evidence linked to UI rule IDs. Use when asked to "verify this in the browser", "reproduce this UI finding", or "re-measure after the UI fix".
compatibility
Requires access to the target app and browser automation. Bundled JavaScript recipes use the Playwright page API.

UI Verification

Owns the browser session. Every other UI skill in this repo reasons about source and infers what the user will see; this one loads the page and measures it.

  • IS: booting the app, driving it with a browser, and running the probes that decide a rule at runtime: computed boxes, injected failures, observed layout shift, a scripted Tab walk, an axe scan per theme. Output is findings keyed to rule ids with reproducible evidence, plus the clearing re-run after a fix.
  • IS NOT: finding defects by reading source or deciding their tier and ship verdict (ui-design Audit mode owns both); building or restyling UI (ui-design Build); authoring a durable test suite (write Playwright tests); building or maintaining a whole app's persistent verify CLI, doctor command, and feature map (app-verification, call it by name); pixel-diff regression against a baseline (Chromatic, Percy); field performance (RUM or CrUX; Lighthouse is a lab tool).

The division of labour is the point. A static audit reports what the code will probably do; it cannot see a 40px control whose hit area a pseudo-element already expands to 44, or a retry button wired to nothing. This skill reproduces or kills each of those, so a finding arrives with a measurement instead of a confidence.

Contents

When to run

SituationWhat this skill does
A ui-design audit emitted findings and the user wants them confirmedRun only the probes those rule ids map to, in references/rule-coverage.md
No audit ran; the user points at a route or a running appDetect features, run the full battery on the resolved routes
A fix just landed for a previously reproduced findingRun the clearing re-run only (step 5)
The user asks for captures across themes or widthsprobes/theme-locale-matrix.md alone
A design-system.md claims a scale and someone needs it checkedRead the claimed values off computed styles on a real page; a theme value the build overrides never reaches the browser

Nothing here assigns a tier or a ship verdict. Hand reproduced findings back with their evidence and let ui-design's references/ship-readiness.md tier them; two skills tiering the same finding is how the tiers drift apart.

Probe catalogue

Each file is one probe: what it measures, the driver calls, the false positives it must guard, and the shape of the evidence it returns.

ProbeMeasuresPrimary for
probes/axe-scan.mdaxe-core violations per route and theme, including computed contrastcontrast, accessible names, landmarks
probes/target-size.mdBounding box and effective hit area of every visible interactive elementinteraction-target-size
probes/focus-walk.mdScripted Tab traversal, focus-ring pixel delta, dialog trap and restorationfocus-*, interaction-focus-visible, interaction-keyboard-operable
probes/layout-shift.mdAttributed layout-shift entries with the data response held openstates-layout-shift, perf-image-dimensions-and-priority
probes/viewport-stress.mdHorizontal overflow and clipped text at 320px and up, and under tripled stringslayout-long-content-safety, mobile-viewport-scaling
probes/failure-injection.mdWhat renders when the data request returns 500, [], malformed, or nothingstates-no-error-state, states-no-empty-state, async-*, microcopy-*
probes/theme-locale-matrix.mdThe same route captured across viewport, theme, pseudo-locale, and directiondark-i18n-untested, dark-i18n-rtl-untested
probes/console-network.mdConsole errors, page errors, failed requests, hydration mismatch warningshydration mismatch, silent runtime failures
probes/web-vitals.mdLCP with its attributed element, CLS, INP on a scripted interactionperf attribution, never a budget verdict

Progress checklist

text
Verification progress:
- [ ] Step 1: Establish the session (references/session-setup.md): driver, build mode, base URL, auth, resolved routes
- [ ] Step 2: Select the probe set from handed-over rule ids (references/rule-coverage.md) or from the routes
- [ ] Step 3: Run each probe; record evidence artifacts before interpreting any of them
- [ ] Step 4: Decide each finding reproduced / not-reproduced / unknown; repeat timing-sensitive or inconsistent results when needed
- [ ] Step 5: For each fix applied, re-run the identical probe and record clearedBy
- [ ] Step 6: Emit the verification block per finding (references/evidence-output.md), then render
- [ ] Step 7: List probes skipped and why. A probe that could not run is never a pass

1. Establish the session

Read references/session-setup.md. It resolves four things, and getting any of them wrong invalidates every probe downstream:

  • Driver. Playwright when the repo can run it, the Chrome DevTools browser tools when it cannot. Four probes need request interception and network control, so they do not exist under a driver without them.
  • Build mode. Perf and layout-shift probes run against a production build; dev-mode numbers measure the bundler. Failure injection and focus probes run fine in dev.
  • Auth and seeded data. A route behind a login with no seeded session returns unknown, never a pass.
  • Routes. Map the diff to URLs. A component with no reachable route falls back to the repo's Storybook, and to unknown if there is none.

Stop here if the app will not boot. A verification run with no session produces no findings, which is a reportable outcome and not a clean bill of health.

2. Select the probe set

Handed a list of rule ids (the normal case, from a ui-design audit): open references/rule-coverage.md, take the probe each id maps to, and run only those. Ids with no probe stay source-only findings and pass through untouched, marked as such.

Given only routes: detect features the way ui-design's feature playbooks do (form, list, modal, dashboard, checkout), then run the battery that surface earns. probes/axe-scan.md, probes/console-network.md, and probes/viewport-stress.md run on every route regardless: they are cheap, and they are the three that find things nobody suspected.

Budget the matrix before running it. Routes multiplied by viewports multiplied by themes grows fast, and a run that takes twenty minutes gets skipped next time. Two viewports (360 and 1280) and two themes cover the ground; add widths only where a probe already found an edge.

3. Run the probes

Each probe file carries its own recipe. Three rules hold across all of them:

  • Capture evidence before interpreting it. Write the screenshot, the JSON measurement, and the console log to disk first. A finding whose evidence was never written is unverifiable by the person reading the report, which puts it back where the static audit left it.
  • One probe, one route, one viewport, one theme. Never fold two conditions into one run: when the result surprises you, you need to know which axis produced it.
  • Repeat uncertain measurements. Retry timing-sensitive or inconsistent results under controlled conditions. A deterministic failure with a captured trigger needs no duplicate run. A result that flips remains inconclusive until its conditions are understood.

4. Decide each finding

Three outcomes, and the middle one is the one that earns this skill its keep.

OutcomeMeaningWhat it does to the handed-over finding
reproducedThe probe measured the defectStays a fail, now carrying observed from the measurement rather than from the source read
not-reproducedThe tested conditions did not exhibit the defectWithdraw only if the probe exercised the alleged trigger; otherwise retain the candidate with the remaining coverage gap
unknownThe probe could not run or could not decideFinding survives as unknown with the probe's reason. Never converts to a pass

A not-reproduced result is a real deliverable, not a wasted run. It is what removes the false positives a large rule corpus asserts with file:line confidence, and it feeds the rejection section ui-design already requires.

Where a probe finds something no static rule predicted, emit it as a new finding against the rule id the probe is primary for. Where no rule covers it (an axe violation with no ui-design counterpart, a console error), emit it under axe:<violation-id> or runtime:<signature> and say plainly that it came from the browser and not the corpus.

Show full SKILL.md (688 more words)Show less

5. Clearing re-run

A fix is not verified by reading the diff. Re-run the identical probe: same route, same viewport, same theme, same seed, same injected failure. Record it as clearedBy on the finding, keep the before-artifact, and never overwrite it with the after.

Three outcomes worth naming:

  • The probe now passes: the finding is applied and cleared.
  • The probe still fails: the fix did not work. Report it as applied-unverified and say so; do not report the fix and stay silent about the re-run.
  • The probe passes but a different probe on the same route now fails: the fix caused a regression, which is a new finding, not a footnote on the old one.

6. Report

references/evidence-output.md owns the shape: a verification block appended to each finding, a top-level session block, and artifact paths. It defines only the delta on ui-design's references/output-adapters.md schema, which stays the single owner of the finding object, the three counts, and the verdict.

Where this skill runs standalone, render the same terminal adapter with SHIP VERDICT omitted: the verdict is a property of a tiered audit, and printing one from probe results alone invents a tier assignment nobody made.

Honesty rules

  • A skipped probe is not a passed probe. Report every probe that did not run, with the reason (no route, auth required, driver lacks interception, app would not build). A report that silently omits them reads as coverage it does not have.
  • A screenshot is evidence, not a verdict. Probes decide on measurements. Captures exist so a human can check the measurement, and for the two questions no measurement settles: whether the dark theme looks right, and whether the pseudo-locale broke the layout or merely the prose.
  • Do not widen into a redesign. This skill reports what it measured. A finding that needs a new type scale names the mode to run next, exactly as an audit does.
  • The app is the subject, not the harness. A failure caused by the probe itself (a selector that never resolved, a route interception that swallowed the wrong request) is a harness bug. Fix the probe and re-run; never report it as a defect in the app.
  • Numbers carry their units and their conditions. 44x44px at 360px width with touch emulation on is a measurement. too small is the inference this skill exists to replace.

Gotchas

The ones that cut across probes. Each probe file carries its own false positives.

  • page.emulateMedia({ colorScheme: 'dark' }) does nothing for an app that themes with a class="dark" or data-theme attribute, which is most Tailwind apps. The probe reports a clean dark pass while never having left light mode. Drive the app's own toggle, and assert the attribute landed before capturing.
  • Running perf or layout-shift probes against next dev measures on-demand compilation. The first navigation to a route can spend seconds in the bundler, which lands in LCP and dwarfs anything real.
  • Animations that never settle keep a screenshot probe waiting until timeout. Set prefers-reduced-motion: reduce for captures, then run motion-sensitive checks in a separate pass with it off, because reduced motion is also a code path that can be broken.
  • ui-design: reads source, produces the findings this skill reproduces, and owns tiering, the ship verdict, and the finding schema.
  • typography-audit: type findings that need a rendered measure or leading value can be handed here for the measurement.
  • ax-audit: agentic surfaces. Its runtime questions use the same session and probes.
  • ui-animation: motion craft. This skill can capture the timing, but judging the curve is that skill's.
  • app-verification: a durable, per-app harness (a verify CLI, a doctor command, a feature map, worktree isolation) that a repo keeps and every agent reuses; this skill's probes are the kind of check that harness's verify can call for the UI paths it covers. Reach for this skill for a one-off browser probe against a fixed UI rule; reach for app-verification (call it by name) to give the repo a lasting way to run and prove itself session after session.

Maintenance only: evals/evals.json holds the behavioural scenarios and routing prompts for anyone changing this skill. It never loads during a verification run.

© mblode, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (references) in skills/ui-verification of mblode/agent-skills.

  • SKILL.md
  • evals/evals.json
  • probes/axe-scan.md
  • probes/console-network.md
  • probes/failure-injection.md
  • probes/focus-walk.md
  • probes/layout-shift.md
  • probes/target-size.md
  • probes/theme-locale-matrix.md
  • probes/viewport-stress.md
  • probes/web-vitals.md
  • references/evidence-output.md
  • references/rule-coverage.md
  • references/session-setup.md

Open the folder on GitHubat commit cef4cfa

Compare with similar skills

UI Verification next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

UI Verification compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
UI Verification this skillmblode/agent-skills143—~3.5kAutomated safety check: PassMIT
Fieldworks Uia2 Parity Testingsillsdev/FieldWorks110—~1kAutomated safety check: PassCustom licence
Cm UI Engineerkingxiaozhe/cm-workflow104—~1.1kAutomated safety check: PassMIT
Phone HarnessShawnPana/phone-harness3.2k—~4.3kAutomated safety check: PassMIT
Maa Issue Log AnalysisMaaAssistantArknights/MaaAssistantArknights24k—~4kAutomated safety check: PassAGPL-3.0
Mobile QAtloncorp/tlon-apps107—~2.4kAutomated safety check: PassMIT

Similar skills

  • Design or review FieldWorks UI automation and accessibility tests: UIA2, FlaUI, Appium, WinAppDriver, Avalonia.Headless, keyboard, focus, IME, and automation-id strategy.

    110 GitHub stars~1k tokensUpdated today
    MobileAuto-check passed
  • Cm UI Engineer

    kingxiaozhe/cm-workflow

    UI 还原工程师 Skill,把已确认的设计基准像素级还原为生产代码(token 先行、原子顺序、按交付形态量化验收:Web 用 BackstopJS、App 用 Maestro+模拟器截图);有基准才出场,不做业务逻辑

    104 GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed
  • Phone Harness

    ShawnPana/phone-harness

    Control the user's phone — an iPhone through the Mac's iPhone Mirroring window, an Android over adb, or a rented cloud Android: open apps, tap, type, swipe, read the screen.

    3.2k GitHub stars~4.3k tokensUpdated 10 days ago
    MobileAuto-check passed
  • Maa Issue Log Analysis

    MaaAssistantArknights/MaaAssistantArknights

    分析 MaaAssistantArknights 上游仓库公开 Issue(https://github.com/MaaAssistantArknights/MaaAssistantArknights/issues/...

    24k GitHub stars~4k tokensUpdated today
    MobileAuto-check passed
  • Mobile QA

    tloncorp/tlon-apps

    Run a mobile QA checklist on a physical Android device over adb for tlon-apps, then triage what fails into fixes.

    107 GitHub stars~2.4k tokensUpdated today
    MobileAuto-check passed
  • Store Listing Screenshots

    therxmv/Telegram-Themer

    Generate TelegramThemer's Play Store listing images — capture the 8 required app screenshots on a running emulator/device by driving the real UI with adb, then composite them into the final…

    119 GitHub stars~2.5k tokensUpdated 26 days ago
    MobileAuto-check passed

More from mblode/agent-skills

All 28 skills in this repo
  • Agent Ready

    mblode/agent-skills

    Implements agent-readiness on public sites and docs from Mintlify Agent Score, AFDocs, Is Agentic, Is It Agent Ready, or url-discovery-bench reports, or from server logs of agents 404ing on guessed…

    143 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • Agent Skills Creator

    mblode/agent-skills

    Creates and improves portable Agent Skills with a validator, routing scenarios, and evidence-based keep, cut, merge, or retire decisions.

    143 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • Chat History

    mblode/agent-skills

    Recovers decisions, previous fixes, research, and what followed a prompt from past AI conversations, with source evidence.

    143 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • CI Speedup

    mblode/agent-skills

    Cuts the wait from push to green by measuring a pipeline's critical path from run timestamps, then splitting, sharding, trimming setup and sharing test module state, with a before/after ledger.

    143 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • PR Babysitter

    mblode/agent-skills

    Monitors or repairs an open GitHub PR: CI failures, conflicts, review threads, and merge readiness, reporting state changes.

    143 GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • App Verification

    mblode/agent-skills

    Builds and maintains a repo's own verification harness (verify CLI, doctor, worktree isolation, feature map, seed data) and a reproduce-first bug handoff.

    143 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed

Questions about UI Verification

What does UI Verification do?

Runs scoped browser probes for focus, hit targets, overflow, themes, request failures, and performance attribution, with evidence linked to UI rule IDs. UI Verification is an agent skill from mblode/agent-skills. Runs scoped browser probes for focus, hit targets, overflow, themes, request failures, and performance attribution, with evidence linked to UI rule IDs.

When should I use UI Verification?

UI Verification fits situations like: asked to verify this in the browser; reproduce this UI finding; re-measure after the UI fix.

How do I install UI Verification in Claude Code?

Run `npx skills add mblode/agent-skills --skill ui-verification -a claude-code`. Or copy the skill folder (skills/ui-verification in mblode/agent-skills) into .claude/skills/ui-verification in your project. Claude Code loads it when a task matches its description.

How do I install UI Verification in Codex?

Run `npx skills add mblode/agent-skills --skill ui-verification -a codex`. Or copy the skill folder (skills/ui-verification in mblode/agent-skills) into .agents/skills/ui-verification in your project. Codex loads it when a task matches its description.

Can I use UI Verification in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mblode/agent-skills --skill ui-verification -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ui-verification, .gemini/skills/ui-verification, .github/skills/ui-verification and .opencode/skills/ui-verification in your project.

What does UI Verification need to run?

SKILL.md names no scripts, command-line tools or credentials: UI Verification is instructions for the agent only. Compatibility (from SKILL.md): Requires access to the target app and browser automation. Bundled JavaScript recipes use the Playwright page API..

Does UI Verification access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is UI Verification safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does UI Verification use?

UI Verification is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does UI Verification use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.

What are the alternatives to UI Verification?

Skills that share tags, products or a category with UI Verification: Fieldworks Uia2 Parity Testing (sillsdev/FieldWorks, 110 stars), Cm UI Engineer (kingxiaozhe/cm-workflow, 104 stars), Phone Harness (ShawnPana/phone-harness, 3.2k stars) and Maa Issue Log Analysis (MaaAssistantArknights/MaaAssistantArknights, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains UI Verification?

mblode (a GitHub user) maintains it in mblode/agent-skills, which has 143 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on October 6, 2026.

Source: mblode/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.