Agent skill

Visual Verdict

by vibeeval in vibeeval/vibecosystem

Screenshot comparison QA for frontend development. An agent skill from vibeeval/vibecosystem.

MITAuto-check passedFrontend & Design

Install Visual Verdict

skills CLI
$ npx skills add vibeeval/vibecosystem --skill visual-verdict -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vibeeval/vibecosystem visual-verdict --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/visual-verdict .claude/skills/visual-verdict && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
visual-verdict
GitHub stars
531
Token cost
~2.4k tokens
SKILL.md length
635 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Screenshot comparison QA for frontend development. An agent skill from vibeeval/vibecosystem.

  • Implementing UI from a design reference
  • SKILL.md covers When to Activate, How It Works, Prerequisites and Scoring Dimensions, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Verifying visual correctness

What it does

Visual Verdict is an agent skill from vibeeval/vibecosystem. Screenshot comparison QA for frontend development. Takes a screenshot of the current implementation, scores it across multiple visual dimensions, and returns a structured PASS/REVISE/FAIL verdict with concrete fixes. Use when implementing UI from a design reference or verifying visual correctness.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Frontend & Design, covering Visual regression testing, Frontend development and Responsive design. It works with Model Context Protocol. The repository describes itself as: AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution. The licence is MIT.

When your agent uses it

  • Implementing UI from a design reference
  • Verifying visual correctness

Example prompts

  • “/visual-verdict”

What it can do on your machine

Read from SKILL.md and the folder at commit 3b763b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown, javascript and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Visual Verdict loads about 2.4k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 635 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vibeeval/vibecosystem at commit 3b763b1, republished under its MIT licence (© vibeeval). 635 words, ~2,367 tokens.

Download SKILL.mdSave it as .claude/skills/visual-verdict/SKILL.md (or your agent's skills folder).
name
visual-verdict
description
Screenshot comparison QA for frontend development. Takes a screenshot of the current implementation, scores it across multiple visual dimensions, and returns a structured PASS/REVISE/FAIL verdict with concrete fixes. Use when implementing UI from a design reference or verifying visual correctness.

Visual Verdict — Screenshot QA Skill

Pixel-imprecise human eyes miss visual bugs. This skill provides structured, scored visual review using screenshots, returning actionable verdicts rather than vague impressions.

When to Activate

  • After a frontend-dev agent implements a UI component or page
  • When cloning an external website (pairs with clone-website skill)
  • After CSS changes to verify no visual regressions
  • When a designer provides a Figma export or reference screenshot
  • Before marking any UI task as complete

How It Works

1. Take screenshot of current implementation
2. Load reference (design mockup, Figma export, or previous version)
3. Score against 6 dimensions (0-100 each)
4. Compute weighted total score
5. Issue verdict: PASS (90+) / REVISE (60-89) / FAIL (<60)
6. List concrete mismatches with file + line fix hints
7. Loop: frontend-dev fixes → new screenshot → rescore → repeat until PASS

Prerequisites

  • browser-use MCP installed and active (~/.mcp.json)
  • Reference image available (Figma PNG export, screenshot, or URL)
  • Implementation running (local dev server or deployed URL)

Scoring Dimensions

Dimension Weights
DimensionWeightDescription
Layout accuracy0.25Positioning, spacing, alignment of all elements
Typography0.15Font family, size, weight, line-height, letter-spacing
Color accuracy0.15Background, text, border, shadow, gradient colors
Responsive behavior0.15Breakpoint transitions, mobile/tablet rendering
Interactive states0.15Hover, focus, active, disabled, loading states
Content completeness0.15All text, images, icons, labels present
Weighted Score Formula
total = (layout * 0.25) + (typography * 0.15) + (color * 0.15)
      + (responsive * 0.15) + (interactive * 0.15) + (content * 0.15)
Scoring Each Dimension (0-100)

Layout accuracy

  • All elements exactly in position, spacing matches: 90-100
  • Minor spacing deviation (< 4px): 75-89
  • Noticeable misalignment (4-16px): 50-74
  • Layout clearly wrong (columns swapped, elements missing): 0-49

Typography

  • Font, size, weight, line-height all match: 90-100
  • One attribute off (wrong weight only): 70-89
  • Two attributes off: 50-69
  • Wrong font family entirely: 0-49

Color accuracy

  • All colors within 5% of reference: 90-100
  • One element color clearly off: 65-89
  • Multiple elements wrong color: 40-64
  • Dark/light mode inverted: 0-39

Responsive behavior

  • Tested at 3 breakpoints, all correct: 90-100
  • One breakpoint breaks layout: 65-89
  • Two breakpoints break: 40-64
  • Not responsive at all: 0-39

Interactive states

  • All states implemented and visible: 90-100
  • Missing one minor state (loading): 70-89
  • Missing hover or focus: 50-69
  • No interactive states at all: 0-49

Content completeness

  • All text, images, icons present: 90-100
  • One non-critical item missing: 75-89
  • One critical item missing (CTA, heading): 50-74
  • Multiple items missing: 0-49

Verdict Thresholds

VerdictScoreMeaning
PASS90-100Acceptable for production
REVISE60-89Needs fixes, not shippable
FAIL0-59Major rework required

Verdict Output Format

markdown
## Visual Verdict: [Component/Page Name]

**Score: 85/100 — REVISE**
**Attempt: 2/5**
**Previous score: 72 (+13 this iteration)**

### Dimension Scores
| Dimension | Score | Weight | Weighted | Issues |
|-----------|-------|--------|----------|--------|
| Layout | 92 | 0.25 | 23.0 | Minor padding on mobile |
| Typography | 68 | 0.15 | 10.2 | Wrong font-weight on H2 |
| Color | 96 | 0.15 | 14.4 | - |
| Responsive | 80 | 0.15 | 12.0 | Card grid breaks at 768px |
| Interactive | 82 | 0.15 | 12.3 | Missing hover on CTA |
| Content | 91 | 0.15 | 13.7 | - |
| **TOTAL** | | | **85.6** | |

### Concrete Mismatches (prioritized by impact)

1. **[Typography / HIGH]** H2 hero heading
   Expected: `font-weight: 700`
   Actual: `font-weight: 400` (normal weight)
   Fix: `src/components/Hero.tsx` → add `font-bold` class to `<h2>`

2. **[Responsive / MEDIUM]** Card grid layout
   Expected: 3-column grid collapses at 1024px
   Actual: Collapses at 768px — causes overflow between 768-1024
   Fix: `src/components/CardGrid.tsx` → change `md:grid-cols-3` to `lg:grid-cols-3`

3. **[Interactive / MEDIUM]** Primary CTA button
   Expected: Darker background on hover (`bg-blue-700`)
   Actual: No hover state change
   Fix: `src/components/Hero.tsx` → add `hover:bg-blue-700 transition-colors` to CTA

4. **[Layout / LOW]** Mobile padding
   Expected: 16px horizontal padding on mobile
   Actual: 12px
   Fix: `src/components/Layout.tsx` → change `px-3` to `px-4`

### Next Steps
Fix issues 1-3 above, then call visual-verdict again.
Estimated score after fixes: ~94 (PASS)

Screenshot Workflow with browser-use MCP

Taking the Current Implementation Screenshot
Use browser-use MCP:
1. Navigate to implementation URL (e.g., http://localhost:3000/dashboard)
2. Wait for page fully loaded (no spinners)
3. Take full-page screenshot
4. Save to /tmp/verdict-current-{timestamp}.png
Loading the Reference

Three reference types accepted:

Figma PNG export: Designer exports component/page as PNG at 2x

Reference path: ~/Desktop/designs/dashboard-v2.png

Previous production screenshot: Regression check

Reference path: ~/.claude/benchmarks/screenshots/dashboard-prod.png

Live URL comparison: Compare two running versions

Reference URL A: http://localhost:3000/dashboard (current)
Reference URL B: https://staging.app.com/dashboard (baseline)
Show full SKILL.md (259 more words)Show less
Responsive Testing Breakpoints

Always test at these three widths unless told otherwise:

NameWidthTarget
Mobile375pxiPhone SE
Tablet768pxiPad portrait
Desktop1440pxStandard laptop

Integration with frontend-dev Agent

The typical iteration loop:

frontend-dev implements component
  → visual-verdict takes screenshot + scores
  → REVISE: sends concrete fix list to frontend-dev
  → frontend-dev applies fixes
  → visual-verdict rescores
  → Repeat until PASS (max 5 iterations)
  → PASS: task marked complete
Iteration Limit

Maximum 5 iterations per component. If PASS is not reached after 5 attempts:

ESCALATION: Visual QA failed after 5 iterations
Component: [name]
Best score achieved: [score]
Blocker: [specific issue that keeps failing]
Recommendation: [designer clarification needed / fundamentally wrong approach]

Integration with clone-website Skill

When cloning an external site, visual-verdict is the quality gate:

1. clone-website skill produces initial implementation
2. visual-verdict screenshots both original and clone
3. Scores similarity across all 6 dimensions
4. REVISE verdict triggers targeted fixes
5. PASS verdict confirms clone is production-ready

Color Comparison Method

Color is compared by extracting computed CSS values and comparing:

Reference color:      #1D4ED8  (rgb 29, 78, 216)
Actual color:         #2563EB  (rgb 37, 99, 235)
Delta E (perceived):  6.2 — noticeable, score -15

Delta E thresholds:

  • Delta E < 2: Imperceptible — full score
  • Delta E 2-5: Barely noticeable — minor deduction
  • Delta E 5-10: Noticeable — moderate deduction
  • Delta E > 10: Clearly wrong — major deduction

Typography Extraction

Typography is checked by reading computed styles, not pixels:

javascript
// Extract from running implementation
const element = document.querySelector('.hero-title')
const styles = window.getComputedStyle(element)

{
  fontFamily: styles.fontFamily,        // Compare against reference
  fontSize: styles.fontSize,            // "32px" vs expected "36px"
  fontWeight: styles.fontWeight,        // "400" vs expected "700"
  lineHeight: styles.lineHeight,        // "1.5" vs expected "1.25"
  letterSpacing: styles.letterSpacing   // "normal" vs expected "-0.02em"
}

Common Visual Bugs Caught

  • Font weight 400 instead of 700 (bold not applying)
  • Tailwind class not included in safelist (purged in build)
  • Media query breakpoint uses px when rem expected
  • Hover state not added after copy-paste from design
  • Icon imported but wrong size prop
  • Color defined in CSS variable but variable not set in theme
  • Padding/margin using px-3 when design shows px-4
  • Z-index issue hiding interactive element
  • Border-radius value wrong (rounded vs rounded-lg vs rounded-full)
  • Line-clamp not applied, text overflows on mobile

Saving Screenshots for Regression History

After a PASS verdict, save the screenshot as the new baseline:

bash
cp /tmp/verdict-current-{timestamp}.png \
   ~/.claude/benchmarks/screenshots/{component-name}-{date}.png

On the next change, this saved screenshot becomes the reference for regression comparison.


Remember: A visual that looks "pretty close" in a quick glance often has 3-5 compounding issues that add up to a broken experience. Score it. Trust the number, not the impression.

© vibeeval, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/visual-verdict of vibeeval/vibecosystem.

Open the folder on GitHubat commit 3b763b1

Compare with similar skills

Visual Verdict next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Visual Verdict compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Visual Verdict this skillvibeeval/vibecosystem531—~2.4kAutomated safety check: PassMIT
UI CraftMemTensor/memmy-agent2.1k—~2.9kAutomated safety check: PassMIT
Browser QAaffaan-m/ECC275k2 repos~1kAutomated safety check: PassMIT
Common Web Visual TestingHoangNguyen0403/agent-skills-standard571—~760Automated safety check: PassMIT
Frontend Design Routercode-yeongyu/oh-my-openagent70k—~5.3kAutomated safety check: PassCustom licence
Refero Designreferodesign/refero_skill295—~5.3kAutomated safety check: PassMIT

Similar skills

  • UI Craft

    MemTensor/memmy-agent

    Design, build, redesign, or modify browser-visible web interfaces with product-quality composition, visual systems, interaction states, responsive behavior, accessibility, and browser visual QA.

    2.1k GitHub stars~2.9k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Browser QA

    affaan-m/ECC

    Run automated post-deploy UI verification with a browser automation MCP (claude-in-chrome, Playwright, or Puppeteer): console-error and Core Web Vitals smoke checks, form and auth-flow interaction…

    275k GitHub starsUsed in 2 repos~1k tokens
    Frontend & DesignAuto-check passed
  • Common Web Visual Testing

    HoangNguyen0403/agent-skills-standard

    Standardizes visual audits, responsive design, and behavioral testing for web apps.

    571 GitHub stars~760 tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Frontend Design Router

    code-yeongyu/oh-my-openagent

    Routes frontend and UI work through design reference rulesets, with a design-system gate plus layout and print guidance, before any UI code is written.

    70k GitHub stars~5.3k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Refero Design

    referodesign/refero_skill

    Primary/default skill for UI design, product design, web design, landing pages, dashboards, product screens, redesigns, visual polish, frontend/CSS styling, design systems, components, responsive…

    295 GitHub stars~5.3k tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed
  • Better Design

    marvkr/better-design

    Build, improve, and review production interfaces with the Better Design MCP.

    254 GitHub stars~1.2k tokensUpdated 15 days ago
    Frontend & DesignAuto-check passed

More from vibeeval/vibecosystem

All 12 skills in this repo
  • Agent Benchmark

    vibeeval/vibecosystem

    Framework for measuring and tracking agent response quality over time.

    531 GitHub stars~2.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Differential Review

    vibeeval/vibecosystem

    Security-focused differential code review with blast radius analysis, risk-adaptive depth (DEEP/FOCUSED/SURGICAL), git history correlation, and structured finding format.

    531 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Factcheck Guard

    vibeeval/vibecosystem

    A skill your agent uses when making any factual claim about the codebase — existence, absence, or behavior.

    531 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Fp Check

    vibeeval/vibecosystem

    Systematic false positive verification for security findings.

    531 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • N8n Workflows

    vibeeval/vibecosystem

    n8n otomasyon workflow'lari. An agent skill from vibeeval/vibecosystem.

    531 GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Notepad System

    vibeeval/vibecosystem

    A skill your agent uses when context compression is imminent, when resuming a session, or when preserving critical decisions across long tasks.

    531 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Visual Verdict

What does Visual Verdict do?

Screenshot comparison QA for frontend development. An agent skill from vibeeval/vibecosystem. Visual Verdict is an agent skill from vibeeval/vibecosystem. Screenshot comparison QA for frontend development.

When should I use Visual Verdict?

Visual Verdict fits situations like: implementing UI from a design reference; verifying visual correctness.

How do I install Visual Verdict in Claude Code?

Run `npx skills add vibeeval/vibecosystem --skill visual-verdict -a claude-code`. Or copy the skill folder (skills/visual-verdict in vibeeval/vibecosystem) into .claude/skills/visual-verdict in your project. Claude Code loads it when a task matches its description.

How do I install Visual Verdict in Codex?

Run `npx skills add vibeeval/vibecosystem --skill visual-verdict -a codex`. Or copy the skill folder (skills/visual-verdict in vibeeval/vibecosystem) into .agents/skills/visual-verdict in your project. Codex loads it when a task matches its description.

Can I use Visual Verdict in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vibeeval/vibecosystem --skill visual-verdict -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/visual-verdict, .gemini/skills/visual-verdict, .github/skills/visual-verdict and .opencode/skills/visual-verdict in your project.

What does Visual Verdict need to run?

SKILL.md names no scripts, command-line tools or credentials: Visual Verdict is instructions for the agent only.

Does Visual Verdict access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Visual Verdict safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Visual Verdict use?

Visual Verdict is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Visual Verdict use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Visual Verdict?

Skills that share tags, products or a category with Visual Verdict: UI Craft (MemTensor/memmy-agent, 2.1k stars), Browser QA (affaan-m/ECC, 275k stars), Common Web Visual Testing (HoangNguyen0403/agent-skills-standard, 571 stars) and Frontend Design Router (code-yeongyu/oh-my-openagent, 70k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Visual Verdict?

vibeeval (a GitHub user) maintains it in vibeeval/vibecosystem, which has 531 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on August 8, 2026.

Source: vibeeval/vibecosystem on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.