Agent skill

Compare Screenshots

by dzhng in dzhng/skills

Compare screenshots against the intended design, distinguishing approved references from historical baselines.

MITAuto-check passed

Install Compare Screenshots

skills CLI
$ npx skills add dzhng/skills --skill compare-screenshots -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dzhng/skills compare-screenshots --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/visual/compare-screenshots .claude/skills/compare-screenshots && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compare-screenshots
GitHub stars
1k
Token cost
~2.6k tokens
SKILL.md length
1,449 words
Files
5 (incl. scripts, references)
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Compare screenshots against the intended design, distinguishing approved references from historical baselines.

  • Works in 8 steps: Establish the target. Use the… → Confirm comparability so the differences… → Measure the approved design. For… → …
  • Iterative implementation matching
  • SKILL.md covers Workflow, Useful Metrics, Single-Image Metrics and Distance Score, plus 3 more sections
  • Runs JavaScript scripts from its folder

What it does

Compare Screenshots is an agent skill from dzhng/skills. Compare screenshots against the intended design, distinguishing approved references from historical baselines. Use for iterative implementation matching, before/after visual review, objective image telemetry, or checking a lone capture for flat, empty, or badly framed content.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/reference-landmarks.md` and `references/subagent-visual-review.md`).

The repository describes itself as: Reusable AI agent skills for software factories: explore ideas, write specs, implement, review, and run autonomous research. Works with Claude Code, Codex, and other… The licence is MIT.

When your agent uses it

  • Iterative implementation matching
  • Before/after visual review
  • Objective image telemetry
  • Checking a lone capture for flat

Example prompts

  • “/compare-screenshots”

Requirements

  • Node.js

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Establish the target. Use the user-approved reference and stated design
  2. Confirm comparability so the differences you see are real, not capture
  3. Measure the approved design. For reference matching, read and complete
  4. Inspect boundaries before judging the whole. For every changed visual
  5. Judge each divergence against the target. For every place the two images
  6. Get a neutral second opinion for disputed or high-stakes calls: a fresh
  7. Resolve mismatches to an approved target. When implementing a selected
  8. Conclude with one verdict: candidate is less wrong (accept, and re-bless

What it can do on your machine

Read from SKILL.md and the folder at commit d513228. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (JavaScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compare Screenshots loads about 2.6k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 1,449 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from dzhng/skills at commit d513228, republished under its MIT licence (© dzhng). 1,449 words, ~2,631 tokens.

Download SKILL.mdSave it as .claude/skills/compare-screenshots/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
compare-screenshots
description
Compare screenshots against the intended design, distinguishing approved references from historical baselines. Use for iterative implementation matching, before/after visual review, objective image telemetry, or checking a lone capture for flat, empty, or badly framed content.

Compare Screenshots

Judge images against the intended result. A user-approved design is the target; a historical baseline is only an earlier attempt and may be wrong. Metrics locate differences, never decide correctness. Use design-with-images for the full exploration-to-implementation loop.

Workflow

  1. Establish the target. Use the user-approved reference and stated design requirements when available; do not replace them with your own taste. Otherwise derive the target from the visual requirement, what the thing depicts in reality and the domain skill that owns the look. Write it down in one or two concrete sentences ("low sun should cast long shadows east; trees fill the canopy; labels stay legible at this zoom").
    • If the right answer isn't clear — competing valid readings, a taste or product-intent call, a tradeoff only the owner can settle — stop and ask the user what the correct answer should be. Show them the comparison. Do not quietly default to the baseline to avoid asking; that bakes in whatever the baseline got wrong.
  2. Confirm comparability so the differences you see are real, not capture artifacts: same viewport, DPR, route/page, frozen time/tick, camera intent, UI state, data, fonts/assets where they matter. If not comparable, fix capture setup or compare only a crop/feature where the mismatch is harmless.
  3. Measure the approved design. For reference matching, read and complete Reference Landmarks before changing code or accepting a candidate. Then generate artifacts sized to the question: side-by-side, key-feature crops/zooms, grayscale, absolute grayscale heatmap, pixelmatch diff, per-image Sobel/edge maps, edge-difference heatmap, JSON metrics.
  4. Inspect boundaries before judging the whole. For every changed visual effect, inspect all sides at native scale and in matched detail crops. Include the effect's full fade and surrounding space; a crop ending at the component box hides spill. Compare top/right/bottom/left extents separately, anchored to visible text, rules or silhouettes rather than inferred CSS bounds. Check every foreground feature crossed by the effect (lines, icons, text, adjacent panels): brightness, color, sharpness and continuity must match the target. A readable line can still be incorrectly dimmed. Record each check as reference observation → candidate observation → pass/fix/uncertain, with its crop. Use the landmark table to resolve local distances and contrast; full-frame averages cannot settle a local defect.
  5. Judge each divergence against the target. For every place the two images differ, name what is actually there in plain terms — missing content, wrong camera, bad hierarchy, weak contrast, wrong depth, text overlap, layout shift, clipped edge, unexpected blur, style mismatch — and decide which side is closer to correct. The answer can be the candidate, the baseline, both wrong, or a genuine toss-up.
  6. Get a neutral second opinion for disputed or high-stakes calls: a fresh subagent given only the two images and neutral labels, per references/subagent-visual-review.md.
  7. Resolve mismatches to an approved target. When implementing a selected design, record material differences in spacing, shape, softness, typography and hierarchy; revise, recapture and repeat until resolved or the user changes the target. Do not silently exempt a difference because the code is simpler.
  8. Conclude with one verdict: candidate is less wrong (accept, and re-bless the baseline if one exists), baseline is less wrong (reject), both wrong (another pass needed — say what's still off), or unclear (ask the user). Accept only when every boundary/overlap check passes or has an explicit user-approved deviation; uncertainty requires closer evidence, not a pass. Overall resemblance or a positive second opinion cannot cancel a local defect. Never accept on a lower score alone or reject on a higher one. Never hide content, blur detail, crop away differences, or make the capture less truthful to move a number.

Useful Metrics

Pick metrics that answer the question. For full visual comparisons, report:

  • mae: mean absolute grayscale difference, 0..255, lower is closer.
  • rmse: grayscale root mean square error, lower is closer.
  • diffRatio16, diffRatio32, diffRatio64: fraction of pixels over each grayscale delta threshold.
  • pixelmatchRatio: mismatch ratio from pixelmatch over grayscale images.
  • edgeEnergyCurrent and edgeEnergyCandidate: average Sobel edge strength.
  • edgeEnergyRatio: candidate/current. Far below 1 usually means missing geometry, props, labels, or terrain; far above 1 usually means noisy or incorrect detail.
  • edgeDiffRatio32: fraction of pixels whose Sobel edge differs materially.
  • avgLuminanceCurrent, avgLuminanceCandidate, avgLuminanceDelta: average brightness and delta. Use when a render is visibly too dark/light even if a broader distance score improves.
  • Content proxies relevant to the scene: black/void ratio, terrain-like ratio, water-like ratio, team-color ratio, label/text mask ratio.

For UI/document/layout reviews, also use crop bounds, text/foreground mask coverage, contrast checks, edge clipping, element positions, and before/after dimensions when those beat global pixel distance.

Show full SKILL.md (699 more words)Show less

Single-Image Metrics

Every metric above measures one image against another, so none of them can answer "is this capture worth anything" when there is nothing to compare it to — and the pair score is symmetric, so an enormous distance never says which side is the empty frame. A few absolute numbers do, computed on a coarse grid from a single PNG:

  • colorEntropyBits under ~3.0, or dominantColorShare over ~0.6: one colour owns the frame. A sparse scene, an unlit one, or a subject that never drew.
  • edgeDensity under ~0.04: almost no form anywhere. Empty framing, a primitive-dominant scene, or the subject sitting outside the crop.
  • luminanceContrast under ~60: fog, darkness, or haze compressing the whole frame into one band.
  • transparentShare above 0 on a capture that should be opaque: the capture itself is wrong. A transparent pixel keeps whatever RGB it was left with, so an invisible frame can look rich until it is composited — the scene metrics composite before measuring, and name the invisible share rather than letting you infer it. The pair metrics above still read stored RGB, so this field is where transparency gets told either way.

These are thresholds for suspicion, not gates. A deliberately minimal design, a night scene, an empty-state screen, and a whiteboard all trip them honestly. Use them to decide where to look, then say what the frame is actually doing — never adjust a capture to raise a number, which is the same failure as cropping away a difference.

Distance Score

When a single fixed-pair number is useful, this default works for structural changes:

distance = 0.35 * diffRatio32 + 0.25 * pixelmatchRatio + 0.25 * edgeDiffRatio32 + 0.15 * min(1, abs(log2(edgeEnergyRatio)))

It measures distance from the other image, nothing more. Because the baseline can be wrong, a distance of 0 is not success and a large distance is not failure — a richer scene, clearer models, stronger labels, real depth, or better lighting all legitimately raise it. Use the score to find where the images move; decide who is right in step 5. Name the field for what it measures (distance, not "parity") so no one reads it as a verdict.

Report the full-frame score and, when UI dominates the shot, a labeled world-crop score. Use the world-crop score to locate renderer movement and keep the full-frame score so UI/camera mistakes stay visible.

Score Discipline

  • Quote the previous and new distance for the same pair each iteration, then say whether the movement is toward the target, away from it, or diagnostic noise.
  • Prefer edge metrics for missing-content bugs. A flat top-down map can show a deceptively moderate grayscale diff while edge energy proves trees, roads, city forms, or army silhouettes are absent.
  • Segment out stable UI when it dominates and the question is the world render; keep a full-frame score too, labeled.
  • If the camera is wrong, pixel scores are diagnostic only. Fix camera intent first, then judge the render.

Tooling

Keep comparison scripts inside the skill or a temporary workspace, not in product code, unless the product genuinely needs screenshot comparison at runtime.

  • scripts/visual-parity-diff.mjs is a reusable local helper. Run it with REFERENCE_DIR=<png-folder>, CANDIDATE_DIR=<png-folder>, and optional OUT_DIR=<artifact-folder>. REPORT_ORDER=a,b,c pins ordering; CROPS_JSON=<file> adds labeled crops (keyed by image id, each crop in pixels or { "unit": "ratio" } normalized bounds). Pair reports carry a sceneMetrics block per side.
  • Drop REFERENCE_DIR to run the same helper on a folder with no counterpart: it writes scene-metrics.json with the single-image numbers above and no diff artifacts. Use it on a lone screenshot, on a full capture set before anyone reviews it, or to find which side of a large distance is the empty one.
  • scripts/visual-parity-diff.eval.mjs is the helper's own eval: it generates fixtures whose correct answer is known by construction — empty, transparent, primitive-dominant, authored, degenerate — and asserts the classification, the invariances, and the CLI contract. Run it with the same REPO_ROOT after changing the script. Never satisfy a failing check by loosening a threshold until you have shown the fixture, not the code, is what's wrong.
  • For other tasks, adapt the same artifact set rather than adding one-off scripts to the application. Extend the helper if a needed pair is uncovered.

References

  • references/subagent-visual-review.md: neutral subagent prompt/config for an independent judgment when history could bias you.

© dzhng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/visual/compare-screenshots of dzhng/skills.

  • SKILL.md
  • references/reference-landmarks.md
  • references/subagent-visual-review.md
  • scripts/visual-parity-diff.eval.mjs
  • scripts/visual-parity-diff.mjs

Open the folder on GitHubat commit d513228

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in dzhng/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Compare Screenshots next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compare Screenshots compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compare Screenshots this skilldzhng/skills1k—~2.6kAutomated safety check: PassMIT
Screenshotnexu-io/open-design100k—~291Automated safety check: PassApache-2.0
Tabler Screenshot Makertabler/tabler42k—~1.3kAutomated safety check: PassMIT
Screenshots Marketingnexu-io/open-design100k—~301Automated safety check: PassApache-2.0
Operator Approval Loopaffaan-m/ECC276k—~3.3kAutomated safety check: PassMIT
PR Screenshotsgithub/awesome-copilot40k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Screenshot

    nexu-io/open-design

    Capture desktop, app windows, or pixel regions across OS platforms.

    100k GitHub stars~291 tokensUpdated today
    Testing & QAAuto-check passed
  • Makes reproducible screenshots of Tabler components and screens for release notes, the website and social posts, using the repository's screenshots app.

    42k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Screenshots Marketing

    nexu-io/open-design

    Generate marketing screenshots with Playwright. An agent skill from nexu-io/open-design.

    100k GitHub stars~301 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Operator approval contract with internal filing notices for agent-drafted outbound messages, hashed drafts, epoch-keyed decisions, durable delivery claims and receipts, and a pre-draft baseline gate.

    276k GitHub stars~3.3k tokensUpdated 4 days ago
    Auto-check passed
  • PR Screenshots

    github/awesome-copilot

    Official

    Embed before/after screenshots and annotated images in pull request descriptions.

    40k GitHub stars~1.2k tokensUpdated today
    DevelopmentAuto-check passed
  • UI Screenshots

    github/awesome-copilot

    Official

    Capture screenshots of web apps during development using Playwright and PIL.

    40k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed

More from dzhng/skills

All 27 skills in this repo
  • Claude

    dzhng/skills

    Use Claude Code as an independent claude -p subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude.

    1k GitHub stars~1.3k tokensUpdated 3 days ago
    Auto-check passed
  • Refactor Clean

    dzhng/skills

    Refactor cleanly instead of layering sediment. An agent skill from dzhng/skills.

    1k GitHub stars~3.1k tokensUpdated 3 days ago
    Auto-check passed
  • Write Skills

    dzhng/skills

    Create or revise agent skills. An agent skill from dzhng/skills.

    1k GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed
  • Codex

    dzhng/skills

    Use the local Codex CLI as an independent second agent. An agent skill from dzhng/skills.

    1k GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check: warnings
  • Audit Agents

    dzhng/skills

    Audit or rewrite AGENTS.md so it holds only lasting principles.

    1k GitHub stars~847 tokensUpdated 3 days ago
    Auto-check passed
  • Audit Choices

    dzhng/skills

    Audit the choices an implementing agent made, not its diff — a pure decision audit that traces the session's history into a choices ledger, changes no code, and never blocks an unsupervised run.

    1k GitHub stars~2.5k tokensUpdated 3 days ago
    Auto-check passed

Questions about Compare Screenshots

What does Compare Screenshots do?

Compare screenshots against the intended design, distinguishing approved references from historical baselines. Compare Screenshots is an agent skill from dzhng/skills. Compare screenshots against the intended design, distinguishing approved references from historical baselines.

When should I use Compare Screenshots?

Compare Screenshots fits situations like: iterative implementation matching; before/after visual review; objective image telemetry; checking a lone capture for flat.

How do I install Compare Screenshots in Claude Code?

Run `npx skills add dzhng/skills --skill compare-screenshots -a claude-code`. Or copy the skill folder (skills/visual/compare-screenshots in dzhng/skills) into .claude/skills/compare-screenshots in your project. Claude Code loads it when a task matches its description.

How do I install Compare Screenshots in Codex?

Run `npx skills add dzhng/skills --skill compare-screenshots -a codex`. Or copy the skill folder (skills/visual/compare-screenshots in dzhng/skills) into .agents/skills/compare-screenshots in your project. Codex loads it when a task matches its description.

Can I use Compare Screenshots in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dzhng/skills --skill compare-screenshots -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compare-screenshots, .gemini/skills/compare-screenshots, .github/skills/compare-screenshots and .opencode/skills/compare-screenshots in your project.

What does Compare Screenshots need to run?

Going by SKILL.md and its folder, Compare Screenshots needs JavaScript for the scripts in its folder. Our summary lists: Node.js.

Does Compare Screenshots access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compare Screenshots safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Compare Screenshots use?

Compare Screenshots is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compare Screenshots use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Compare Screenshots?

Skills that share tags, products or a category with Compare Screenshots: Screenshot (nexu-io/open-design, 100k stars), Tabler Screenshot Maker (tabler/tabler, 42k stars), Screenshots Marketing (nexu-io/open-design, 100k stars) and Operator Approval Loop (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compare Screenshots?

dzhng (a GitHub user) maintains it in dzhng/skills, which has 1,020 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 5, 2026.

Source: dzhng/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.