Agent skill

Running Visual Regression Tests

by jaktestowac in jaktestowac/awesome-copilot-for-testers

Sets up and maintains visual regression testing: what to snapshot, baseline strategy, masking dynamic regions, threshold tuning, containerized baselines, and the review-and-update workflow.

MITAuto-check passedTesting & QA

Install Running Visual Regression Tests

skills CLI
$ npx skills add jaktestowac/awesome-copilot-for-testers --skill running-visual-regression-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaktestowac/awesome-copilot-for-testers running-visual-regression-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaktestowac/awesome-copilot-for-testers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/running-visual-regression-tests .claude/skills/running-visual-regression-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
running-visual-regression-tests
GitHub stars
116
Token cost
~2.6k tokens
SKILL.md length
1,347 words
Files
5
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Sets up and maintains visual regression testing: what to snapshot, baseline strategy, masking dynamic regions, threshold tuning, containerized baselines, and the review-and-update workflow.

  • Works in 7 steps: Decide whether pixels are the right tool → Stabilize the page → Choose the snapshot surface → …
  • Styling regressions escape to production
  • SKILL.md covers When to Use, Operating Principles, Workflow and Common Failure Modes, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Running Visual Regression Tests is an agent skill from jaktestowac/awesome-copilot-for-testers. Sets up and maintains visual regression testing: what to snapshot, baseline strategy, masking dynamic regions, threshold tuning, containerized baselines, and the review-and-update workflow. Use when styling regressions escape to production, when snapshots fail on every machine or every run, when baselines are being updated without being looked at, or when deciding whether visual testing is the right tool at all.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `resources/baseline-policy.md`, `resources/masking-and-stabilization.md` and `resources/visual-config-recipes.md`).

It sits in Testing & QA, covering Visual regression testing. The repository describes itself as: 👨💻 Instructions, prompts, and chat modes to help You with test automation for GitHub Copilot 🤖. The licence is MIT.

When your agent uses it

  • Styling regressions escape to production
  • Snapshots fail on every machine
  • Baselines are being updated without being looked at
  • Deciding whether visual testing is the right tool at all

Example prompts

  • “/running-visual-regression-tests”

Requirements

  • Docker

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Decide whether pixels are the right tool
  2. Stabilize the page
  3. Choose the snapshot surface
  4. Set the baseline policy
  5. Configure comparison
  6. Run the review workflow
  7. Keep the suite affordable

What it can do on your machine

Read from SKILL.md and the folder at commit 8910672. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Running Visual Regression Tests loads about 2.6k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 1,347 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaktestowac/awesome-copilot-for-testers at commit 8910672, republished under its MIT licence (© jaktestowac). 1,347 words, ~2,564 tokens.

Download SKILL.mdSave it as .claude/skills/running-visual-regression-tests/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
running-visual-regression-tests
description
Sets up and maintains visual regression testing: what to snapshot, baseline strategy, masking dynamic regions, threshold tuning, containerized baselines, and the review-and-update workflow. Use when styling regressions escape to production, when snapshots fail on every machine or every run, when baselines are being updated without being looked at, or when deciding whether visual testing is the right tool at all.
argument-hint
Pages or components to cover, the current snapshot setup if any, the CI platform, and the failure being investigated
user-invocable
true

Running Visual Regression Tests

Use this skill when appearance is the thing under test and functional assertions cannot express it: a broken layout at one breakpoint, a theme token that stopped applying, a component that shifted three pixels and swallowed a button.

Visual testing has one failure mode that dominates all others: a suite so noisy that --update-snapshots becomes reflex. At that point the baselines record whatever the code currently does, the diffs are never read, and the suite costs money while proving nothing. Everything below exists to prevent that.

When to Use

  • styling and layout regressions reach production while the functional suite stays green
  • snapshots pass locally and fail in CI on every run
  • baselines get regenerated wholesale because nobody can tell signal from noise
  • a design system or theme change needs a blast-radius check
  • deciding whether a case belongs in a visual test, an assertion, or an accessibility check

Operating Principles

  • A visual test earns its place only if a functional assertion cannot do the job. Text content, disabled state, and element presence are assertions, not pixels.
  • Determinism before comparison. Fix fonts, animation, time, data, and viewport first. Threshold tuning applied to an unstable page hides real diffs along with the noise.
  • Baselines are generated in one environment. The container that produces them is the container that compares them.
  • Every diff is looked at by a human before it becomes a baseline. A bulk update with no review is the failure mode, not a shortcut.
  • Snapshot the smallest surface that carries the risk. Component snapshots localize failures; full-page snapshots turn one change into forty red tests.
  • A masked region is a documented decision. Masking is how the suite stays quiet, and also how coverage silently disappears.

Workflow

Phase 0: Decide whether pixels are the right tool

Run through this before writing anything.

The riskRight tool
Wrong text, wrong label, wrong countAssertion
Element missing or disabledAssertion
Contrast, focus order, ARIAauditing-accessibility
Layout broken at a breakpointVisual
Theme token stopped applyingVisual
Component spacing driftedVisual, component-level
Chart or canvas renderingVisual, with a fixed data seed
Third-party embed changedNeither; stub it and test your side

If most rows land outside the visual column, say so and stop. A visual suite added to compensate for missing assertions makes both worse.

Phase 1: Stabilize the page

Nothing about thresholds until this phase is finished. ./resources/masking-and-stabilization.md has the code for each item.

  • Fonts: wait for document.fonts.ready, or self-host and preload. A late webfont swap is the single most common cause of a one-off diff.
  • Animation and transitions: disable globally via injected CSS or animations: 'disabled'.
  • Time: pin the clock so relative timestamps stop moving. See mocking-network-and-time.
  • Data: seed deterministically. A list ordered by an unstable sort key produces a real diff every run.
  • Network: stub anything that varies, including avatars, ads, and telemetry.
  • Viewport and device scale: set explicitly per project.
  • Scrollbars and caret: hide them; they differ across platforms and focus states.
  • Randomized content: ids, session tokens, generated names.

The completion criterion for this phase is concrete: the same test run three times in a row on the same commit produces zero diffs. Do not proceed until that holds.

Phase 2: Choose the snapshot surface
SurfaceUse forCost of a change
Component / storyDesign system, shared UIOne test per affected component
Region (locator.screenshot)A card, a header, a tableLocalized
Full pageLayout and page compositionOne shared-component change reddens every page
Full page, fullPage: trueLong-scroll layoutsHighest noise; lazy-loading and sticky headers add diffs

Default to region-level. Add full-page snapshots for a small set of layout-critical routes, and know that they will be the noisiest tests in the suite.

Phase 3: Set the baseline policy

Decide and record, using ./resources/baseline-policy.md:

  • Where baselines are generated: a container image pinned by digest, matching the CI runner. Locally generated baselines and CI comparison is the second most common cause of permanent redness.
  • Which platforms have baselines: one per project (browser, viewport, theme, locale). Every added axis multiplies the file count and the review burden.
  • Where baselines live: in the repository with the tests, or in a hosted service. In-repo is simpler and grows the repository; state which you chose.
  • Who may update them: the rule that stops reflex updates. A pull request that changes baselines carries the rendered before-and-after in its description.
Phase 4: Configure comparison

Start strict and loosen only against observed noise.

ts
expect: {
  toHaveScreenshot: {
    maxDiffPixelRatio: 0.001,
    threshold: 0.2,          // per-pixel colour tolerance
    animations: 'disabled',
    caret: 'hide',
    scale: 'css',
  },
},

Rules for loosening:

  • change one knob at a time and record which noise it addressed
  • prefer masking a known-dynamic region over raising the global ratio
  • a maxDiffPixelRatio above roughly 0.01 hides a missing button; if you need that much, go back to Phase 1
  • never set a per-test tolerance without a comment naming the source of the noise

Config snippets, including the anti-aliasing and platform-difference cases, are in ./resources/visual-config-recipes.md.

Show full SKILL.md (539 more words)Show less
Phase 5: Run the review workflow

The workflow in ./resources/visual-review-workflow.md, in short:

  1. CI produces the diff triplet (expected, actual, diff) as artifacts on failure.
  2. A human opens the diff and reaches a verdict: intended change, regression, or noise.
  3. Intended changes get their baselines updated in the same pull request as the code change, with the diff image visible to the reviewer.
  4. Regressions become bugs. Hand off to reporting-bugs.
  5. Noise goes back to Phase 1. Raising a threshold to silence noise is the last resort, and it is recorded.

The verdict is mandatory per failing test. "Updated all snapshots" without per-test verdicts is the anti-pattern this skill is written to prevent.

Phase 6: Keep the suite affordable

Review periodically:

  • snapshots that have never caught a regression and never failed: candidates for deletion
  • snapshots that fail on more than a third of unrelated pull requests: too broad, split or mask
  • total suite runtime and artifact storage
  • baseline count per platform axis; drop axes that never diverge

A visual suite is worth its cost when it has caught something. Write down what it has caught.

Common Failure Modes

  • running --update-snapshots across the whole suite to clear a red build
  • generating baselines on a developer laptop and comparing them on a Linux CI runner
  • raising the diff ratio until the suite goes quiet, then discovering it no longer detects a missing element
  • full-page snapshots of every route, so one header change produces forty failures with one cause
  • masking a region and never recording why, until nobody knows what stopped being covered
  • visual tests standing in for assertions the suite should have had
  • baselines committed without the diff image visible in the pull request
  • no artifact upload on failure, so the only way to see the diff is to reproduce it locally

Resource Map

  • ./resources/masking-and-stabilization.md - fonts, animation, time, data, scrollbars, masking, and the three-identical-runs gate
  • ./resources/baseline-policy.md - where baselines are generated and stored, platform axes, update authority, containerized generation
  • ./resources/visual-config-recipes.md - Playwright config and per-test options, Docker baseline generation, CI artifact upload
  • ./resources/visual-review-workflow.md - the per-failure verdict process, pull request checklist, and escalation to a bug
  • mocking-network-and-time - when snapshot instability comes from live data or a moving clock
  • stabilizing-flaky-tests (planned) - when the instability is timing rather than rendering
  • auditing-accessibility - when the concern is contrast, focus, or semantics rather than appearance
  • ui-playwright-test-developer (planned) - when the visual checks sit inside a wider browser suite
  • automating-ci-test-pipelines (planned) - for baseline containers, artifact retention, and sharding the visual suite
  • reporting-bugs - when a diff is confirmed as a regression

Definition of Done

This skill is complete when:

  • every visual test exists because an assertion could not cover the risk, and the alternative was considered
  • three consecutive runs on an unchanged commit produce zero diffs
  • baselines are generated in a pinned container that matches the comparison environment
  • the snapshot surface is as small as the risk allows, and full-page snapshots are a deliberate short list
  • comparison starts strict; every loosened threshold names the noise it addresses
  • every masked region has a recorded reason and a note of what it stops covering
  • failures produce expected, actual, and diff artifacts in CI
  • each failing snapshot carries a per-test verdict of intended, regression, or noise before any baseline is written

© jaktestowac, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/running-visual-regression-tests of jaktestowac/awesome-copilot-for-testers.

  • SKILL.md
  • resources/baseline-policy.md
  • resources/masking-and-stabilization.md
  • resources/visual-config-recipes.md
  • resources/visual-review-workflow.md

Open the folder on GitHubat commit 8910672

Compare with similar skills

Running Visual Regression Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Running Visual Regression Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Running Visual Regression Tests this skilljaktestowac/awesome-copilot-for-testers116—~2.6kAutomated safety check: PassMIT
Review Local UI Screenshotsxiaocang/easydict_win32101—~1.1kAutomated safety check: PassGPL-3.0
Handsontable Visual Test Demoshandsontable/handsontable22k—~1.3kAutomated safety check: PassCustom licence
Kc Screenshotimran31415/kube-coder388—~870Automated safety check: NotesMIT
Visual QAliangdabiao/Godogen127—~1.4kAutomated safety check: PassMIT
Meticulous FixFlintSH/Flare135—~2.3kAutomated safety check: PassMIT

Similar skills

  • Review Local UI Screenshots

    xiaocang/easydict_win32

    Review Easydict UI automation screenshot artifacts already present in local artifacts/ui-screenshots, screenshots, or a user-provided artifact directory.

    101 GitHub stars~1.1k tokensUpdated 16 days ago
    Testing & QAAuto-check passed
  • Handsontable Visual Test Demos

    handsontable/handsontable

    Explains how to add or change the demo pages that Handsontable's visual regression suite photographs, including per-feature routes in the js demo and the shared grid.

    22k GitHub stars~1.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Kc Screenshot

    imran31415/kube-coder

    Capture desktop + mobile, dark + light screenshots of the kube-coder dashboard SPA for visual QA of a UI change.

    388 GitHub stars~870 tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Visual QA

    liangdabiao/Godogen

    Visual quality assurance — analyze game screenshots for defects, compare against reference, check motion in frame sequences.

    127 GitHub stars~1.4k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Meticulous Fix

    FlintSH/Flare

    Fix the visual diffs that have been reviewed and rejected on a Meticulous test run, following their review comments if given.

    135 GitHub stars~2.3k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Visual QA For Web And Terminal UIs

    code-yeongyu/oh-my-openagent

    Checks a built UI against a human-interface checklist and any reference image, pairing day and night screenshots at two screen widths.

    70k GitHub stars~9.5k tokensUpdated today
    Testing & QAAuto-check passed

More from jaktestowac/awesome-copilot-for-testers

All 13 skills in this repo
  • API Playwright Test Developer

    jaktestowac/awesome-copilot-for-testers

    Writes and reviews API automation tests with Playwright Test, covering setup/teardown, assertions, data management, and hybrid API+UI flows.

    116 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Assessing Comprehension Debt

    jaktestowac/awesome-copilot-for-testers

    Measures the risk that code shipped without anyone understanding it: a teach-back attestation on high-risk changes, a risk band from changed-code complexity, diff size and whether a human…

    116 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Orchestration Packs

    jaktestowac/awesome-copilot-for-testers

    Creates agent orchestration packs: cooperating .agent.md files with an orchestrator, subagents, matched handoffs, minimal tool grants, and a shared handoff packet contract.

    116 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Plugins

    jaktestowac/awesome-copilot-for-testers

    Packages repository skills as installable Copilot plugins: marketplace registration, plugin.json manifests, generated skill copies, and the sync check CI enforces.

    116 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Governing Quality Waivers

    jaktestowac/awesome-copilot-for-testers

    Turns "we will skip this check for now" into a dated, attributed, expiring waiver with a stated reason and owner, inventories the silent skips already hiding in a repo - skipped tests, disabled lint…

    116 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Recording Change Intent

    jaktestowac/awesome-copilot-for-testers

    Requires an externalised rationale for high-risk changes - new public exports, new endpoints, auth edits, migrations, removed guards - recorded as an Intent commit trailer, an ADR reference, or a…

    116 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Running Visual Regression Tests

What does Running Visual Regression Tests do?

Sets up and maintains visual regression testing: what to snapshot, baseline strategy, masking dynamic regions, threshold tuning, containerized baselines, and the review-and-update workflow. Running Visual Regression Tests is an agent skill from jaktestowac/awesome-copilot-for-testers. Sets up and maintains visual regression testing: what to snapshot, baseline strategy, masking dynamic regions, threshold tuning, containerized baselines, and the review-and-update workflow.

When should I use Running Visual Regression Tests?

Running Visual Regression Tests fits situations like: styling regressions escape to production; snapshots fail on every machine; baselines are being updated without being looked at; deciding whether visual testing is the right tool at all.

How do I install Running Visual Regression Tests in Claude Code?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill running-visual-regression-tests -a claude-code`. Or copy the skill folder (skills/running-visual-regression-tests in jaktestowac/awesome-copilot-for-testers) into .claude/skills/running-visual-regression-tests in your project. Claude Code loads it when a task matches its description.

How do I install Running Visual Regression Tests in Codex?

Run `npx skills add jaktestowac/awesome-copilot-for-testers --skill running-visual-regression-tests -a codex`. Or copy the skill folder (skills/running-visual-regression-tests in jaktestowac/awesome-copilot-for-testers) into .agents/skills/running-visual-regression-tests in your project. Codex loads it when a task matches its description.

Can I use Running Visual Regression Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaktestowac/awesome-copilot-for-testers --skill running-visual-regression-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/running-visual-regression-tests, .gemini/skills/running-visual-regression-tests, .github/skills/running-visual-regression-tests and .opencode/skills/running-visual-regression-tests in your project.

What does Running Visual Regression Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Running Visual Regression Tests is instructions for the agent only. Our summary lists: Docker.

Does Running Visual Regression Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Running Visual Regression Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Running Visual Regression Tests use?

Running Visual Regression Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Running Visual Regression Tests use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Running Visual Regression Tests?

Skills that share tags, products or a category with Running Visual Regression Tests: Review Local UI Screenshots (xiaocang/easydict_win32, 101 stars), Handsontable Visual Test Demos (handsontable/handsontable, 22k stars), Kc Screenshot (imran31415/kube-coder, 388 stars) and Visual QA (liangdabiao/Godogen, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Running Visual Regression Tests?

jaktestowac (a GitHub user) maintains it in jaktestowac/awesome-copilot-for-testers, which has 116 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on August 26, 2026.

Source: jaktestowac/awesome-copilot-for-testers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.