Visual Regression
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Use visual regression testing.
Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses regression-finder --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/regression-finder .claude/skills/regression-finder && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "regression-finder" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/regression-finder into .claude/skills/regression-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "regression-finder", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/regression-finderType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses regression-finder --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/regression-finder .agents/skills/regression-finder && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "regression-finder" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/regression-finder into .agents/skills/regression-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "regression-finder", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses regression-finder --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/regression-finder .cursor/skills/regression-finder && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "regression-finder" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/regression-finder into .cursor/skills/regression-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "regression-finder", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/RyanAlberts/best-of-Agent-Harnesses.git --path skills/regression-finder--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses regression-finder --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/regression-finder .gemini/skills/regression-finder && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "regression-finder" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/regression-finder into .gemini/skills/regression-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "regression-finder", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses regression-finderInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/regression-finder .github/skills/regression-finder && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "regression-finder" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/regression-finder into .github/skills/regression-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "regression-finder", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install RyanAlberts/best-of-Agent-Harnesses regression-finder --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/regression-finder .opencode/skills/regression-finder && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "regression-finder" agent skill from https://github.com/RyanAlberts/best-of-Agent-Harnesses/tree/main/skills/regression-finder into .opencode/skills/regression-finder/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "regression-finder", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
regression-finderRegression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…
Regression Finder is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse, dumber, lazier, or degraded since an update or since they upgraded; asks whether a new Claude Code or Codex version or model made it worse than the old one; wants to know which release, version, or week it regressed in; or wants numbers to…
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `README.md`, `references/metrics.md` and `references/statistics.md`).
The repository describes itself as: 🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4fa20bc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Regression Finder loads about 3k tokens when it runs, and up to ~7.2k if it reads all its reference files. Until then it costs about 175 tokens; SKILL.md has 1,734 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from RyanAlberts/best-of-Agent-Harnesses at commit 4fa20bc, republished under its MIT licence (© RyanAlberts). 1,734 words, ~2,994 tokens.
.claude/skills/regression-finder/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.When a coding agent seems worse after an update, the user's own session history can show whether it changed and when. This skill splits the user's Claude Code or Codex sessions by harness version, model, or week, measures the same behavior in each (reads before the first edit, reads per edit, edits to files not read first, interrupts, corrections, failed tool calls, output, cost), and places the change at the update, or within the span of versions, where the numbers moved, along with anything else that changed at the same point. It reads local transcripts, prints counts and rates, never prints prompt text, and sends nothing anywhere.
session-waste-report.runaway-guard.harness-test-drive.claim-check.rules-to-guards.Tell the user this when they ask what the check looks at:
~/.claude/projects/, Codex under ~/.codex/sessions/), only those changed within the window. Prompt text is read on this machine only to spot corrections such as "no, that's wrong".--out or --svg, which write the one file named.<skill-dir> means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Keep the quotes around the script path in every command: skill folders can sit under paths with spaces.
Run the default check. Pick the harness the user asks about (claude-code by default, or codex) and run:
python3 "<skill-dir>/scripts/regress.py" --harness claude-codeIt reads the last 90 days and splits by harness version. Match the flags to the question:
| The user says | Add |
|---|---|
| "since the last Codex update" | --harness codex |
| "since I switched to the new model" | --by model |
| "worse these past few weeks" | --by week |
| "in my api repo" | --project <that folder> |
| "since the spring" | --since 180d |
It takes a few seconds per thousand sessions and exits 0 even when it finds nothing. Done when the output starts with a bold headline, or you have told the user the exact error.
Answer the user's own question first. If the user named an update or a time such as last week, find it in the By version table first. If it was not tested, lead with that (for example: 2.1.280 had 7 sessions, and the test needs 20 on each side), then give the headline as background, not as the answer. The notes list every update that was not tested, with its sessions.
Then read the headline. It is one of these kinds:
2.1.270, your agent reads 41% less before it edits...", or "After an update between Claude Code 2.1.260 and 2.1.270, ..." when the test had to borrow neighboring versions. The change lies somewhere in that span; say the span, not one version. Go to step 3.--since 180d, when the history goes back that far), a split by time (--by week), and last a lower bar (--min-sessions 12, the floor). A lower bar tests more updates, but each test can only catch larger changes.Done when the user's own update or time is answered, or you know which kind of headline it is.
Follow each confounder. A confounder is anything else that changed at the same update and could explain the numbers. The section "What else changed at the same point" lists them, and the headline ends with "but ... changed at the same point" or "both sides were in use at the same time" when they exist. For each line:
--by model and see whether the change follows the model instead.--by version.--project argument the line prints. Copy the --project argument exactly as printed, quotes included.Done when each confounder line has a rerun result or a one-sentence explanation for the user.
Offer the chart when something is flagged and the user wants to see it or share it:
python3 "<skill-dir>/scripts/regress.py" --harness claude-code --svg regression.svgIt writes one small line chart per flagged number to the path given, and nothing else. Done when you have given the user the path, or the user declined.
Report in the shape below.
--json prints the same result for a program: headline, totals, thresholds, slices, tests (every comparison, flagged or not), flagged, stand_outs, confounders, untested, left_out, and notes.Open references/metrics.md to explain what a number measures and why it matters, and references/statistics.md when the user asks how sure the result is or why a visible change was not flagged.
--project <path> or --by model.--json for the data and --svg for the chart.session-waste-report.An example of the shape (the numbers are invented):
After Claude Code
2.1.270, your agent reads 41% less before it edits and gets interrupted twice as often.
Where What changed Before After Change Sessions after 2.1.270Reads before the first edit (session median) 5 3 -41% 34, 29 after 2.1.270Interrupts per 100 turns (session mean) 4.1 8.3 +102% 34, 29 Nothing else changed at that update: same model, same projects. Next: read the
2.1.270entry in the Claude Code changelog, and if it matches, file it with--jsonand--svg.
When nothing was flagged, say what was compared (versions, sessions, turns) and that small histories cannot show small changes; quote no percentages as findings.
Quote versions, models, and paths exactly as the report prints them, inside inline code: it shows the home folder as ~, and it has already made any text taken from transcripts safe to display.
scripts/regress.py: the check. Flags: --harness, --by version|model|week, --since 90d, --project, --min-turns 30, --min-sessions 20 (at least 12), --svg <path>, --json, --out <path>, --fail-on worse|any (exit 1 when a flagged change usually means worse, or when anything is flagged).scripts/transcripts.py, scripts/pricing.py, scripts/safe.py: the shared reader for session files, the price table, and the text cleaner that puts transcript text in the report inside inline code; several skills in this repository use them.references/metrics.md: each number's definition and why it matters, next to the method of issue #42796.references/statistics.md: the test, the thresholds, the minimum samples, how changes are placed, and the limits.© RyanAlberts, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in skills/regression-finder of RyanAlberts/best-of-Agent-Harnesses.
Open the folder on GitHubat commit 4fa20bc
Regression Finder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Regression Finder this skillRyanAlberts/best-of-Agent-Harnesses | 1.1k | — | ~3k | Automated safety check: Pass | MIT | |
| Visual Regressionthedaviddias/Front-End-Checklist | 74k | — | ~493 | Automated safety check: Pass | MIT | |
| AI Regression Testingaffaan-m/ECC | 277k | 5 repos | ~2.9k | Automated safety check: Pass | MIT | |
| AI Regression Testingaffaan-m/ECC | 277k | 1 repos | ~2.2k | Automated safety check: Pass | MIT | |
| AI Regression Testingaffaan-m/ECC | 276k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Regression Suite MaintenanceDonchitos/Claude-Code-Game-Studios | 26k | — | ~3.5k | Automated safety check: Pass | MIT |
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Use visual regression testing.
affaan-m/ECC
Regression testing strategies for AI-assisted development. An agent skill from affaan-m/ECC.
affaan-m/ECC
AI辅助开发的回归测试策略。沙盒模式API测试,无需依赖数据库,自动化的缺陷检查工作流程,以及捕捉AI盲点的模式,其中同一模型编写和审查代码。
affaan-m/ECC
AI 支援開発のためのリグレッションテスト戦略。データベース依存なしのサンドボックスモード API テスト、自動化されたバグチェックワークフロー、同じモデルがコードを書いてレビューする AI のブラインドスポットを捕捉するパターン。
Donchitos/Claude-Code-Game-Studios
Maps existing tests to a game's critical paths, finds fixed bugs that lack regression tests and flags coverage drift as new features arrive.
sanity-io/sanity
Add, review, and maintain Chromatic visual regression coverage in the Sanity monorepo via dev/storybook stories, the vitest browser-mode suite, and Playwright e2e snapshots.
RyanAlberts/best-of-Agent-Harnesses
Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot instructions) each coding agent loads from a repo, what gets cut or skipped, and whether the commands those…
RyanAlberts/best-of-Agent-Harnesses
Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed…
RyanAlberts/best-of-Agent-Harnesses
Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including…
RyanAlberts/best-of-Agent-Harnesses
Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests…
RyanAlberts/best-of-Agent-Harnesses
Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding agent keeps breaking, counts every violation in recent Claude Code, Codex, Gemini CLI, and OpenCode…
RyanAlberts/best-of-Agent-Harnesses
Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on the same tool call, keeps failing, or exceeds a dollar cap.
Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…. Regression Finder is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed.
Regression Finder fits situations like: the user says the agents quality dropped; degraded since an update; since they upgraded; asks whether a new Claude Code.
Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a claude-code`. Or copy the skill folder (skills/regression-finder in RyanAlberts/best-of-Agent-Harnesses) into .claude/skills/regression-finder in your project. Claude Code loads it when a task matches its description.
Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a codex`. Or copy the skill folder (skills/regression-finder in RyanAlberts/best-of-Agent-Harnesses) into .agents/skills/regression-finder in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regression-finder, .gemini/skills/regression-finder, .github/skills/regression-finder and .opencode/skills/regression-finder in your project.
Going by SKILL.md and its folder, Regression Finder needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Regression Finder is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Regression Finder: Visual Regression (thedaviddias/Front-End-Checklist, 74k stars), AI Regression Testing (affaan-m/ECC, 277k stars), AI Regression Testing (affaan-m/ECC, 277k stars) and AI Regression Testing (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
RyanAlberts (a GitHub user) maintains it in RyanAlberts/best-of-Agent-Harnesses, which has 1,133 GitHub stars. The repository was last updated on October 9, 2026.
Source: RyanAlberts/best-of-Agent-Harnesses on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.