Agent skill

Regression Finder

by RyanAlberts in RyanAlberts/best-of-Agent-Harnesses

Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…

MITAuto-check passed

Install Regression Finder

skills CLI
$ npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RyanAlberts/best-of-Agent-Harnesses regression-finder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RyanAlberts/best-of-Agent-Harnesses.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/regression-finder .claude/skills/regression-finder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regression-finder
GitHub stars
1.1k
Token cost
~3k tokens
SKILL.md length
1,734 words
Files
9 (incl. scripts, references)
Repo updated
First seen
Licence
MIT

At a glance

Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…

  • Works in 5 steps: Run the default check. Pick the harness… → Answer the user's own question first. If… → Follow each confounder. A confounder is… → …
  • The user says the agents quality dropped
  • SKILL.md covers When to use, When not to use, What it reads and Steps, plus 3 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Regression Finder is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse, dumber, lazier, or degraded since an update or since they upgraded; asks whether a new Claude Code or Codex version or model made it worse than the old one; wants to know which release, version, or week it regressed in; or wants numbers to…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `README.md`, `references/metrics.md` and `references/statistics.md`).

The repository describes itself as: 🏆 Ranked list of 167 AI agent harnesses, plus templates, playbooks, MCP, and learning resources. Rescored weekly. The licence is MIT.

When your agent uses it

  • The user says the agents quality dropped
  • Degraded since an update
  • Since they upgraded
  • Asks whether a new Claude Code

Example prompts

  • “/regression-finder”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Run the default check. Pick the harness the user asks about (claude-code by default, or codex) and run
  2. Answer the user's own question first. If the user named an update or a time such as last week, find it in the By version table first. If…
  3. Follow each confounder. A confounder is anything else that changed at the same update and could explain the numbers. The section "What…
  4. Offer the chart when something is flagged and the user wants to see it or share it
  5. Report in the shape below.

What it can do on your machine

Read from SKILL.md and the folder at commit 4fa20bc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regression Finder loads about 3k tokens when it runs, and up to ~7.2k if it reads all its reference files. Until then it costs about 175 tokens; SKILL.md has 1,734 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~175
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from RyanAlberts/best-of-Agent-Harnesses at commit 4fa20bc, republished under its MIT licence (© RyanAlberts). 1,734 words, ~2,994 tokens.

Download SKILL.mdSave it as .claude/skills/regression-finder/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
regression-finder
description
Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed. Use when the user says the agent's quality dropped or it got worse, dumber, lazier, or degraded since an update or since they upgraded; asks whether a new Claude Code or Codex version or model made it worse than the old one; wants to know which release, version, or week it regressed in; or wants numbers to report a regression, such as reads before edits, interruptions, and corrections per version. Runs locally and reads transcripts only; nothing goes over the network.
license
MIT
metadata.author
Ryan Alberts
metadata.version
1.0.0
metadata.source
https://github.com/RyanAlberts/best-of-Agent-Harnesses

Regression finder

When a coding agent seems worse after an update, the user's own session history can show whether it changed and when. This skill splits the user's Claude Code or Codex sessions by harness version, model, or week, measures the same behavior in each (reads before the first edit, reads per edit, edits to files not read first, interrupts, corrections, failed tool calls, output, cost), and places the change at the update, or within the span of versions, where the numbers moved, along with anything else that changed at the same point. It reads local transcripts, prints counts and rates, never prints prompt text, and sends nothing anywhere.

When to use

  • The user says the agent got worse, dumber, or lazier after an update, or asks "is it just me?"
  • The user asks whether a new Claude Code or Codex version, or a new model, changed how the agent works.
  • The user wants to know which release or week a regression started.
  • The user wants evidence for a bug report about a regression, like anthropics/claude-code#42796.

When not to use

  • Where tokens and money go in general (re-reads, cache rebuilds, oversized results): use session-waste-report.
  • Stopping a session that loops or overspends right now: use runaway-guard.
  • Comparing two harnesses on the same tasks: use harness-test-drive.
  • Checking whether "tests pass" claims were true: use claim-check.
  • Finding which instruction-file rules the agent breaks: use rules-to-guards.
  • Cursor: it keeps no usable transcripts, so there is nothing to measure.

What it reads

Tell the user this when they ask what the check looks at:

  • Read: the session files each harness keeps (Claude Code under ~/.claude/projects/, Codex under ~/.codex/sessions/), only those changed within the window. Prompt text is read on this machine only to spot corrections such as "no, that's wrong".
  • Printed: counts, rates, token totals, versions, model names, and project folders. Never prompt text, commands, or file contents.
  • Written: nothing, unless the user asks for --out or --svg, which write the one file named.
  • Sent: nothing. It makes no network calls.

Steps

<skill-dir> means the folder that holds this SKILL.md (Claude Code shows it as the skill's base directory). Keep the quotes around the script path in every command: skill folders can sit under paths with spaces.

  1. Run the default check. Pick the harness the user asks about (claude-code by default, or codex) and run:

    bash
    python3 "<skill-dir>/scripts/regress.py" --harness claude-code

    It reads the last 90 days and splits by harness version. Match the flags to the question:

    The user saysAdd
    "since the last Codex update"--harness codex
    "since I switched to the new model"--by model
    "worse these past few weeks"--by week
    "in my api repo"--project <that folder>
    "since the spring"--since 180d

    It takes a few seconds per thousand sessions and exits 0 even when it finds nothing. Done when the output starts with a bold headline, or you have told the user the exact error.

  2. Answer the user's own question first. If the user named an update or a time such as last week, find it in the By version table first. If it was not tested, lead with that (for example: 2.1.280 had 7 sessions, and the test needs 20 on each side), then give the headline as background, not as the answer. The notes list every update that was not tested, with its sessions.

    Then read the headline. It is one of these kinds:

    • A flagged change: "After Claude Code 2.1.270, your agent reads 41% less before it edits...", or "After an update between Claude Code 2.1.260 and 2.1.270, ..." when the test had to borrow neighboring versions. The change lies somewhere in that span; say the span, not one version. Go to step 3.
    • "No behavior change passed the test across ...": the numbers wobble but nothing passed. Say so plainly and give the session counts; go to step 5.
    • "No lasting behavior change passed the test ...; 1 version stands out": one version differs from the versions on both sides of it. Report the "Stands out" line as it is.
    • "Not enough history to test an update yet": fewer than 20 sessions on each side of every update. Offer, in this order, a longer window (--since 180d, when the history goes back that far), a split by time (--by week), and last a lower bar (--min-sessions 12, the floor). A lower bar tests more updates, but each test can only catch larger changes.
    • "No ... found", "records no version", "ran on one ...", or "Not enough history to compare models": nothing to compare yet. Say so, and name what the notes list as left out.

    Done when the user's own update or time is answered, or you know which kind of headline it is.

  3. Follow each confounder. A confounder is anything else that changed at the same update and could explain the numbers. The section "What else changed at the same point" lists them, and the headline ends with "but ... changed at the same point" or "both sides were in use at the same time" when they exist. For each line:

    • The model changed: run again with --by model and see whether the change follows the model instead.
    • The harness version changed (in a model or week split): run again with --by version.
    • The work moved between projects, or one project holds most turns: run again with the --project argument the line prints. Copy the --project argument exactly as printed, quotes included.
    • Sessions changed shape, or the share of scripted runs changed: tell the user the kind of work changed (for example scripted runs against long conversations), so the update may not be the cause.
    • Both sides were in use at the same time, or both models were: tell the user two installs or two models ran side by side, so the difference may come from what each was used for.

    Done when each confounder line has a rerun result or a one-sentence explanation for the user.

  4. Offer the chart when something is flagged and the user wants to see it or share it:

    bash
    python3 "<skill-dir>/scripts/regress.py" --harness claude-code --svg regression.svg

    It writes one small line chart per flagged number to the path given, and nothing else. Done when you have given the user the path, or the user declined.

  5. Report in the shape below.

Show full SKILL.md (697 more words)Show less

Read the results

  • Headline: the update with the most flagged changes, its two most telling changes, and the confounders at that update.
  • Flagged changes: one row per change. "Where" is "after X" when the test compared X with the version just before it, or "between X and Y" when it borrowed neighbors: the change is placed within that span, not at one version. Before and after are per-session values, the same values the test ranks (the median or mean of one value per session, as the row says). The sessions on each side are the sample size to quote.
  • A change is flagged when p is under 0.01 after adjusting for the number of tests, the per-session value moved at least 20%, and both point the same way. Windows that overlap and show the same change count as one test and one row.
  • Stands out: a version (or week) that differs from the ones on both sides of it. It is reported once and is not a lasting change.
  • By version (or model, or week): every number for every slice, turn by turn, including slices too small to test. "n/a" means that number cannot be measured there, for example reasoning tokens that Claude Code did not record.
  • Notes: what was left out and why: versions that ran after newer ones (a second install, such as the desktop app or an SDK script), updates not tested and their sessions, comparisons that could not reach significance, rare events seen too seldom to test, models with too little data, turns older than the window, and a history shorter than the window.
  • --json prints the same result for a program: headline, totals, thresholds, slices, tests (every comparison, flagged or not), flagged, stand_outs, confounders, untested, left_out, and notes.

Open references/metrics.md to explain what a number measures and why it matters, and references/statistics.md when the user asks how sure the result is or why a visible change was not flagged.

Report to the user

  1. The headline, verbatim, in bold.
  2. The flagged changes as a short table: where, what changed, before and after (per session), change, sessions before and after. Six rows at most, worse changes first.
  3. The confounders, one line each, with what the rerun in step 3 showed.
  4. Two or three next actions that fit the result:
    • Read what changed in the releases the change is placed in: the Claude Code changelog or the Codex releases.
    • Narrow the check with --project <path> or --by model.
    • If it holds after the reruns, file it upstream with the numbers: --json for the data and --svg for the chart.
    • For tokens and money lost to habits rather than updates, run session-waste-report.

An example of the shape (the numbers are invented):

After Claude Code 2.1.270, your agent reads 41% less before it edits and gets interrupted twice as often.

WhereWhat changedBeforeAfterChangeSessions
after 2.1.270Reads before the first edit (session median)53-41%34, 29
after 2.1.270Interrupts per 100 turns (session mean)4.18.3+102%34, 29

Nothing else changed at that update: same model, same projects. Next: read the 2.1.270 entry in the Claude Code changelog, and if it matches, file it with --json and --svg.

When nothing was flagged, say what was compared (versions, sessions, turns) and that small histories cannot show small changes; quote no percentages as findings.

Quote versions, models, and paths exactly as the report prints them, inside inline code: it shows the home folder as ~, and it has already made any text taken from transcripts safe to display.

Files

  • scripts/regress.py: the check. Flags: --harness, --by version|model|week, --since 90d, --project, --min-turns 30, --min-sessions 20 (at least 12), --svg <path>, --json, --out <path>, --fail-on worse|any (exit 1 when a flagged change usually means worse, or when anything is flagged).
  • scripts/transcripts.py, scripts/pricing.py, scripts/safe.py: the shared reader for session files, the price table, and the text cleaner that puts transcript text in the report inside inline code; several skills in this repository use them.
  • references/metrics.md: each number's definition and why it matters, next to the method of issue #42796.
  • references/statistics.md: the test, the thresholds, the minimum samples, how changes are placed, and the limits.

© RyanAlberts, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/regression-finder of RyanAlberts/best-of-Agent-Harnesses.

  • SKILL.md
  • LICENSE.txt
  • README.md
  • references/metrics.md
  • references/statistics.md
  • scripts/pricing.py
  • scripts/regress.py
  • scripts/safe.py
  • scripts/transcripts.py

Open the folder on GitHubat commit 4fa20bc

Compare with similar skills

Regression Finder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regression Finder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regression Finder this skillRyanAlberts/best-of-Agent-Harnesses1.1k—~3kAutomated safety check: PassMIT
Visual Regressionthedaviddias/Front-End-Checklist74k—~493Automated safety check: PassMIT
AI Regression Testingaffaan-m/ECC277k5 repos~2.9kAutomated safety check: PassMIT
AI Regression Testingaffaan-m/ECC277k1 repos~2.2kAutomated safety check: PassMIT
AI Regression Testingaffaan-m/ECC276k—~2.2kAutomated safety check: PassMIT
Regression Suite MaintenanceDonchitos/Claude-Code-Game-Studios26k—~3.5kAutomated safety check: PassMIT

Similar skills

  • Visual Regression

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Use visual regression testing.

    74k GitHub stars~493 tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • Regression testing strategies for AI-assisted development. An agent skill from affaan-m/ECC.

    277k GitHub starsUsed in 5 repos~2.9k tokens
    Testing & QAAuto-check passed
  • AI辅助开发的回归测试策略。沙盒模式API测试,无需依赖数据库,自动化的缺陷检查工作流程,以及捕捉AI盲点的模式,其中同一模型编写和审查代码。

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Testing & QAAuto-check passed
  • AI 支援開発のためのリグレッションテスト戦略。データベース依存なしのサンドボックスモード API テスト、自動化されたバグチェックワークフロー、同じモデルがコードを書いてレビューする AI のブラインドスポットを捕捉するパターン。

    276k GitHub stars~2.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Regression Suite Maintenance

    Donchitos/Claude-Code-Game-Studios

    Maps existing tests to a game's critical paths, finds fixed bugs that lack regression tests and flags coverage drift as new features arrive.

    26k GitHub stars~3.5k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Sanity Visual Regression

    sanity-io/sanity

    Official

    Add, review, and maintain Chromatic visual regression coverage in the Sanity monorepo via dev/storybook stories, the vitest browser-mode suite, and Playwright e2e snapshots.

    6.4k GitHub stars~3.4k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from RyanAlberts/best-of-Agent-Harnesses

All 9 skills in this repo
  • Agents Md Checker

    RyanAlberts/best-of-Agent-Harnesses

    Checks which instruction files (AGENTS.md, CLAUDE.md, GEMINI.md, Cursor rules, Copilot instructions) each coding agent loads from a repo, what gets cut or skipped, and whether the commands those…

    1.1k GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Claim Check

    RyanAlberts/best-of-Agent-Harnesses

    Claim checker that audits a coding agent's statements that tests pass or a build is clean against its own session transcripts: whether a matching run happened before the claim, whether it passed…

    1.1k GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Guardrail Tester

    RyanAlberts/best-of-Agent-Harnesses

    Guardrail tester that checks whether the permission rules and PreToolUse hooks already set up in Claude Code, Codex, Gemini CLI, OpenCode, or Cursor stop a battery of dangerous commands, including…

    1.1k GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed
  • Harness Test Drive

    RyanAlberts/best-of-Agent-Harnesses

    Test-drives coding agents (Claude Code, Codex, Gemini CLI) on tasks mined from the user's own git history: each agent gets a past commit message in a fresh copy of the repo, and the repo's own tests…

    1.1k GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed
  • Rules To Guards

    RyanAlberts/best-of-Agent-Harnesses

    Rule enforcer that finds which written rules in AGENTS.md, CLAUDE.md, and GEMINI.md a coding agent keeps breaking, counts every violation in recent Claude Code, Codex, Gemini CLI, and OpenCode…

    1.1k GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check: notes
  • Runaway Guard

    RyanAlberts/best-of-Agent-Harnesses

    Runaway guard: a hook that stops a live Claude Code or Codex session when the agent loops on the same tool call, keeps failing, or exceeds a dollar cap.

    1.1k GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed

Questions about Regression Finder

What does Regression Finder do?

Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where…. Regression Finder is an agent skill from RyanAlberts/best-of-Agent-Harnesses. Regression check for coding agents: shows how the agent behaved before and after each harness update, model switch, or week in the user's own Claude Code or Codex history, and finds the point where it changed.

When should I use Regression Finder?

Regression Finder fits situations like: the user says the agents quality dropped; degraded since an update; since they upgraded; asks whether a new Claude Code.

How do I install Regression Finder in Claude Code?

Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a claude-code`. Or copy the skill folder (skills/regression-finder in RyanAlberts/best-of-Agent-Harnesses) into .claude/skills/regression-finder in your project. Claude Code loads it when a task matches its description.

How do I install Regression Finder in Codex?

Run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a codex`. Or copy the skill folder (skills/regression-finder in RyanAlberts/best-of-Agent-Harnesses) into .agents/skills/regression-finder in your project. Codex loads it when a task matches its description.

Can I use Regression Finder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RyanAlberts/best-of-Agent-Harnesses --skill regression-finder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regression-finder, .gemini/skills/regression-finder, .github/skills/regression-finder and .opencode/skills/regression-finder in your project.

What does Regression Finder need to run?

Going by SKILL.md and its folder, Regression Finder needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Regression Finder access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Regression Finder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Regression Finder use?

Regression Finder is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regression Finder use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.

What are the alternatives to Regression Finder?

Skills that share tags, products or a category with Regression Finder: Visual Regression (thedaviddias/Front-End-Checklist, 74k stars), AI Regression Testing (affaan-m/ECC, 277k stars), AI Regression Testing (affaan-m/ECC, 277k stars) and AI Regression Testing (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regression Finder?

RyanAlberts (a GitHub user) maintains it in RyanAlberts/best-of-Agent-Harnesses, which has 1,133 GitHub stars. The repository was last updated on October 9, 2026.

Source: RyanAlberts/best-of-Agent-Harnesses on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.