Agent skill

Regression Consistency Checker

by ArabelaTso in ArabelaTso/Skills-4-SE

Checks whether a new version of a repository preserves the behavior observed by tests on the old version.

Apache-2.0Auto-check passedDevelopment

Install Regression Consistency Checker

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill regression-consistency-checker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE regression-consistency-checker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/regression-consistency-checker .claude/skills/regression-consistency-checker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regression-consistency-checker
GitHub stars
253
Token cost
~2.1k tokens
SKILL.md length
451 words
Files
3 (incl. scripts, references)
Skills in repo
170
Repo updated
First seen
Licence
Apache-2.0

At a glance

Checks whether a new version of a repository preserves the behavior observed by tests on the old version.

  • Works in 8 steps: Prepare Versions → Run Tests on Old Version → Run Tests on New Version → …
  • Comparing two versions of code to detect regressions
  • SKILL.md covers Workflow, Quick Reference, Helper Script and Best Practices
  • Runs Python scripts from its folder; calls git, python and pytest

What it does

Regression Consistency Checker is an agent skill from ArabelaTso/Skills-4-SE. Checks whether a new version of a repository preserves the behavior observed by tests on the old version. Use this skill when comparing two versions of code to detect regressions, verify refactoring safety, validate bug fixes don't break existing functionality, or ensure backward compatibility. Detects differences in function outputs, exceptions, observable states, and performance between versions. Generates reports highlighting potential regressions (critical, high, medium, low severity), improvements, and areas…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/detection_strategies.md` and `scripts/compare_results.py`).

It sits in Development, covering Debugging and Refactoring. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • Comparing two versions of code to detect regressions
  • Verify refactoring safety
  • Validate bug fixes dont break existing functionality
  • Ensure backward compatibility

Example prompts

  • “Use the regression-consistency-checker skill to check whether a new version of a repository preserves the behavior observed by tests on the old…”
  • “/regression-consistency-checker”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Prepare Versions
  2. Run Tests on Old Version
  3. Run Tests on New Version
  4. Compare Results
  5. Analyze Regressions
  6. Investigate Root Causes
  7. Document Findings
  8. Fix or Accept Changes

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • python
    • pytest
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regression Consistency Checker loads about 2.1k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 191 tokens; SKILL.md has 451 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~191
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 451 words, ~2,092 tokens.

Download SKILL.mdSave it as .claude/skills/regression-consistency-checker/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
regression-consistency-checker
description
Checks whether a new version of a repository preserves the behavior observed by tests on the old version. Use this skill when comparing two versions of code to detect regressions, verify refactoring safety, validate bug fixes don't break existing functionality, or ensure backward compatibility. Detects differences in function outputs, exceptions, observable states, and performance between versions. Generates reports highlighting potential regressions (critical, high, medium, low severity), improvements, and areas requiring verification. Triggers when users ask to check for regressions between versions, compare test behavior across versions, verify behavior preservation, or validate that changes don't break existing tests.

Regression Consistency Checker

Check whether a new version of a repository preserves the behavior observed by tests on the old version.

Workflow

1. Prepare Versions

Set up old version:

bash
# Tag or note the old version
git tag old-version

# Or checkout specific commit
git checkout <old-commit-hash>

Set up new version:

bash
# Tag the new version
git tag new-version

# Or checkout new commit
git checkout <new-commit-hash>

Ensure clean environment:

  • Same dependencies installed
  • Same test configuration
  • Same environment variables
  • Deterministic test execution (fix random seeds, mock time)
2. Run Tests on Old Version

Capture baseline results:

bash
# Python (pytest with JSON report)
git checkout old-version
pytest --json-report --json-report-file=old_results.json

# JavaScript (Jest with JSON report)
git checkout old-version
npm test -- --json --outputFile=old_results.json

# Run multiple times to check stability
pytest --json-report --json-report-file=old_results_1.json
pytest --json-report --json-report-file=old_results_2.json
# Compare to ensure deterministic

Verify baseline stability:

  • All tests should pass (or document known failures)
  • Results should be consistent across runs
  • No flaky tests
3. Run Tests on New Version

Capture new results:

bash
# Python
git checkout new-version
pytest --json-report --json-report-file=new_results.json

# JavaScript
git checkout new-version
npm test -- --json --outputFile=new_results.json

Note any immediate failures:

  • Tests that now fail
  • New errors or exceptions
  • Changed behavior
4. Compare Results

Use comparison script:

bash
python scripts/compare_results.py old_results.json new_results.json

# With custom tolerance for floats
python scripts/compare_results.py old_results.json new_results.json --tolerance 0.001

# Save detailed report
python scripts/compare_results.py old_results.json new_results.json --output regression_report.json

Script detects:

  • 🔴 Critical: Tests that passed now fail, missing tests
  • 🟠 High: Different outputs for same inputs
  • 🟡 Medium: Different exception types
  • 🔵 Low: Changed error messages
  • ✅ Improvements: Tests that now pass, bug fixes
5. Analyze Regressions

For each regression, determine:

Is it a true regression?

  • Unintended behavior change
  • Bug introduced
  • Performance degradation
  • Breaking change

Or is it expected?

  • Intentional behavior change
  • Bug fix that changes output
  • Improved error handling
  • Refactoring with equivalent behavior

Review strategies in detection_strategies.md.

6. Investigate Root Causes

For critical regressions:

bash
# Find commits that caused regression
git bisect start
git bisect bad new-version
git bisect good old-version
# Test each commit
git bisect run pytest path/to/failing_test.py

For output differences:

  • Compare function inputs/outputs
  • Check for changed algorithms
  • Verify data transformations
  • Review calculation logic

For exception changes:

  • Check error handling code
  • Verify exception types
  • Review validation logic
7. Document Findings

Create regression report:

REGRESSION ANALYSIS REPORT
==========================

Version Comparison: v1.0.0 → v1.1.0
Date: 2024-01-15
Tests Run: 156

SUMMARY
-------
Critical Regressions: 2
High Severity: 5
Medium Severity: 3
Low Severity: 8
Improvements: 4
Unchanged: 134

CRITICAL REGRESSIONS
--------------------
1. test_user_authentication
   - Status: PASS → FAIL
   - Error: KeyError: 'user_id'
   - Root Cause: Removed field from response
   - Action: Restore field or update API contract

2. test_payment_processing
   - Status: PASS → FAIL
   - Error: AssertionError: expected 100.00, got 100.01
   - Root Cause: Rounding change in calculation
   - Action: Fix rounding logic

HIGH SEVERITY REGRESSIONS
--------------------------
1. test_data_export
   - Output changed: CSV format → JSON format
   - Impact: Breaking change for consumers
   - Action: Maintain backward compatibility

[... continue for all regressions ...]

EXPECTED CHANGES
----------------
1. test_error_messages
   - Error messages now include more context
   - Intentional improvement
   - Action: Update baseline

RECOMMENDATIONS
---------------
1. Fix critical regressions before release
2. Review high severity changes with team
3. Document breaking changes in changelog
4. Update tests for intentional changes
8. Fix or Accept Changes

Fix true regressions:

bash
# Fix the code
git checkout new-version
# Make fixes
git commit -m "Fix: regression in user authentication"

# Re-run tests
pytest --json-report --json-report-file=fixed_results.json

# Verify fix
python scripts/compare_results.py old_results.json fixed_results.json

Accept intentional changes:

bash
# Update baseline
cp new_results.json baseline_results.json

# Document in changelog
echo "- Changed: CSV export now returns JSON" >> CHANGELOG.md

Quick Reference

Regression Types

Output Regressions:

  • Function returns different values
  • Data format changes
  • Calculation differences

Exception Regressions:

  • New exceptions raised
  • Different exception types
  • Changed error messages

State Regressions:

  • Different database state
  • Different files created
  • Different side effects

Performance Regressions:

  • Slower execution
  • Higher memory usage
  • More API calls
Show full SKILL.md (165 more words)Show less
Severity Levels

Critical (block release):

  • Test passed → failed
  • Data corruption
  • Security issues
  • Crashes

High (fix before release):

  • Wrong outputs
  • Breaking API changes
  • Major performance degradation (>2x)

Medium (review and decide):

  • Minor output changes
  • Moderate performance degradation (50-100%)
  • Changed error messages

Low (document):

  • Cosmetic changes
  • Minor performance changes (<50%)
  • Log message changes
Comparison Strategies

Exact comparison:

python
old_output == new_output

Approximate comparison (floats):

python
abs(old_output - new_output) < tolerance

Structural comparison (ignore fields):

python
# Ignore timestamps, IDs
compare_ignoring_fields(old, new, ['timestamp', 'id'])

Semantic comparison (order-independent):

python
# Compare as sets
set(old_list) == set(new_list)

Helper Script

The compare_results.py script automates comparison:

bash
# Basic comparison
python scripts/compare_results.py old_results.json new_results.json

# Custom float tolerance
python scripts/compare_results.py old_results.json new_results.json --tolerance 0.001

# Save detailed report
python scripts/compare_results.py old_results.json new_results.json --output report.json

Supported formats:

  • pytest JSON report
  • Jest JSON report
  • Generic JSON format

Output includes:

  • Categorized regressions by severity
  • Specific test failures
  • Output diffs
  • Exception changes
  • Improvements

Best Practices

Ensure deterministic tests:

  • Fix random seeds
  • Mock current time
  • Mock external APIs
  • Sort non-deterministic outputs

Run multiple times:

  • Verify baseline stability
  • Catch flaky tests
  • Ensure reproducibility

Isolate changes:

  • Test one change at a time
  • Use git bisect for root cause
  • Compare specific commits

Document expectations:

  • Maintain changelog
  • Note intentional changes
  • Update test baselines

Automate checks:

  • Run in CI/CD pipeline
  • Block on critical regressions
  • Generate reports automatically

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/regression-consistency-checker of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • references/detection_strategies.md
  • scripts/compare_results.py

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Regression Consistency Checker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regression Consistency Checker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regression Consistency Checker this skillArabelaTso/Skills-4-SE253—~2.1kAutomated safety check: PassApache-2.0
Code Review Graph Navigatorhandsontable/handsontable22k—~939Automated safety check: PassCustom licence
Analyze Projectlllllllama/RigorPilot-Skills4971 repos~519Automated safety check: PassMIT
Andrej Karpathy Skillduolahypercho/andrej-karpathy-skills249—~793Automated safety check: PassMIT
Verdaccio Change Implementationverdaccio/verdaccio18k—~1.3kAutomated safety check: PassMIT
RoamCranot/roam-code517—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Code Review Graph Navigator

    handsontable/handsontable

    Queries a pre-built, Tree-sitter-based code graph of the whole monorepo instead of grepping call chains, for exploring, debugging, refactoring or reviewing code.

    22k GitHub stars~939 tokensUpdated today
    DevelopmentAuto-check passed
  • Analyze Project

    lllllllama/RigorPilot-Skills

    Rigor Analyze / Rigor Audit read-only skill for deep learning research repositories.

    497 GitHub starsUsed in 1 repo~519 tokens
    DevelopmentAuto-check passed
  • Andrej Karpathy Skill

    duolahypercho/andrej-karpathy-skills

    Apply Andrej Karpathy-inspired coding-agent guidelines in Codex.

    249 GitHub stars~793 tokensUpdated 4 mo ago
    DevelopmentAuto-check passed
  • A workflow for implementing a Verdaccio bug fix, feature or refactor: pick the release lines, check existing options, edit the owning layer, test and add a changeset.

    18k GitHub stars~1.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Roam

    Cranot/roam-code

    Codebase comprehension via roam-code CLI. An agent skill from Cranot/roam-code.

    517 GitHub stars~2.4k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Odoo Workflow

    unclecatvn/agent-skills

    Mandatory pre-code gate and definition-of-done for ANY Odoo change (add field, override method, inherit view/xpath, OWL/JS patch, wizard, cron, controller, report, security, migration, bug fix…

    143 GitHub stars~4.7k tokensUpdated 15 days ago
    DevelopmentAuto-check passed

More from ArabelaTso/Skills-4-SE

All 170 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Regression Consistency Checker

What does Regression Consistency Checker do?

Checks whether a new version of a repository preserves the behavior observed by tests on the old version. Regression Consistency Checker is an agent skill from ArabelaTso/Skills-4-SE. Checks whether a new version of a repository preserves the behavior observed by tests on the old version.

When should I use Regression Consistency Checker?

Regression Consistency Checker fits situations like: comparing two versions of code to detect regressions; verify refactoring safety; validate bug fixes dont break existing functionality; ensure backward compatibility.

How do I install Regression Consistency Checker in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill regression-consistency-checker -a claude-code`. Or copy the skill folder (skills/regression-consistency-checker in ArabelaTso/Skills-4-SE) into .claude/skills/regression-consistency-checker in your project. Claude Code loads it when a task matches its description.

How do I install Regression Consistency Checker in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill regression-consistency-checker -a codex`. Or copy the skill folder (skills/regression-consistency-checker in ArabelaTso/Skills-4-SE) into .agents/skills/regression-consistency-checker in your project. Codex loads it when a task matches its description.

Can I use Regression Consistency Checker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill regression-consistency-checker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regression-consistency-checker, .gemini/skills/regression-consistency-checker, .github/skills/regression-consistency-checker and .opencode/skills/regression-consistency-checker in your project.

What does Regression Consistency Checker need to run?

Going by SKILL.md and its folder, Regression Consistency Checker needs Python for the scripts in its folder and the command-line tools its instructions call (git, python, pytest and npm). Our summary lists: Python 3.

Does Regression Consistency Checker access the network?

SKILL.md contains no URLs. Its commands use git and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Regression Consistency Checker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Regression Consistency Checker use?

Regression Consistency Checker is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regression Consistency Checker use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Regression Consistency Checker?

Skills that share tags, products or a category with Regression Consistency Checker: Code Review Graph Navigator (handsontable/handsontable, 22k stars), Analyze Project (lllllllama/RigorPilot-Skills, 497 stars), Andrej Karpathy Skill (duolahypercho/andrej-karpathy-skills, 249 stars) and Verdaccio Change Implementation (verdaccio/verdaccio, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regression Consistency Checker?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 170 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.