Agent skill

AI Regression Testing

by affaan-m in affaan-m/ECC

Regression testing strategies for AI-assisted development. An agent skill from affaan-m/ECC.

MITAuto-check passedTesting & QA

Install AI Regression Testing

skills CLI
$ npx skills add affaan-m/ECC --skill ai-regression-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC ai-regression-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-regression-testing .claude/skills/ai-regression-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-regression-testing
GitHub stars
275k
Used in
5 other repos
Token cost
~2.9k tokens
SKILL.md length
475 words
Files
1
Skills in repo
645
Repo updated
First seen
Licence
MIT

At a glance

Regression testing strategies for AI-assisted development. An agent skill from affaan-m/ECC.

  • Works in 4 steps: AI tends to make the same category of… → Bugs cluster in complex areas (auth,… → Once tested, that exact regression… → …
  • Adding regression coverage to AI-assisted code
  • SKILL.md covers When to Activate, The Core Problem, Sandbox-Mode API Testing and Integrating Tests into…, plus 4 more sections
  • Needs NEXT_PUBLIC_SUPABASE_ANON_KEY

What it does

AI Regression Testing is an agent skill from affaan-m/ECC. Regression testing strategies for AI-assisted development. Sandbox-mode API testing without database dependencies, automated bug-check workflows, and patterns to catch AI blind spots where the same model writes and reviews code. Use when adding regression coverage to AI-assisted code, or when the same model both wrote and reviewed a change.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports and API testing. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Adding regression coverage to AI-assisted code
  • The same model both wrote and reviewed a change

Example prompts

  • “/ai-regression-testing”

Requirements

  • A credential in NEXT_PUBLIC_SUPABASE_ANON_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. AI tends to make the same category of mistake repeatedly
  2. Bugs cluster in complex areas (auth, multi-path logic, state management)
  3. Once tested, that exact regression cannot happen again
  4. Test count grows organically with bug fixes — no wasted effort

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NEXT_PUBLIC_SUPABASE_ANON_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Regression Testing loads about 2.9k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 475 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 475 words, ~2,878 tokens.

Download SKILL.mdSave it as .claude/skills/ai-regression-testing/SKILL.md (or your agent's skills folder).
name
ai-regression-testing
description
Regression testing strategies for AI-assisted development. Sandbox-mode API testing without database dependencies, automated bug-check workflows, and patterns to catch AI blind spots where the same model writes and reviews code. Use when adding regression coverage to AI-assisted code, or when the same model both wrote and reviewed a change.
metadata.origin
ECC

AI Regression Testing

Testing patterns specifically designed for AI-assisted development, where the same model writes code and reviews it — creating systematic blind spots that only automated tests can catch.

When to Activate

  • AI agent (Claude Code, Cursor, Codex) has modified API routes or backend logic
  • A bug was found and fixed — need to prevent re-introduction
  • Project has a sandbox/mock mode that can be leveraged for DB-free testing
  • Running /bug-check or similar review commands after code changes
  • Multiple code paths exist (sandbox vs production, feature flags, etc.)

The Core Problem

When an AI writes code and then reviews its own work, it carries the same assumptions into both steps. This creates a predictable failure pattern:

AI writes fix → AI reviews fix → AI says "looks correct" → Bug still exists

Real-world example (observed in production):

Fix 1: Added notification_settings to API response
  → Forgot to add it to the SELECT query
  → AI reviewed and missed it (same blind spot)

Fix 2: Added it to SELECT query
  → TypeScript build error (column not in generated types)
  → AI reviewed Fix 1 but didn't catch the SELECT issue

Fix 3: Changed to SELECT *
  → Fixed production path, forgot sandbox path
  → AI reviewed and missed it AGAIN (4th occurrence)

Fix 4: Test caught it instantly on first run PASS:

The pattern: sandbox/production path inconsistency is the #1 AI-introduced regression.

Sandbox-Mode API Testing

Most projects with AI-friendly architecture have a sandbox/mock mode. This is the key to fast, DB-free API testing.

Setup (Vitest + Next.js App Router)
typescript
// vitest.config.ts
import { defineConfig } from "vitest/config";
import path from "path";

export default defineConfig({
  test: {
    environment: "node",
    globals: true,
    include: ["__tests__/**/*.test.ts"],
    setupFiles: ["__tests__/setup.ts"],
  },
  resolve: {
    alias: {
      "@": path.resolve(__dirname, "."),
    },
  },
});
typescript
// __tests__/setup.ts
// Force sandbox mode — no database needed
process.env.SANDBOX_MODE = "true";
process.env.NEXT_PUBLIC_SUPABASE_URL = "";
process.env.NEXT_PUBLIC_SUPABASE_ANON_KEY = "";
Test Helper for Next.js API Routes
typescript
// __tests__/helpers.ts
import { NextRequest } from "next/server";

export function createTestRequest(
  url: string,
  options?: {
    method?: string;
    body?: Record<string, unknown>;
    headers?: Record<string, string>;
    sandboxUserId?: string;
  },
): NextRequest {
  const { method = "GET", body, headers = {}, sandboxUserId } = options || {};
  const fullUrl = url.startsWith("http") ? url : `http://localhost:3000${url}`;
  const reqHeaders: Record<string, string> = { ...headers };

  if (sandboxUserId) {
    reqHeaders["x-sandbox-user-id"] = sandboxUserId;
  }

  const init: { method: string; headers: Record<string, string>; body?: string } = {
    method,
    headers: reqHeaders,
  };

  if (body) {
    init.body = JSON.stringify(body);
    reqHeaders["content-type"] = "application/json";
  }

  return new NextRequest(fullUrl, init);
}

export async function parseResponse(response: Response) {
  const json = await response.json();
  return { status: response.status, json };
}
Writing Regression Tests

The key principle: write tests for bugs that were found, not for code that works.

typescript
// __tests__/api/user/profile.test.ts
import { describe, it, expect } from "vitest";
import { createTestRequest, parseResponse } from "../../helpers";
import { GET, PATCH } from "@/app/api/user/profile/route";

// Define the contract — what fields MUST be in the response
const REQUIRED_FIELDS = [
  "id",
  "email",
  "full_name",
  "phone",
  "role",
  "created_at",
  "avatar_url",
  "notification_settings",  // ← Added after bug found it missing
];

describe("GET /api/user/profile", () => {
  it("returns all required fields", async () => {
    const req = createTestRequest("/api/user/profile");
    const res = await GET(req);
    const { status, json } = await parseResponse(res);

    expect(status).toBe(200);
    for (const field of REQUIRED_FIELDS) {
      expect(json.data).toHaveProperty(field);
    }
  });

  // Regression test — this exact bug was introduced by AI 4 times
  it("notification_settings is not undefined (BUG-R1 regression)", async () => {
    const req = createTestRequest("/api/user/profile");
    const res = await GET(req);
    const { json } = await parseResponse(res);

    expect("notification_settings" in json.data).toBe(true);
    const ns = json.data.notification_settings;
    expect(ns === null || typeof ns === "object").toBe(true);
  });
});
Testing Sandbox/Production Parity

The most common AI regression: fixing production path but forgetting sandbox path (or vice versa).

typescript
// Test that sandbox responses match the expected contract
describe("GET /api/user/messages (conversation list)", () => {
  it("includes partner_name in sandbox mode", async () => {
    const req = createTestRequest("/api/user/messages", {
      sandboxUserId: "user-001",
    });
    const res = await GET(req);
    const { json } = await parseResponse(res);

    // This caught a bug where partner_name was added
    // to production path but not sandbox path
    if (json.data.length > 0) {
      for (const conv of json.data) {
        expect("partner_name" in conv).toBe(true);
      }
    }
  });
});

Integrating Tests into Bug-Check Workflow

Custom Command Definition
markdown
<!-- .claude/commands/bug-check.md -->
# Bug Check

## Step 1: Automated Tests (mandatory, cannot skip)

Run these commands FIRST before any code review:

    npm run test       # Vitest test suite
    npm run build      # TypeScript type check + build

- If tests fail → report as highest priority bug
- If build fails → report type errors as highest priority
- Only proceed to Step 2 if both pass

## Step 2: Code Review (AI review)

1. Sandbox / production path consistency
2. API response shape matches frontend expectations
3. SELECT clause completeness
4. Error handling with rollback
5. Optimistic update race conditions

## Step 3: For each bug fixed, propose a regression test
The Workflow
User: "バグチェックして" (or "/bug-check")
  │
  ├─ Step 1: npm run test
  │   ├─ FAIL → Bug found mechanically (no AI judgment needed)
  │   └─ PASS → Continue
  │
  ├─ Step 2: npm run build
  │   ├─ FAIL → Type error found mechanically
  │   └─ PASS → Continue
  │
  ├─ Step 3: AI code review (with known blind spots in mind)
  │   └─ Findings reported
  │
  └─ Step 4: For each fix, write a regression test
      └─ Next bug-check catches if fix breaks

Common AI Regression Patterns

Pattern 1: Sandbox/Production Path Mismatch

Frequency: Most common (observed in 3 out of 4 regressions)

typescript
// FAIL: AI adds field to production path only
if (isSandboxMode()) {
  return { data: { id, email, name } };  // Missing new field
}
// Production path
return { data: { id, email, name, notification_settings } };

// PASS: Both paths must return the same shape
if (isSandboxMode()) {
  return { data: { id, email, name, notification_settings: null } };
}
return { data: { id, email, name, notification_settings } };

Test to catch it:

typescript
it("sandbox and production return same fields", async () => {
  // In test env, sandbox mode is forced ON
  const res = await GET(createTestRequest("/api/user/profile"));
  const { json } = await parseResponse(res);

  for (const field of REQUIRED_FIELDS) {
    expect(json.data).toHaveProperty(field);
  }
});
Pattern 2: SELECT Clause Omission

Frequency: Common with Supabase/Prisma when adding new columns

typescript
// FAIL: New column added to response but not to SELECT
const { data } = await supabase
  .from("users")
  .select("id, email, name")  // notification_settings not here
  .single();

return { data: { ...data, notification_settings: data.notification_settings } };
// → notification_settings is always undefined

// PASS: Use SELECT * or explicitly include new columns
const { data } = await supabase
  .from("users")
  .select("*")
  .single();
Pattern 3: Error State Leakage

Frequency: Moderate — when adding error handling to existing components

typescript
// FAIL: Error state set but old data not cleared
catch (err) {
  setError("Failed to load");
  // reservations still shows data from previous tab!
}

// PASS: Clear related state on error
catch (err) {
  setReservations([]);  // Clear stale data
  setError("Failed to load");
}
Pattern 4: Optimistic Update Without Proper Rollback
typescript
// FAIL: No rollback on failure
const handleRemove = async (id: string) => {
  setItems(prev => prev.filter(i => i.id !== id));
  await fetch(`/api/items/${id}`, { method: "DELETE" });
  // If API fails, item is gone from UI but still in DB
};

// PASS: Capture previous state and rollback on failure
const handleRemove = async (id: string) => {
  const prevItems = [...items];
  setItems(prev => prev.filter(i => i.id !== id));
  try {
    const res = await fetch(`/api/items/${id}`, { method: "DELETE" });
    if (!res.ok) throw new Error("API error");
  } catch {
    setItems(prevItems);  // Rollback
    alert("削除に失敗しました");
  }
};
Show full SKILL.md (211 more words)Show less

Strategy: Test Where Bugs Were Found

Don't aim for 100% coverage. Instead:

Bug found in /api/user/profile     → Write test for profile API
Bug found in /api/user/messages    → Write test for messages API
Bug found in /api/user/favorites   → Write test for favorites API
No bug in /api/user/notifications  → Don't write test (yet)

Why this works with AI development:

  1. AI tends to make the same category of mistake repeatedly
  2. Bugs cluster in complex areas (auth, multi-path logic, state management)
  3. Once tested, that exact regression cannot happen again
  4. Test count grows organically with bug fixes — no wasted effort

Quick Reference

AI Regression PatternTest StrategyPriority
Sandbox/production mismatchAssert same response shape in sandbox modeHigh
SELECT clause omissionAssert all required fields in responseHigh
Error state leakageAssert state cleanup on errorMedium
Missing rollbackAssert state restored on API failureMedium
Type cast masking nullAssert field is not undefinedMedium

DO / DON'T

DO:

  • Write tests immediately after finding a bug (before fixing it if possible)
  • Test the API response shape, not the implementation
  • Run tests as the first step of every bug-check
  • Keep tests fast (< 1 second total with sandbox mode)
  • Name tests after the bug they prevent (e.g., "BUG-R1 regression")

DON'T:

  • Write tests for code that has never had a bug
  • Trust AI self-review as a substitute for automated tests
  • Skip sandbox path testing because "it's just mock data"
  • Write integration tests when unit tests suffice
  • Aim for coverage percentage — aim for regression prevention

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ai-regression-testing of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Used in 5 other repositories

We found 10 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

AI Regression Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Regression Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Regression Testing this skillaffaan-m/ECC275k5 repos~2.9kAutomated safety check: PassMIT
Diff-Driven QASkyvern-AI/skyvern23k—~4.7kAutomated safety check: WarnAGPL-3.0
QA Test and Fixgarrytan/gstack136k—~13kAutomated safety check: NotesMIT
QA Report Onlygarrytan/gstack136k—~11kAutomated safety check: NotesMIT
Bugbashav/mi102—~1.4kAutomated safety check: PassNone
Use Yaakmountain-loop/yaak19k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Diff-Driven QA

    Skyvern-AI/skyvern

    Reads your git diff, decides whether the change needs browser QA, API checks or repo tests, runs that validation and reports pass or fail with evidence.

    23k GitHub stars~4.7k tokensUpdated today
    Testing & QAAuto-check: warnings
  • QA Test and Fix

    garrytan/gstack

    Tests a site or service for bugs, fixes what it finds with one atomic commit per fix, and reports fix evidence and a ship-readiness summary at one of three depths.

    136k GitHub stars~13k tokensUpdated today
    Testing & QAAuto-check: notes
  • QA Report Only

    garrytan/gstack

    Tests a browser app, API, CLI, job, worker or webhook and writes a structured bug report with repro steps, without changing any code.

    136k GitHub stars~11k tokensUpdated today
    Testing & QAAuto-check: notes
  • Bugbash

    av/mi

    Systematically explores a CLI, API, backend or library to find bugs and edge cases, then writes a report with reproduction evidence for each issue.

    102 GitHub stars~1.4k tokensUpdated 12 days ago
    Testing & QAAuto-check passed
  • Use Yaak

    mountain-loop/yaak

    A skill your agent uses when the user mentions Yaak, a Yaak workspace, or the yaak command, or asks to call, hit, or smoke test HTTP/REST endpoints, save or organize API requests for reuse or manual…

    19k GitHub stars~1.9k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Reproduce Chat States

    different-ai/openwork

    Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.

    24k GitHub stars~673 tokensUpdated today
    Testing & QAAuto-check passed

More from affaan-m/ECC

All 645 skills in this repo
  • Videodb

    affaan-m/ECC

    Ingest, index, search, edit, and monitor video and audio with the VideoDB Python SDK — upload from files, URLs, or RTSP feeds, build spoken and scene indexes with timestamped search and playable…

    275k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    275k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    275k GitHub stars~2.9k tokensUpdated 3 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    275k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    275k GitHub stars~3.5k tokensUpdated 3 days ago
    Auto-check passed
  • Adds one optional external Codex critique that tries to break a council's decision draft, sent to OpenAI only after you consent.

    275k GitHub stars~1.5k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about AI Regression Testing

What does AI Regression Testing do?

Regression testing strategies for AI-assisted development. An agent skill from affaan-m/ECC. AI Regression Testing is an agent skill from affaan-m/ECC. Regression testing strategies for AI-assisted development.

When should I use AI Regression Testing?

AI Regression Testing fits situations like: adding regression coverage to AI-assisted code; the same model both wrote and reviewed a change.

How do I install AI Regression Testing in Claude Code?

Run `npx skills add affaan-m/ECC --skill ai-regression-testing -a claude-code`. Or copy the skill folder (skills/ai-regression-testing in affaan-m/ECC) into .claude/skills/ai-regression-testing in your project. Claude Code loads it when a task matches its description.

How do I install AI Regression Testing in Codex?

Run `npx skills add affaan-m/ECC --skill ai-regression-testing -a codex`. Or copy the skill folder (skills/ai-regression-testing in affaan-m/ECC) into .agents/skills/ai-regression-testing in your project. Codex loads it when a task matches its description.

Can I use AI Regression Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill ai-regression-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-regression-testing, .gemini/skills/ai-regression-testing, .github/skills/ai-regression-testing and .opencode/skills/ai-regression-testing in your project.

What does AI Regression Testing need to run?

Going by SKILL.md and its folder, AI Regression Testing needs credentials named NEXT_PUBLIC_SUPABASE_ANON_KEY. Our summary lists: A credential in NEXT_PUBLIC_SUPABASE_ANON_KEY.

Does AI Regression Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Regression Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Regression Testing use?

AI Regression Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Regression Testing use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Regression Testing?

Skills that share tags, products or a category with AI Regression Testing: Diff-Driven QA (Skyvern-AI/skyvern, 23k stars), QA Test and Fix (garrytan/gstack, 136k stars), QA Report Only (garrytan/gstack, 136k stars) and Bugbash (av/mi, 102 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Regression Testing?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,023 GitHub stars. The repository holds 645 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.