Agent skill

Prompt Regression Suite

by mohitagw15856 in mohitagw15856/pm-claude-skills

Design a regression test suite that catches an LLM feature getting worse when the prompt, model, or context changes.

MITAuto-check passedTesting & QA

Install Prompt Regression Suite

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill prompt-regression-suite -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills prompt-regression-suite --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/prompt-regression-suite .claude/skills/prompt-regression-suite && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-regression-suite
GitHub stars
1.4k
Token cost
~1.3k tokens
SKILL.md length
681 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Design a regression test suite that catches an LLM feature getting worse when the prompt, model, or context changes.

  • Works in 3 steps: Exact / structural — JSON parses,… → Property checks — output… → LLM-as-judge with a rubric — only where…
  • Asked to stop prompt changes breaking production
  • SKILL.md covers What This Skill Produces, Required Inputs, Building the Golden Set and Scoring Per Case, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Prompt Regression Suite is an agent skill from mohitagw15856/pm-claude-skills. Design a regression test suite that catches an LLM feature getting worse when the prompt, model, or context changes. Use when asked to stop prompt changes breaking production, set up golden tests or CI gates for an LLM feature, or test a model/prompt upgrade before shipping it. Produces a golden case set, per-case pass criteria, CI gate thresholds, and a triage protocol for failures. For designing first-time evaluation of a new feature use ai-eval-plan instead.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to stop prompt changes breaking production
  • Set up golden tests
  • CI gates for an LLM feature
  • Test a model/prompt upgrade before shipping it

Example prompts

  • “/prompt-regression-suite”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Exact / structural — JSON parses, required fields present, enum values legal. Free and deterministic; use wherever the contract is…
  2. Property checks — output contains/never-contains X, length bounds, citation count. Deterministic proxies for quality.
  3. LLM-as-judge with a rubric — only where judgement is unavoidable. Pin the judge model + rubric version, score against the baseline output…

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Regression Suite loads about 1.3k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 681 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 681 words, ~1,315 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-regression-suite/SKILL.md (or your agent's skills folder).
name
prompt-regression-suite
description
Design a regression test suite that catches an LLM feature getting worse when the prompt, model, or context changes. Use when asked to stop prompt changes breaking production, set up golden tests or CI gates for an LLM feature, or test a model/prompt upgrade before shipping it. Produces a golden case set, per-case pass criteria, CI gate thresholds, and a triage protocol for failures. For designing first-time evaluation of a new feature use ai-eval-plan instead.

Prompt Regression Suite Skill

Every prompt tweak, model upgrade, and context change is a deploy. This skill designs the suite that runs on each one and answers a single question: did anything that used to work stop working?

What This Skill Produces

  • A golden case set: curated inputs with per-case pass criteria
  • Scoring methods per case class (exact, rubric-judge, property checks)
  • CI gate thresholds — what blocks a merge vs. what warns
  • A failure triage protocol — flaky vs. regressed vs. golden-set-wrong

Required Inputs

Ask for (if not already provided):

  • The feature and its contract — what the LLM step receives and must produce
  • What has broken before (or nearly) — past incidents seed the best cases
  • Real traffic examples — 10-20 representative inputs, including ugly ones
  • What triggers a run — prompt edits, model bumps, retrieval changes, all of the above?

Building the Golden Set

Compose the set from four deliberate classes — not a random sample:

ClassPurposeShare
Core pathsThe 5-10 inputs that represent most real traffic~40%
Past failuresEvery input that caused a bug, complaint, or incident — permanently~25%
Edge & adversarialEmpty/huge inputs, wrong language, injection attempts, off-topic~25%
CanariesCases pinned to behaviours you never want to change (refusals, format, tone)~10%

Keep it small enough to run on every change (30-80 cases beats 500 nobody runs). Version it in git next to the prompt.

Scoring Per Case

Choose the cheapest check that catches the regression:

  1. Exact / structural — JSON parses, required fields present, enum values legal. Free and deterministic; use wherever the contract is structural.
  2. Property checks — output contains/never-contains X, length bounds, citation count. Deterministic proxies for quality.
  3. LLM-as-judge with a rubric — only where judgement is unavoidable. Pin the judge model + rubric version, score against the baseline output, and spot-check judge agreement with a human on ~20 cases before trusting it.

Every case records: input, pass criteria, scoring method, and the baseline output at the time it was added.

CI Gates

  • Block the merge: any past-failure or canary case fails; structural pass rate < 100%; overall pass rate drops more than [X]% vs. baseline.
  • Warn, don't block: judge-scored quality drifts within tolerance; latency/cost moves past its soft budget (pair with llm-cost-latency-budget).
  • Every run logs model ID, prompt version, and per-case results — regressions must be diffable to the exact change.
Show full SKILL.md (300 more words)Show less

Failure Triage Protocol

When a case fails, classify before "fixing":

  1. Flaky — re-run N times; if intermittent, tighten the prompt/temperature or the check, don't ignore it.
  2. Genuine regression — the change made it worse: revert or fix the change.
  3. Golden set wrong — the new behaviour is actually better: update the case via review, never silently, and record why the expectation changed.

Output Format

Prompt Regression Suite: [feature]

Trigger: runs on [prompt edit / model bump / retrieval change] via [CI job].

Golden set ([n] cases):

#ClassInput (summary)Pass criteriaMethod

Gates: merge blocks when [conditions]. Warnings on [conditions].

Triage: [the three-way protocol, with who owns updates to the golden set]

Maintenance: every production incident adds a case within [period]; the set is reviewed for dead cases each [quarter].

Quality Checks

  • Every past production failure appears as a permanent case
  • Canary cases cover the behaviours that must never change (refusals, format, safety)
  • No case relies on an LLM judge where a structural or property check would do
  • Gate thresholds are numbers, not "significant degradation"
  • The suite is fast and cheap enough that it actually runs on every change — state its runtime and cost

Anti-Patterns

  • Do not test only happy paths — the suite exists for the inputs that hurt you
  • Do not let anyone update golden expectations in the same PR that broke them, without review
  • Do not use an unpinned judge model — a judge that upgrades itself moves your baseline silently
  • Do not treat pass-rate-vs-baseline as the only gate — one dead canary matters more than 2% aggregate drift
  • Do not grow the set unboundedly — a suite too slow to run on every change protects nothing

Example Trigger Phrases

  • "Stop prompt changes breaking production."
  • "Set up golden tests or CI gates for an LLM feature."
  • "Test a model/prompt upgrade before shipping it."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/prompt-regression-suite of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Prompt Regression Suite next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Regression Suite compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Regression Suite this skillmohitagw15856/pm-claude-skills1.4k—~1.3kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Diagnosing Bugsfossasia/eventyay-interpretation1.6k32 repos~2.1kAutomated safety check: PassApache-2.0
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Diagnosing Bugs

    fossasia/eventyay-interpretation

    Diagnosis loop for hard bugs and performance regressions. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 32 repos~2.1k tokens
    Testing & QAAuto-check passed
  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    381 GitHub starsUsed in 9 repos~2.9k tokens
    Testing & QAAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Prompt Regression Suite

What does Prompt Regression Suite do?

Design a regression test suite that catches an LLM feature getting worse when the prompt, model, or context changes. Prompt Regression Suite is an agent skill from mohitagw15856/pm-claude-skills. Design a regression test suite that catches an LLM feature getting worse when the prompt, model, or context changes.

When should I use Prompt Regression Suite?

Prompt Regression Suite fits situations like: asked to stop prompt changes breaking production; set up golden tests; CI gates for an LLM feature; test a model/prompt upgrade before shipping it.

How do I install Prompt Regression Suite in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill prompt-regression-suite -a claude-code`. Or copy the skill folder (skills/prompt-regression-suite in mohitagw15856/pm-claude-skills) into .claude/skills/prompt-regression-suite in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Regression Suite in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill prompt-regression-suite -a codex`. Or copy the skill folder (skills/prompt-regression-suite in mohitagw15856/pm-claude-skills) into .agents/skills/prompt-regression-suite in your project. Codex loads it when a task matches its description.

Can I use Prompt Regression Suite in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill prompt-regression-suite -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-regression-suite, .gemini/skills/prompt-regression-suite, .github/skills/prompt-regression-suite and .opencode/skills/prompt-regression-suite in your project.

What does Prompt Regression Suite need to run?

SKILL.md names no scripts, command-line tools or credentials: Prompt Regression Suite is instructions for the agent only.

Does Prompt Regression Suite access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prompt Regression Suite safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prompt Regression Suite use?

Prompt Regression Suite is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Regression Suite use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prompt Regression Suite?

Skills that share tags, products or a category with Prompt Regression Suite: Web Application Testing (anthropics/skills, 180k stars), Diagnosing Bugs (fossasia/eventyay-interpretation, 1.6k stars), TDD (pietheinstrengholt/rssmonster, 564 stars) and TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Regression Suite?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.