Agent skill

AI Prompt Regression Testing

by sickn33 in sickn33/agentic-awesome-skills

Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.

MITAuto-check passedTesting & QA

Install AI Prompt Regression Testing

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill ai-prompt-regression-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills ai-prompt-regression-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-prompt-regression-testing .claude/skills/ai-prompt-regression-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-prompt-regression-testing
GitHub stars
47k
Used in
1 other repo
Token cost
~1.3k tokens
SKILL.md length
504 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.

  • Works in 3 steps: Define the parameters, thresholds, and… → Select appropriate boundary enforcement… → Export standardized artifacts (CSV…
  • Tasks that involve QA and bug reports
  • SKILL.md covers Overview, When to Use This Skill, How It Works and Field Reference, plus 9 more sections
  • Calls claude and gemini

What it does

AI Prompt Regression Testing is an agent skill from sickn33/agentic-awesome-skills. Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports. It works with SQL. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve QA and bug reports

Example prompts

  • “/ai-prompt-regression-testing”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Define the parameters, thresholds, and identity bindings required for the target operational register.
  2. Select appropriate boundary enforcement values from validated enum select sets.
  3. Export standardized artifacts (CSV table, SQL DDL, JSON Schema) to integrate into validation CI pipelines.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude
    • gemini

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Prompt Regression Testing loads about 1.3k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 504 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 504 words, ~1,340 tokens.

Download SKILL.mdSave it as .claude/skills/ai-prompt-regression-testing/SKILL.md (or your agent's skills folder).
name
ai-prompt-regression-testing
description
Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.
category
engineering
risk
safe
source
self
source_type
self
date_added
2026-10-01
author
Ranjeet2063
tags
ai, prompt-engineering, testing, evaluation, llm, benchmarking
source_repo
Ranjeet2063/agentic-awesome-skills

AI Prompt Regression Test Matrix

What it is: Tracks regression baselines, evaluation rubrics, and automated judge verdicts to prevent output degradation across prompt revisions.

Overview

Provides a standardized, auditable framework and data model for AI Prompt Regression Test Matrix operations across distributed engineering and decentralized application systems.

When to Use This Skill

  • When formalizing architectural contracts, security invariants, or operational limits for AI Prompt Regression Test Matrix.
  • When cross-functional review is required between protocol developers, smart contract auditors, and AI engineering agents.
  • When generating reproducible CSV, SQL DDL, JSON Schema, and Notion property registers for tracking compliance.

How It Works

  1. Define the parameters, thresholds, and identity bindings required for the target operational register.
  2. Select appropriate boundary enforcement values from validated enum select sets.
  3. Export standardized artifacts (CSV table, SQL DDL, JSON Schema) to integrate into validation CI pipelines.

Field Reference

#Field NameTypeSQL TypeJSON Schema TypeNotion Property TypeExample Value
1Prompt TestCase IDidSERIAL PRIMARY KEYintegerTextPTEST-001
2Prompt IdentifiertextVARCHAR(64)stringTextsoroban_code_refactor_v2
3Target LLM Model FamilyselectVARCHAR(64)stringSelectClaude 3.5 Sonnet
4Evaluation MetricselectVARCHAR(64)stringSelectAST Code Correctness
5Semantic Drift ThresholdnumberNUMERIC(5,2)numberNumber0.05
6Golden Baseline Match %numberNUMERIC(5,2)numberNumber98.50
7Judge Model EvaluatortextVARCHAR(64)stringTextGemini 1.5 Pro
8Zero-Shot Reasoning VerifiedselectVARCHAR(16)stringSelectYes
9Latency Bound SecondsnumberNUMERIC(6,2)numberNumber3.20
10Test Suite VerdictselectVARCHAR(32)stringSelectPassed
11Benchmarking DatedateDATEstring, format: dateDate2026-10-01

Select Options

Target LLM Model Family

Claude 3.5 Sonnet | GPT-4o | Gemini 1.5 Pro | DeepSeek Coder

Evaluation Metric

AST Code Correctness | Semantic Embedding Cosine | Exact Match | Rubric Scoring

Zero-Shot Reasoning Verified

Yes | No

Test Suite Verdict

Passed | Degraded | Failed Regression

Relations

  • Audit Reference -> links to the formal review documentation or test repository.
  • Target Architecture -> links to the deployed contract or autonomous agent runtime component.
Show full SKILL.md (203 more words)Show less

Examples

Prompt

How do I configure and track AI Prompt Regression Test Matrix for our production environment?

Recommended Next Step

Generate the unified field schema, SQL DDL migration, and JSON validation schema to register into your system catalog.

Workflow: Define criteria -> Run automated verification -> Record baseline -> Monitor invariants.

Best Practices

  • Enforce strict typing on numerical bounds and currency amounts; avoid unstructured free-text fields for critical states.
  • Re-run validation test suites on every state-altering commit or parameter change.
  • Keep example data synthetic and isolated from production cryptographic keys or private endpoints.

Limitations

  • Provides architectural specifications, data models, and verification schemas; does not execute direct transaction signing without authorized external tooling.
  • Requires network connectivity and valid RPC credentials when querying on-chain states.

Security & Safety Notes

  • All parameters declare risk: safe. No unauthorized state modification or privileged credential access is performed.
  • Use synthetic dummy keys and mock addresses in test suites and local verification scripts.

Common Pitfalls

  • Problem: Mismatched decimal precision between contract runtime and database register. Solution: Always verify decimals using the explicit field mapping in this reference.
  • Problem: Missing authorization checks prior to state update. Solution: Cross-validate against the Security Audit register before deployment.
  • @ai-agent-tool-routing - covers tool schema registration and retry policy.
  • @cross-chain-relayer-audit - covers message hashes, nonces and quorum proofs.
  • @smart-contract-formal-verification - verifies state invariants mathematically.

Reusable Prompt

I want to establish a verified AI Prompt Regression Test Matrix register for our production protocol.
Guide me through the required field parameters and output the corresponding SQL DDL and JSON Schema.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ai-prompt-regression-testing of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

AI Prompt Regression Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Prompt Regression Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Prompt Regression Testing this skillsickn33/agentic-awesome-skills47k1 repos~1.3kAutomated safety check: PassMIT
Dogfood Exploratory QAvercel-labs/agent-browser44k8 repos~2.7kAutomated safety check: PassApache-2.0
DeerFlow Smoke Testbytedance/deer-flow84k—~2.5kAutomated safety check: NotesMIT
CodexBar Live QAsteipete/CodexBar22k—~1.2kAutomated safety check: PassMIT
Diagnose Playwright Failure as Product Bugappsmithorg/appsmith41k—~1.5kAutomated safety check: PassApache-2.0
Acceptance Evidence for Deliverieslobehub/lobehub83k—~9.7kAutomated safety check: PassApache-2.0

Similar skills

  • Dogfood Exploratory QA

    vercel-labs/agent-browser

    Official

    Explores a web app with the agent-browser CLI to find bugs and UX problems, then writes a report with screenshots, repro videos and step-by-step reproduction for each issue.

    44k GitHub starsUsed in 8 repos~2.7k tokens
    Testing & QAAuto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    84k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • CodexBar Live QA

    steipete/CodexBar

    Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

    22k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

    41k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check passed
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Codex Plugin QA

    code-yeongyu/oh-my-openagent

    Tests the omo Codex plugin in an isolated CODEX_HOME with a local mock model, proving hooks fired through app-server notifications without touching ~/.codex.

    70k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Categories

Questions about AI Prompt Regression Testing

What does AI Prompt Regression Testing do?

Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset. AI Prompt Regression Testing is an agent skill from sickn33/agentic-awesome-skills. Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.

When should I use AI Prompt Regression Testing?

AI Prompt Regression Testing fits situations like: tasks that involve QA and bug reports.

How do I install AI Prompt Regression Testing in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill ai-prompt-regression-testing -a claude-code`. Or copy the skill folder (skills/ai-prompt-regression-testing in sickn33/agentic-awesome-skills) into .claude/skills/ai-prompt-regression-testing in your project. Claude Code loads it when a task matches its description.

How do I install AI Prompt Regression Testing in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill ai-prompt-regression-testing -a codex`. Or copy the skill folder (skills/ai-prompt-regression-testing in sickn33/agentic-awesome-skills) into .agents/skills/ai-prompt-regression-testing in your project. Codex loads it when a task matches its description.

Can I use AI Prompt Regression Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill ai-prompt-regression-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-prompt-regression-testing, .gemini/skills/ai-prompt-regression-testing, .github/skills/ai-prompt-regression-testing and .opencode/skills/ai-prompt-regression-testing in your project.

What does AI Prompt Regression Testing need to run?

Going by SKILL.md and its folder, AI Prompt Regression Testing needs the command-line tools its instructions call (claude and gemini).

Does AI Prompt Regression Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Prompt Regression Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Prompt Regression Testing use?

AI Prompt Regression Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Prompt Regression Testing use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Prompt Regression Testing?

Skills that share tags, products or a category with AI Prompt Regression Testing: Dogfood Exploratory QA (vercel-labs/agent-browser, 44k stars), DeerFlow Smoke Test (bytedance/deer-flow, 84k stars), CodexBar Live QA (steipete/CodexBar, 22k stars) and Diagnose Playwright Failure as Product Bug (appsmithorg/appsmith, 41k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Prompt Regression Testing?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.