Agent skill

Prompt Regression

by agentscope-ai in agentscope-ai/OpenJudge

A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Prompt Regression

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge prompt-regression --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval_pipeline/06-prompt-regression .claude/skills/prompt-regression && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prompt-regression
GitHub stars
871
Token cost
~2.8k tokens
SKILL.md length
822 words
Files
2 (incl. scripts)
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

  • Works in 6 steps: Load and Analyze Prompts → Derive Comparison Dimensions → Select Graders → …
  • The user has changed a prompt (system prompt
  • SKILL.md covers When to Activate, Checklist, Fast path: run the bundled… and Step 1: Load and Analyze Prompts, plus 7 more sections
  • Runs Python scripts from its folder; calls python

What it does

Prompt Regression is an agent skill from agentscope-ai/OpenJudge. Use when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. Also use when the user mentions prompt A/B testing, prompt comparison, prompt optimization validation, "did my prompt change help," or prompt regression testing. Outputs per-dimension win rates with statistical significance using OpenJudge PairwiseAnalyzer.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/pairwise.py`).

It sits in AI & LLM Engineering, covering Prompt engineering, A/B testing and QA and bug reports. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user has changed a prompt (system prompt
  • Agent instruction
  • Etc.) and wants to know whether the candidate is better
  • Worse than the baseline

Example prompts

  • “did my prompt change help,”
  • “/prompt-regression”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Load and Analyze Prompts
  2. Derive Comparison Dimensions
  3. Select Graders
  4. Run Position-Debiased Comparison
  5. Compute Statistics
  6. Present Results

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prompt Regression loads about 2.8k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 822 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 822 words, ~2,811 tokens.

Download SKILL.mdSave it as .claude/skills/prompt-regression/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
prompt-regression
description
Use when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. Also use when the user mentions prompt A/B testing, prompt comparison, prompt optimization validation, "did my prompt change help," or prompt regression testing. Outputs per-dimension win rates with statistical significance using OpenJudge PairwiseAnalyzer.
<HARD-GATE>
NO conclusion about which prompt is better WITHOUT bootstrap 95% CI reported.
NO candidate declared "better" WITHOUT position-debiased (swap-aggregate) comparison.
NO comparison with fewer than 10 samples per axis — CI is too wide to be meaningful.
</HARD-GATE>

Prompt Regression

Compare two prompts head-to-head and determine, with statistical rigor, whether the candidate is better, worse, or tied on each evaluation dimension.

When to Activate

  • You changed the system prompt and want to verify it's actually better
  • You're iterating on RAG answer templates
  • You're optimizing agent step-by-step instructions
  • You want data to support a prompt change decision

Checklist

You MUST create a task for each item and complete them in order:

  1. Load and analyze prompts — diff the baseline vs candidate
  2. Derive comparison dimensions — from the prompt changes + task type
  3. Select graders per dimension — pairwise, judge, or rule
  4. Run position-debiased comparison — swap-aggregate to eliminate order bias
  5. Compute statistics — win rates + bootstrap 95% CI per dimension
  6. Present results — per-dimension verdict with confidence intervals

Fast path: run the bundled script

Don't hand-write the win-rate + bootstrap math (the swap-aggregation and CI are easy to get wrong). Run the bundled, tested script (scripts/pairwise.py, standard library only, no OpenJudge dependency):

bash
python scripts/pairwise.py --comparisons comparisons.jsonl --candidate candidate --baseline baseline

Each comparison row: {"id","model_a","model_b","score","dimension"?} where score >= 0.5 means model_a won. Emit two rows per query with A/B swapped to debias position. The script reports per-dimension candidate/baseline/tie rates, bootstrap 95% CI, and a verdict (BETTER / WORSE / TIED / INSUFFICIENT_EVIDENCE / INCONCLUSIVE; exit 0 only if better). --self-test to verify it.

Steps below explain how to derive dimensions and produce the comparisons (with OpenJudge or any judge); the inline snippets are the reference behind the script.

Step 1: Load and Analyze Prompts

Read the baseline and candidate prompts. Identify:

  • Task type: chatbot / RAG generation / code review / translation / summarization / agent instruction / other
  • What changed: added constraints, changed tone, new examples, different output format, expanded/shortened instructions
  • Intent of change: what problem was the user trying to fix?

Step 2: Derive Comparison Dimensions

Based on the task type and what changed, derive 3-5 comparison dimensions.

Dimension templates by task type

Chatbot / Conversational:

  • Answer relevance — does it address the user's question?
  • Tone appropriateness — does the tone match context?
  • Factual accuracy — no fabricated information
  • Conciseness — doesn't ramble or over-explain
  • Instruction following — obeys system prompt constraints

RAG Generation:

  • Faithfulness — grounded in retrieved documents
  • Citation accuracy — correctly references sources
  • Completeness — covers all aspects of the query
  • No hallucination — no claims beyond documents

Code Review / Generation:

  • Bug detection — finds real issues
  • False positive rate — doesn't flag correct code
  • Actionability — suggestions are specific and implementable
  • Code style — follows conventions

Agent Instructions:

  • Tool selection — picks the right tool
  • Step efficiency — minimal steps to goal
  • Error recovery — handles failures gracefully
  • Output format — follows specified structure

Each dimension gets:

  • An id (slug)
  • A one-sentence description
  • A grader type: pairwise or judge or rule

Step 3: Select Graders

Decision priority:

  1. Can a rule check this? → FunctionGrader or StringMatchGrader. Free, deterministic. Example: output length, keyword presence, JSON validity.
  2. Is there a reference answer? → pairwise against reference.
  3. Subjective quality, no reference? → pairwise A/B comparison.
  4. Single-output judgment needed? → judge (binary pass/fail per output).
Show full SKILL.md (306 more words)Show less

Step 4: Run Position-Debiased Comparison

Pairwise comparison with swap-aggregate

LLM judges have position bias — the first response shown wins 5-15% more often. Swap-aggregate eliminates this: run each comparison twice with swapped positions, keep only consistent wins:

python
from openjudge.graders.llm_grader import LLMGrader
from openjudge.graders.schema import GraderMode
from openjudge.runner.grading_runner import GradingRunner
from openjudge.analyzer.pairwise_analyzer import PairwiseAnalyzer

# Judge prompt for relevance comparison
relevance_judge = LLMGrader(
    model=model,
    name="relevance_compare",
    mode=GraderMode.POINTWISE,
    template="""
Compare Response A and Response B for the query below.
Which response better addresses the user's question?

Query: {query}
Response A: {response_a}
Response B: {response_b}

Score 1.0 if A is better, 0.0 if B is better, 0.5 if tied.
Respond in JSON: {{"score": <float>, "reason": "<explanation>"}}
""",
)

# Build pairwise dataset with position swap
dataset = []
for sample in test_samples:
    # Original order
    dataset.append({
        "query": sample["query"],
        "response_a": baseline_outputs[sample["id"]],
        "response_b": candidate_outputs[sample["id"]],
        "metadata": {"model_a": "baseline", "model_b": "candidate"},
    })
    # Swapped order — critical for debiasing
    dataset.append({
        "query": sample["query"],
        "response_a": candidate_outputs[sample["id"]],
        "response_b": baseline_outputs[sample["id"]],
        "metadata": {"model_a": "candidate", "model_b": "baseline"},
    })

runner = GradingRunner(
    grader_configs={"relevance": relevance_judge},
    max_concurrency=8,
)
results = await runner.arun(dataset)

# Analyze with PairwiseAnalyzer
analyzer = PairwiseAnalyzer(model_names=["baseline", "candidate"])
analysis = analyzer.analyze(dataset, results["relevance"])

print(f"Win rates: {analysis.win_rates}")
# → {'baseline': 0.35, 'candidate': 0.55} → candidate wins 55% of comparisons
print(f"Best model: {analysis.best_model}")

Why swap-aggregate? Without it, if the judge prefers the first response shown, and you always show baseline first, you'll systematically underrate the candidate.

Step 5: Compute Statistics

For each dimension, report:

  • Candidate win rate, baseline win rate, tie rate
  • Bootstrap 95% confidence interval
  • Verdict: better / worse / tied / inconclusive

PairwiseAnalyzer.analyze interprets each comparison as score >= 0.5 → model_a wins, using the row's metadata.model_a / metadata.model_b. So derive a per-comparison winner list from dataset + results, then bootstrap over that list — never index the PairwiseAnalysisResult object (it has no per-sample rows).

python
import numpy as np
from openjudge.graders.schema import GraderScore

def per_comparison_winners(dataset, grader_results):
    """One named winner per comparison row (handles swapped order via metadata)."""
    winners = []
    for sample, result in zip(dataset, grader_results):
        if not isinstance(result, GraderScore):
            continue  # skip errors
        meta = sample.get("metadata", {})
        winners.append(meta["model_a"] if result.score >= 0.5 else meta["model_b"])
    return winners

def bootstrap_win_rate(winners, target, n_iter=1000):
    n = len(winners)
    rates = []
    for _ in range(n_iter):
        idx = np.random.choice(n, n, replace=True)
        rates.append(sum(1 for i in idx if winners[i] == target) / n)
    return float(np.percentile(rates, 2.5)), float(np.percentile(rates, 97.5))

winners = per_comparison_winners(dataset, results["relevance"])
n = len(winners)
candidate_rate = sum(1 for w in winners if w == "candidate") / n
baseline_rate = sum(1 for w in winners if w == "baseline") / n
ci_low, ci_high = bootstrap_win_rate(winners, target="candidate")

if ci_low > 0.5:
    verdict = "candidate BETTER"
elif ci_high < 0.5:
    verdict = "candidate WORSE"
elif (ci_high - ci_low) < 0.3:
    verdict = "TIED (CI brackets 0.5, narrow)"
else:
    verdict = "INCONCLUSIVE (CI too wide — need more samples)"

print({"candidate_win_rate": candidate_rate, "baseline_win_rate": baseline_rate,
       "ci_95": [ci_low, ci_high], "verdict": verdict})

Note: with swap-aggregate each query produces 2 comparison rows. Bootstrapping over rows (above) is the simple approach; for a tighter estimate, bootstrap over queries and average the 2 swapped rows per query so position pairs stay together.

Step 6: Present Results

Prompt Regression: v1 (baseline) vs v2 (candidate)
Task: Customer support chatbot
Samples: 50

Dimension            Candidate  Baseline  Tie   95% CI         Verdict
===========================================================================
Answer relevance        58%       32%      10%   [51%, 65%]   ✓ BETTER
Factual accuracy        48%       44%       8%   [41%, 55%]   = TIED
Tone appropriateness    38%       52%      10%   [31%, 45%]   ✗ WORSE
Conciseness             62%       28%      10%   [55%, 69%]   ✓ BETTER

Summary: v2 is significantly better on relevance and conciseness,
but worse on tone appropriateness. The tone regression likely comes
from the new "be direct" instruction — consider softening it.

Top 3 tone failures (candidate worse):
  1. Query: "I'm really frustrated..." → v2 response too curt
  2. Query: "This is my first time..." → v2 missing empathetic opening
  3. Query: "Can you help me understand..." → v2 skipped explanation

Common Mistakes

  • Not doing position swap. Position bias in LLM judges is 5-15%. Without swap-aggregate, results are systematically skewed.
  • Comparing with < 10 samples. Bootstrap CI at n=10 is ±15%+ half-width. At n=5 it's ±25%+. Results are noise, not signal. Minimum 10, prefer 30+.
  • Single "overall" comparison without dimensions. "V2 is 55% better" hides that it's +20% on relevance but -15% on tone. Always report per-dimension.
  • Accepting ties as "no difference." A true tie and insufficient data look identical without CI. Always report confidence intervals.
  • Not pinning model versions. If baseline and candidate are run on different model versions (even same model, different date), model drift contaminates the prompt comparison. Same model, same version, same temperature.

Next Skills

After 06-prompt-regression:

  • 03-align-human: Calibrate the pairwise judge against human preferences.
  • 02-metric-design: Turn validated dimensions into permanent graders.
  • 04-eval-report: Include prompt comparison results in a comprehensive report.

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/eval_pipeline/06-prompt-regression of agentscope-ai/OpenJudge.

  • SKILL.md
  • scripts/pairwise.py

Open the folder on GitHubat commit d1e0642

Compare with similar skills

Prompt Regression next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prompt Regression compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prompt Regression this skillagentscope-ai/OpenJudge871—~2.8kAutomated safety check: PassApache-2.0
Prompt Governancealirezarezvani/claude-skills28k—~2.8kAutomated safety check: PassMIT
Sap AI Coresecondsky/sap-skills462—~3.3kAutomated safety check: PassGPL-3.0
Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit2603 repos~1.4kAutomated safety check: PassCustom licence
Create Simple Promptpnp/copilot-prompts893—~2.6kAutomated safety check: PassMIT
LLM Application DevMoizIbnYousaf/ai-agent-skills1.1k1 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Prompt Governance

    alirezarezvani/claude-skills

    A skill your agent uses when managing prompts in production at scale: versioning prompts, running A/B tests on prompts, building prompt registries, preventing prompt regressions, or creating eval…

    28k GitHub stars~2.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Sap AI Core

    secondsky/sap-skills

    Guides development with SAP AI Core and SAP AI Launchpad for enterprise AI/ML workloads on SAP BTP.

    462 GitHub stars~3.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Senior Prompt Engineer

    maslennikov-ig/claude-code-orchestrator-kit

    Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.

    260 GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Create Simple Prompt

    pnp/copilot-prompts

    This skill should be used when the user asks to "create a new prompt sample", "add a new prompt sample", "scaffold a new prompt sample", "create a prompt contribution", "add a prompt", or needs to…

    893 GitHub stars~2.6k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Application Dev

    MoizIbnYousaf/ai-agent-skills

    Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.

    1.1k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Context Audit

    undefined-ui/second-brain-os

    Audit an agent's context layout against the four places: system prompt, tools, history, tail.

    1k GitHub stars~802 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    871 GitHub stars~3.1k tokensUpdated 29 days ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    871 GitHub stars~2.4k tokensUpdated 29 days ago
    Auto-check passed
  • Claude Authenticity

    agentscope-ai/OpenJudge

    Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

    871 GitHub starsUsed in 1 repo~5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    871 GitHub stars~2.8k tokensUpdated 29 days ago
    Auto-check: warnings
  • Find Skills Combo

    agentscope-ai/OpenJudge

    Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.

    871 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check: warnings
  • 00 Academic Router

    agentscope-ai/OpenJudge

    A skill your agent uses when the user wants help with academic papers or citations but it's unclear which specific workflow fits — reviewing a paper, checking a BibTeX file for fake references, or…

    871 GitHub stars~976 tokensUpdated 29 days ago
    Auto-check passed

Questions about Prompt Regression

What does Prompt Regression do?

A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. Prompt Regression is an agent skill from agentscope-ai/OpenJudge.) and wants to know whether the candidate is better or worse than the baseline.

When should I use Prompt Regression?

Prompt Regression fits situations like: the user has changed a prompt (system prompt; agent instruction; etc.) and wants to know whether the candidate is better; worse than the baseline.

How do I install Prompt Regression in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a claude-code`. Or copy the skill folder (skills/eval_pipeline/06-prompt-regression in agentscope-ai/OpenJudge) into .claude/skills/prompt-regression in your project. Claude Code loads it when a task matches its description.

How do I install Prompt Regression in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a codex`. Or copy the skill folder (skills/eval_pipeline/06-prompt-regression in agentscope-ai/OpenJudge) into .agents/skills/prompt-regression in your project. Codex loads it when a task matches its description.

Can I use Prompt Regression in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-regression, .gemini/skills/prompt-regression, .github/skills/prompt-regression and .opencode/skills/prompt-regression in your project.

What does Prompt Regression need to run?

Going by SKILL.md and its folder, Prompt Regression needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Prompt Regression access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prompt Regression safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Prompt Regression use?

Prompt Regression is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prompt Regression use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prompt Regression?

Skills that share tags, products or a category with Prompt Regression: Prompt Governance (alirezarezvani/claude-skills, 28k stars), Sap AI Core (secondsky/sap-skills, 462 stars), Senior Prompt Engineer (maslennikov-ig/claude-code-orchestrator-kit, 260 stars) and Create Simple Prompt (pnp/copilot-prompts, 893 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prompt Regression?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.