Prompt Governance
alirezarezvani/claude-skills
A skill your agent uses when managing prompts in production at scale: versioning prompts, running A/B tests on prompts, building prompt registries, preventing prompt regressions, or creating eval…
A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.
$ npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentscope-ai/OpenJudge prompt-regression --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval_pipeline/06-prompt-regression .claude/skills/prompt-regression && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "prompt-regression" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression into .claude/skills/prompt-regression/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-regression", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regressionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentscope-ai/OpenJudge prompt-regression --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval_pipeline/06-prompt-regression .agents/skills/prompt-regression && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "prompt-regression" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression into .agents/skills/prompt-regression/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-regression", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentscope-ai/OpenJudge prompt-regression --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval_pipeline/06-prompt-regression .cursor/skills/prompt-regression && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "prompt-regression" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression into .cursor/skills/prompt-regression/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-regression", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentscope-ai/OpenJudge.git --path skills/eval_pipeline/06-prompt-regression--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentscope-ai/OpenJudge prompt-regression --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval_pipeline/06-prompt-regression .gemini/skills/prompt-regression && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "prompt-regression" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression into .gemini/skills/prompt-regression/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-regression", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentscope-ai/OpenJudge prompt-regressionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval_pipeline/06-prompt-regression .github/skills/prompt-regression && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "prompt-regression" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression into .github/skills/prompt-regression/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-regression", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentscope-ai/OpenJudge prompt-regression --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval_pipeline/06-prompt-regression .opencode/skills/prompt-regression && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "prompt-regression" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/06-prompt-regression into .opencode/skills/prompt-regression/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prompt-regression", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
prompt-regressionA skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.
Prompt Regression is an agent skill from agentscope-ai/OpenJudge. Use when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. Also use when the user mentions prompt A/B testing, prompt comparison, prompt optimization validation, "did my prompt change help," or prompt regression testing. Outputs per-dimension win rates with statistical significance using OpenJudge PairwiseAnalyzer.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/pairwise.py`).
It sits in AI & LLM Engineering, covering Prompt engineering, A/B testing and QA and bug reports. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Prompt Regression loads about 2.8k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 822 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 822 words, ~2,811 tokens.
.claude/skills/prompt-regression/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.<HARD-GATE>
NO conclusion about which prompt is better WITHOUT bootstrap 95% CI reported.
NO candidate declared "better" WITHOUT position-debiased (swap-aggregate) comparison.
NO comparison with fewer than 10 samples per axis — CI is too wide to be meaningful.
</HARD-GATE>
Compare two prompts head-to-head and determine, with statistical rigor, whether the candidate is better, worse, or tied on each evaluation dimension.
You MUST create a task for each item and complete them in order:
Don't hand-write the win-rate + bootstrap math (the swap-aggregation and CI are easy to get
wrong). Run the bundled, tested script (scripts/pairwise.py, standard library only, no
OpenJudge dependency):
python scripts/pairwise.py --comparisons comparisons.jsonl --candidate candidate --baseline baselineEach comparison row: {"id","model_a","model_b","score","dimension"?} where score >= 0.5
means model_a won. Emit two rows per query with A/B swapped to debias position. The
script reports per-dimension candidate/baseline/tie rates, bootstrap 95% CI, and a verdict
(BETTER / WORSE / TIED / INSUFFICIENT_EVIDENCE / INCONCLUSIVE; exit 0 only if
better). --self-test to verify it.
Steps below explain how to derive dimensions and produce the comparisons (with OpenJudge or any judge); the inline snippets are the reference behind the script.
Read the baseline and candidate prompts. Identify:
Based on the task type and what changed, derive 3-5 comparison dimensions.
Chatbot / Conversational:
RAG Generation:
Code Review / Generation:
Agent Instructions:
Each dimension gets:
id (slug)pairwise or judge or ruleDecision priority:
FunctionGrader or StringMatchGrader. Free, deterministic.
Example: output length, keyword presence, JSON validity.pairwise against reference.pairwise A/B comparison.judge (binary pass/fail per output).LLM judges have position bias — the first response shown wins 5-15% more often. Swap-aggregate eliminates this: run each comparison twice with swapped positions, keep only consistent wins:
from openjudge.graders.llm_grader import LLMGrader
from openjudge.graders.schema import GraderMode
from openjudge.runner.grading_runner import GradingRunner
from openjudge.analyzer.pairwise_analyzer import PairwiseAnalyzer
# Judge prompt for relevance comparison
relevance_judge = LLMGrader(
model=model,
name="relevance_compare",
mode=GraderMode.POINTWISE,
template="""
Compare Response A and Response B for the query below.
Which response better addresses the user's question?
Query: {query}
Response A: {response_a}
Response B: {response_b}
Score 1.0 if A is better, 0.0 if B is better, 0.5 if tied.
Respond in JSON: {{"score": <float>, "reason": "<explanation>"}}
""",
)
# Build pairwise dataset with position swap
dataset = []
for sample in test_samples:
# Original order
dataset.append({
"query": sample["query"],
"response_a": baseline_outputs[sample["id"]],
"response_b": candidate_outputs[sample["id"]],
"metadata": {"model_a": "baseline", "model_b": "candidate"},
})
# Swapped order — critical for debiasing
dataset.append({
"query": sample["query"],
"response_a": candidate_outputs[sample["id"]],
"response_b": baseline_outputs[sample["id"]],
"metadata": {"model_a": "candidate", "model_b": "baseline"},
})
runner = GradingRunner(
grader_configs={"relevance": relevance_judge},
max_concurrency=8,
)
results = await runner.arun(dataset)
# Analyze with PairwiseAnalyzer
analyzer = PairwiseAnalyzer(model_names=["baseline", "candidate"])
analysis = analyzer.analyze(dataset, results["relevance"])
print(f"Win rates: {analysis.win_rates}")
# → {'baseline': 0.35, 'candidate': 0.55} → candidate wins 55% of comparisons
print(f"Best model: {analysis.best_model}")Why swap-aggregate? Without it, if the judge prefers the first response shown, and you always show baseline first, you'll systematically underrate the candidate.
For each dimension, report:
PairwiseAnalyzer.analyze interprets each comparison as score >= 0.5 → model_a wins,
using the row's metadata.model_a / metadata.model_b. So derive a per-comparison
winner list from dataset + results, then bootstrap over that list — never index the
PairwiseAnalysisResult object (it has no per-sample rows).
import numpy as np
from openjudge.graders.schema import GraderScore
def per_comparison_winners(dataset, grader_results):
"""One named winner per comparison row (handles swapped order via metadata)."""
winners = []
for sample, result in zip(dataset, grader_results):
if not isinstance(result, GraderScore):
continue # skip errors
meta = sample.get("metadata", {})
winners.append(meta["model_a"] if result.score >= 0.5 else meta["model_b"])
return winners
def bootstrap_win_rate(winners, target, n_iter=1000):
n = len(winners)
rates = []
for _ in range(n_iter):
idx = np.random.choice(n, n, replace=True)
rates.append(sum(1 for i in idx if winners[i] == target) / n)
return float(np.percentile(rates, 2.5)), float(np.percentile(rates, 97.5))
winners = per_comparison_winners(dataset, results["relevance"])
n = len(winners)
candidate_rate = sum(1 for w in winners if w == "candidate") / n
baseline_rate = sum(1 for w in winners if w == "baseline") / n
ci_low, ci_high = bootstrap_win_rate(winners, target="candidate")
if ci_low > 0.5:
verdict = "candidate BETTER"
elif ci_high < 0.5:
verdict = "candidate WORSE"
elif (ci_high - ci_low) < 0.3:
verdict = "TIED (CI brackets 0.5, narrow)"
else:
verdict = "INCONCLUSIVE (CI too wide — need more samples)"
print({"candidate_win_rate": candidate_rate, "baseline_win_rate": baseline_rate,
"ci_95": [ci_low, ci_high], "verdict": verdict})Note: with swap-aggregate each query produces 2 comparison rows. Bootstrapping over rows (above) is the simple approach; for a tighter estimate, bootstrap over queries and average the 2 swapped rows per query so position pairs stay together.
Prompt Regression: v1 (baseline) vs v2 (candidate)
Task: Customer support chatbot
Samples: 50
Dimension Candidate Baseline Tie 95% CI Verdict
===========================================================================
Answer relevance 58% 32% 10% [51%, 65%] ✓ BETTER
Factual accuracy 48% 44% 8% [41%, 55%] = TIED
Tone appropriateness 38% 52% 10% [31%, 45%] ✗ WORSE
Conciseness 62% 28% 10% [55%, 69%] ✓ BETTER
Summary: v2 is significantly better on relevance and conciseness,
but worse on tone appropriateness. The tone regression likely comes
from the new "be direct" instruction — consider softening it.
Top 3 tone failures (candidate worse):
1. Query: "I'm really frustrated..." → v2 response too curt
2. Query: "This is my first time..." → v2 missing empathetic opening
3. Query: "Can you help me understand..." → v2 skipped explanationAfter 06-prompt-regression:
03-align-human: Calibrate the pairwise judge against human preferences.02-metric-design: Turn validated dimensions into permanent graders.04-eval-report: Include prompt comparison results in a comprehensive report.© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in skills/eval_pipeline/06-prompt-regression of agentscope-ai/OpenJudge.
Open the folder on GitHubat commit d1e0642
Prompt Regression next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Prompt Regression this skillagentscope-ai/OpenJudge | 871 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Prompt Governancealirezarezvani/claude-skills | 28k | — | ~2.8k | Automated safety check: Pass | MIT | |
| Sap AI Coresecondsky/sap-skills | 462 | — | ~3.3k | Automated safety check: Pass | GPL-3.0 | |
| Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit | 260 | 3 repos | ~1.4k | Automated safety check: Pass | Custom licence | |
| Create Simple Promptpnp/copilot-prompts | 893 | — | ~2.6k | Automated safety check: Pass | MIT | |
| LLM Application DevMoizIbnYousaf/ai-agent-skills | 1.1k | 1 repos | ~1.3k | Automated safety check: Pass | MIT |
alirezarezvani/claude-skills
A skill your agent uses when managing prompts in production at scale: versioning prompts, running A/B tests on prompts, building prompt registries, preventing prompt regressions, or creating eval…
secondsky/sap-skills
Guides development with SAP AI Core and SAP AI Launchpad for enterprise AI/ML workloads on SAP BTP.
maslennikov-ig/claude-code-orchestrator-kit
Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.
pnp/copilot-prompts
This skill should be used when the user asks to "create a new prompt sample", "add a new prompt sample", "scaffold a new prompt sample", "create a prompt contribution", "add a prompt", or needs to…
MoizIbnYousaf/ai-agent-skills
Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.
undefined-ui/second-brain-os
Audit an agent's context layout against the four places: system prompt, tools, history, tail.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…
agentscope-ai/OpenJudge
A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.
agentscope-ai/OpenJudge
Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.
agentscope-ai/OpenJudge
A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…
agentscope-ai/OpenJudge
Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.
agentscope-ai/OpenJudge
A skill your agent uses when the user wants help with academic papers or citations but it's unclear which specific workflow fits — reviewing a paper, checking a BibTeX file for fake references, or…
Categories
A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. Prompt Regression is an agent skill from agentscope-ai/OpenJudge.) and wants to know whether the candidate is better or worse than the baseline.
Prompt Regression fits situations like: the user has changed a prompt (system prompt; agent instruction; etc.) and wants to know whether the candidate is better; worse than the baseline.
Run `npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a claude-code`. Or copy the skill folder (skills/eval_pipeline/06-prompt-regression in agentscope-ai/OpenJudge) into .claude/skills/prompt-regression in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a codex`. Or copy the skill folder (skills/eval_pipeline/06-prompt-regression in agentscope-ai/OpenJudge) into .agents/skills/prompt-regression in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill prompt-regression -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prompt-regression, .gemini/skills/prompt-regression, .github/skills/prompt-regression and .opencode/skills/prompt-regression in your project.
Going by SKILL.md and its folder, Prompt Regression needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Prompt Regression is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Prompt Regression: Prompt Governance (alirezarezvani/claude-skills, 28k stars), Sap AI Core (secondsky/sap-skills, 462 stars), Senior Prompt Engineer (maslennikov-ig/claude-code-orchestrator-kit, 260 stars) and Create Simple Prompt (pnp/copilot-prompts, 893 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.
Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.