Agent skill

Agentsop Regression Gate

by agentsope in agentsope/SkillAlchemy

Build a held-out eval set, run it on every prompt/model change, and block regressions in CI.

MITAuto-check passedBusiness, Finance & HR

Install Agentsop Regression Gate

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-regression-gate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-regression-gate .claude/skills/agentsop-regression-gate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-regression-gate
GitHub stars
459
Token cost
~6.1k tokens
SKILL.md length
2,999 words
Files
4 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Build a held-out eval set, run it on every prompt/model change, and block regressions in CI.

  • Works in 7 steps: 何时激活 (When to Activate) → 核心心智模型 (Core Mental Model) → SOP (Standard Operating Procedure) → …
  • Tasks that involve Operations and SOPs
  • SKILL.md covers 1. 何时激活 (When to Activate), 2. 核心心智模型 (Core Mental Model), 3. SOP (Standard Operating… and 4. 操作模型 (Operations), plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentsop Regression Gate is an agent skill from agentsope/SkillAlchemy. Build a held-out eval set, run it on every prompt/model change, and block regressions in CI. An LM change is a code change — gate it with a test suite (eval set + metric + threshold). Cross-framework SOP not surfaced by any single base skill.

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).

It sits in Business, Finance & HR, covering Operations and SOPs and Test generation. It works with LlamaIndex. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Tasks that involve Operations and SOPs
  • Tasks that involve Test generation

Example prompts

  • “/agentsop-regression-gate”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 何时激活 (When to Activate)
  2. 核心心智模型 (Core Mental Model)
  3. SOP (Standard Operating Procedure)
  4. 操作模型 (Operations)
  5. 困境决策案例 (Dilemma Cases)
  6. 反模式与边界 (Anti-Patterns & Boundaries)
  7. 跨框架对照 (Cross-Framework Mapping)

What it can do on your machine

Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop Regression Gate loads about 6.1k tokens when it runs, and up to ~7.9k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 2,999 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 2,999 words, ~6,125 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-regression-gate/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
agentsop-regression-gate
description
Build a held-out eval set, run it on every prompt/model change, and block regressions in CI. An LM change is a code change — gate it with a test suite (eval set + metric + threshold). Cross-framework SOP not surfaced by any single base skill.
version
0.1.0
phase
D
tier
core
frequency
high
status
opinionated

regression-gate — Eval Set + Metric + Threshold, Wired Into CI

"Every subsequent change must be gated on these numbers." — Synthesized from [[llamaindex]] Stage 2 (eval loop before optimizing) [llamaindex-sop-skill/SKILL.md:114-126]

"Compiled program beats baseline on a held-out test set (not the val set used in optimization)." — [[dspy]] Stage 3 exit criterion [dspy-sop-skill/SKILL.md:101]

This is an enhancement overlay. The regression-gate SOP exists only as fragments scattered across base skills — [[llamaindex]] OP-10 EvalLoop ("gate every change"), [[dspy]] train/dev/test split + metric — and is never assembled as a standalone cross-framework discipline. It is the discipline that turns a one-off eval into a gate: a test suite that runs in CI on every prompt/model/retriever change and fails the build on regression. It consumes a metric from [[agentsop-metric-design]] and, for domain-specific held-out sets, hands off to [[agentsop-domain-eval-set]].


1. 何时激活 (When to Activate)

Activate when any of these is true:

  • Any prompt change you want to ship safely: a prompt edit, a system-message tweak, a few-shot-demo swap is about to merge and you have no automated way to know if it made things worse.
  • Any model change: swapping GPT-4o → a cheaper/newer model, a temperature change, a provider migration. An LM change silently shifts the whole output distribution. [[dspy]] Case B: "If you optimize a complex pipeline for GPT-4, it usually breaks on a smaller model" [dspy-sop-skill/SKILL.md:184].
  • Any retriever/chunking/reranker change in a RAG pipeline: every such change needs a quantitative gate. [[llamaindex]] OP-10: "Quantitative regression test for every chunking / embedding / retriever / prompt change" [llamaindex-sop-skill/SKILL.md:235].
  • Recurring "it got worse" surprises: the team keeps shipping changes that users report as regressions after the fact. The fix is a gate, not more careful review.
  • Setting up CI for an LLM app and there is no eval job in the pipeline.

Do NOT activate for:

  • One-shot tasks with no production surface — there is nothing to regress. ([[dspy]] boundary: "Summarize this email once → raw API call" [dspy-sop-skill/SKILL.md:259].)
  • The signature/task is still changing daily — gate only after the I/O contract stabilizes, else you re-baseline every commit. ([[dspy]] boundary [dspy-sop-skill/SKILL.md:260].)
  • No willingness to define any success criterion — without a metric there is nothing to gate. Route to [[agentsop-metric-design]] first; if the user refuses, this skill cannot help.

2. 核心心智模型 (Core Mental Model)

"An LM change is a code change. Gate it with a test suite: eval set + metric + threshold."

You already gate code with unit tests in CI: a change that breaks a test fails the build. A prompt edit, a model swap, a chunk-size tweak are also changes to the system's behavior — but they slip through review because their effect is statistical, not a stack trace. The regression gate is the missing unit test for LM behavior.

The gate is exactly three artifacts plus a wiring step:

   eval set        metric           threshold          CI wiring
  (held-out QA)  (ex,pred)->score   (fail if <X / drop>Y)  (block merge)
        │              │                  │                    │
        └──────────────┴──────────────────┴────────────────────┘
                          REGRESSION GATE

Three load-bearing principles:

  1. The eval set is held out and frozen. It is a labelled, version-controlled fixture that the prompt/model under test has never seen. [[dspy]] is explicit: the compiled program must beat baseline on a held-out test set "not the val set used in optimization" [dspy-sop-skill/SKILL.md:101]. The split is train / dev / test; the gate runs on test only. If the eval set leaks into the prompt (few-shot demos, instructions), the gate measures memorization, not quality.

  2. The metric comes from [[agentsop-metric-design]], not invented here. This skill does not design metrics — it consumes one. A bad metric makes the gate theatre: it will pass changes that hurt users and block changes that help them. The metric must be human-calibrated before it gates anything ([[agentsop-metric-design]] OP-M05).

  3. The threshold is a policy, not a number you guess. Two common shapes: an absolute floor (fail if score < X) and a relative no-regression (fail if score drops > Y from the committed baseline). Relative is the regression gate proper; absolute is a quality bar. Most teams use both: a floor for "never ship below this," plus a no-regression delta for "this PR must not make it worse."

Build the eval loop before you optimize anything

[[llamaindex]] Stage 2 is named "Build the eval loop before optimizing anything" [llamaindex-sop-skill/SKILL.md:114]. The anti-pattern it names is A3: "No eval loop; debug by anecdote" [llamaindex-sop-skill/SKILL.md:348]. The gate is the institutional form of that loop — once it exists, every change is debugged by number, not by vibe.


3. SOP (Standard Operating Procedure)

0. Confirm activation (§1); confirm a metric exists or invoke [[agentsop-metric-design]]
1. BUILD eval set:  generate candidates -> curate to a golden set -> freeze + version
2. SPLIT:           train / dev / test; the GATE runs on TEST only
3. PICK metric:     consume from [[agentsop-metric-design]] (do not invent here)
4. SET threshold:   absolute floor AND/OR relative no-regression delta
5. WIRE into CI:    run eval on every prompt/model/retriever PR; fail on regression
6. HANDLE flakiness: pin seeds/temp, average N runs, separate flaky from real drops
Stage 1 — Build the eval set (generate + curate)

Two stages, never one. Generation gives coverage cheaply; curation gives trust.

  • Generate candidates from your corpus. [[llamaindex]] DatasetGenerator.from_documents(docs).generate_dataset_from_nodes(num=50) produces labelled QA pairs from the documents themselves [llamaindex-sop-skill/SKILL.md:121]. promptfoo and synthetic-data generators do the same for non-RAG tasks.
  • Curate the generated set into a golden set: a human reviews, fixes wrong labels, drops ambiguous items, and adds known hard/edge cases the generator missed. A purely generated set inherits the generator model's blind spots and tends to be too easy (Dilemma 1).
  • Freeze and version. The golden set is a committed fixture (eval/golden_v1.jsonl), tagged with the date and the generator model. Changing it is a versioned event, not an edit.

Size: [[dspy]] documents the sweet spot — "30 examples = minimum useful, 300 = recommended" [dspy-sop-skill/SKILL.md:87]. For a held-out gate, 50–200 curated domain examples is the working range; descend to [[agentsop-domain-eval-set]] for domain-specific construction.

Stage 2 — Split: train / dev / test

The gate runs on test only. Keep test sealed from anything that touches the prompt:

  • train — feeds optimizers / few-shot demo selection.
  • dev — tuning and threshold-setting.
  • test — the gate. Never used to author prompts, pick demos, or tune. ([[dspy]] held-out exit criterion [dspy-sop-skill/SKILL.md:101].)

Do not invent a metric here. Consume one from [[agentsop-metric-design]]:

  • RAG → the Faithfulness + Relevancy + Retriever(MRR/hit-rate) triad ([[llamaindex]] OP-10 [llamaindex-sop-skill/SKILL.md:234]).
  • Exact-answer → exact-match ([[dspy]] §4.3 [dspy-sop-skill/SKILL.md:137]).
  • Open-ended → decomposed sub-judges, bool-in-compile/float-in-eval, length penalty ([[agentsop-metric-design]] OP-M01/OP-M02/OP-M03).

A metric that has not been human-calibrated must not gate ([[agentsop-metric-design]] OP-M05). An uncalibrated gate is worse than no gate — it gives false confidence.

Stage 4 — Set the threshold
Threshold shapeRuleUse when
Absolute floorfail if score(test) < X"never ship below this quality bar"
Relative no-regressionfail if baseline − score > Δ"this PR must not make it worse" (the gate proper)
Per-slice floorfail if any slice (e.g. lexical-query subset) drops > Δaggregate hides a regressed minority

Set Δ above measured run-to-run noise (Stage 6), else the gate flaps. Commit the current test score as baseline.json next to the eval set; the gate compares against it.

Stage 5 — Wire into CI
  • Add an eval CI job that runs on every PR touching prompts, model config, retriever/chunking config, or the program graph.
  • Job: load frozen eval set → run pipeline at PR's config → compute metric → compare to baseline.json → exit non-zero on threshold breach → post the before/after table as a PR comment.
  • On merge to main with an intended improvement, bump baseline.json in the same PR (reviewed, not silent).
Stage 6 — Handle flaky evals

LLM outputs are nondeterministic; a naive gate flaps and gets disabled. Mitigations:

  • Pin temperature=0 and seeds where the provider supports them; disable response caching in CI ([[dspy]] AP-10: "Forgetting cache=False in stateless deploys" [dspy-sop-skill/SKILL.md:255]).
  • Average N runs (e.g. 3) and gate on the mean; report variance.
  • Separate flaky from real: if the same config scores differently across reruns by more than Δ, the gate (Δ too tight) or the metric (noisy judge) is the problem — fix those before trusting a single red build. A judge-based metric with high variance should be hardened in [[agentsop-metric-design]], not papered over by widening Δ.

4. 操作模型 (Operations)

OP-01 — GenerateEvalCandidates
  • Trigger: Need a held-out eval set; no labelled fixture exists yet.
  • Action: Generate QA candidates from the corpus — [[llamaindex]] DatasetGenerator.generate_dataset_from_nodes(num=50), promptfoo synthetic generation, or task-specific synthesis. Tag with generator model + date.
  • Output: A raw candidate set (unvalidated), ready for curation.
  • Evidence: [[llamaindex]] Stage 2 [llamaindex-sop-skill/SKILL.md:121]; OP-10 [llamaindex-sop-skill/SKILL.md:234]; external "eval set generation".
OP-02 — CurateGoldenSet
  • Trigger: A generated candidate set exists; it has not been human-reviewed.
  • Action: Human reviews every item: fix wrong labels, drop ambiguous/duplicate items, inject known hard cases and past production failures. Freeze as a versioned fixture (golden_vN.jsonl).
  • Output: A trusted, frozen golden eval set that the gate runs against.
  • Evidence: [[dspy]] metric human-validation discipline [dspy-sop-skill/SKILL.md:212]; [[agentsop-metric-design]] OP-M05; Dilemma 1.
OP-03 — SplitTrainDevTest
  • Trigger: Eval set built; about to use it for both tuning and gating.
  • Action: Partition into train/dev/test. Seal test from all prompt-authoring. The gate reads only test. (Note [[dspy]]'s reversed 20/80 train/val split for prompt optimizers [dspy-sop-skill/SKILL.md:96] — that is an optimizer concern; the gate still needs an untouched test slice.)
  • Output: Three disjoint splits; a sealed test set for the gate.
  • Evidence: [[dspy]] Stage 3 held-out exit criterion [dspy-sop-skill/SKILL.md:101]; AP-9 [dspy-sop-skill/SKILL.md:254].
OP-04 — SetRegressionThreshold
  • Trigger: Have a test score; need a pass/fail policy.
  • Action: Commit current test score as baseline.json. Define absolute floor X and/or relative no-regression Δ (Δ > measured noise). Optionally per-slice floors.
  • Output: A versioned threshold policy the CI job enforces.
  • Evidence: [[dspy]] Stage 3 exit "by ≥ task-relevant delta" [dspy-sop-skill/SKILL.md:101]; [[llamaindex]] gate-every-change [llamaindex-sop-skill/SKILL.md:124].
OP-05 — WireCIGate
  • Trigger: Eval set + metric + threshold exist; not yet enforced automatically.
  • Action: Add a CI job triggered by prompt/model/retriever/graph changes: run eval at PR config → compute metric → compare to baseline → fail on breach → comment the before/after table.
  • Output: Merges that regress quality are blocked; every change carries a number.
  • Evidence: [[llamaindex]] OP-10 regression-test framing [llamaindex-sop-skill/SKILL.md:233-235]; external "llm regression testing CI", "promptfoo".
OP-06 — BumpBaselineOnIntendedWin
  • Trigger: A PR intentionally raises quality; the gate would otherwise pin the old baseline forever.
  • Action: In the same reviewed PR, update baseline.json to the new test score. Never let CI auto-bump silently.
  • Output: Baseline ratchets upward deliberately; future regressions are caught against the new bar.
  • Evidence: [[dspy]] "keep both program.gpt4o.json and program.llama8b.json; A/B" [dspy-sop-skill/SKILL.md:191] (versioned-artifact discipline).
OP-07 — StabilizeFlakyEval
  • Trigger: The gate flaps — same config, different verdicts across reruns.
  • Action: Pin temperature/seed, disable CI caching, average N runs and gate on the mean, report variance. If variance > Δ, fix the metric (harden judge via [[agentsop-metric-design]]) or widen Δ — never disable the gate.
  • Output: A stable gate whose red builds are trustworthy.
  • Evidence: [[dspy]] cache=False AP-10 [dspy-sop-skill/SKILL.md:255]; [[agentsop-metric-design]] judge-bias hardening; [[dspy]] Stage 2 exit "stable across two runs" [dspy-sop-skill/SKILL.md:91].
OP-08 — SliceTheEvalSet
  • Trigger: Aggregate score is flat but a subpopulation (lexical queries, a tenant, a topic) silently regressed.
  • Action: Tag eval items by slice; compute and gate per-slice. A win on the majority must not mask a regression on a minority slice.
  • Output: Slice-level regression detection; aggregate no longer hides harm.
  • Evidence: [[llamaindex]] per-query-type taxonomy [llamaindex-sop-skill/SKILL.md:280]; [[agentsop-metric-design]] triad (no single number).

5. 困境决策案例 (Dilemma Cases)

Show full SKILL.md (1,279 more words)Show less
Dilemma 1 — Generated eval set vs hand-curated golden set

困境: Team needs an eval set fast. DatasetGenerator produces 200 QA pairs in minutes ([[llamaindex]] Stage 2 [llamaindex-sop-skill/SKILL.md:121]). Hand-curating 200 examples costs days of human time. Ship the generated set as the gate, or pay for curation?

约束: The generator is the same model family that powers the pipeline → its questions are answerable by exactly the kind of reasoning the pipeline already does (self-preference leakage). Generated sets skew easy and miss the long-tail failures users actually hit. But zero eval set means shipping blind (anti-pattern).

决策步骤:

  1. Generate for coverage, never gate on raw generation. Use the generated set as a candidate pool, not the gate.
  2. Curate a golden subset (OP-02): a human keeps the good items, fixes labels, drops the trivially-easy and the ambiguous, and injects known production failures and adversarial/edge cases the generator never proposes.
  3. Cross-family generation where possible: generate with a different model family than the task model to reduce self-preference leakage (mirrors [[agentsop-metric-design]] OP-M04).
  4. Size to budget: 50 curated > 200 raw. [[dspy]] floor is 30 useful examples [dspy-sop-skill/SKILL.md:87]; a tight, curated, edge-case-loaded 50 gates better than a bloated easy 200.
  5. Version both: keep the raw generated pool (regeneratable) and the frozen golden set (the gate fixture).

结果: The golden set is the gate; the generated pool is scaffolding. Teams that gate on raw generated sets ship regressions that the easy set never exercised — the gate was green while users churned.

可提取的操作: OP-01 GenerateEvalCandidates, OP-02 CurateGoldenSet. Lesson: generation buys coverage, curation buys trust. A gate needs trust — never gate on an uncurated generated set.

Dilemma 2 — Threshold too strict blocks good changes

困境: A no-regression gate is set at Δ = 0 (any drop fails). A genuinely good refactor — simpler prompt, 40% cheaper model — scores 0.81 vs the 0.83 baseline: a 2-point drop within run-to-run noise. The gate blocks a change that is net-positive (equal quality, far cheaper). The team starts overriding the gate, and soon ignores it entirely.

约束: Run-to-run noise on this judge-based metric is ±1.5 points (measured across 3 reruns). The 2-point "drop" is statistically indistinguishable from noise. A gate that flags noise as regression trains the team to bypass it — a bypassed gate is worse than none.

决策步骤:

  1. Measure noise first (OP-07): rerun the baseline config N times; compute the standard deviation. Here σ ≈ 1.5pp.
  2. Set Δ above noise: Δ = 2σ ≈ 3pp, not 0. A drop must clear the noise band to count as a regression.
  3. Average N runs and gate on the mean to shrink the noise band, rather than just widening Δ.
  4. Separate cost from quality: this PR is a cost win at equal quality. The quality gate should pass (drop within Δ); cost is tracked on its own axis (cross-link [[agentsop-cost-tiered-models]]). Don't let a quality gate block a cost win that doesn't hurt quality.
  5. If the metric is too noisy to set a sane Δ, the metric is the bug — harden it in [[agentsop-metric-design]] (decompose, length penalty, cross-family judge), don't widen Δ to infinity.

结果: Δ tuned to ~2σ passes the cheaper-equal-quality change, still catches real regressions (a 6pp drop), and the team keeps trusting the gate. A gate calibrated to noise survives; a Δ=0 gate gets disabled.

可提取的操作: OP-04 SetRegressionThreshold, OP-07 StabilizeFlakyEval. Lesson: the threshold must clear measured noise. A gate that flags noise as failure gets bypassed, and a bypassed gate protects nothing.


6. 反模式与边界 (Anti-Patterns & Boundaries)

Anti-patterns
#Anti-patternWhy it's wrongFix
AP-1No eval set; ship blindEvery prompt/model change is an uncontrolled experiment on users; "it got worse" is discovered in productionBuild a held-out gate ([[llamaindex]] A3 [llamaindex-sop-skill/SKILL.md:348])
AP-2Eval set leaks into the prompt (few-shot demos / instructions drawn from test)The gate measures memorization, not generalization; green build, real regressionSeal test; demos come from train only (OP-03; [[dspy]] AP-9 [dspy-sop-skill/SKILL.md:254])
AP-3Gate on a raw generated setInherits generator blind spots; too easy; misses real failuresCurate a golden set (OP-02; Dilemma 1)
AP-4Gate on an uncalibrated metricA wrong metric passes harmful changes and blocks good ones — gate is theatreCalibrate via [[agentsop-metric-design]] OP-M05 before gating
AP-5Δ = 0 / threshold below noiseGate flaps on noise, team bypasses itSet Δ > 2σ measured noise (OP-04, OP-07; Dilemma 2)
AP-6Run the gate on the val/dev set used for tuningOptimistic, leaks tuning into evaluationGate on held-out test only ([[dspy]] [dspy-sop-skill/SKILL.md:101])
AP-7Caching on in CIStale cached outputs mask the change under testcache=False ([[dspy]] AP-10 [dspy-sop-skill/SKILL.md:255])
AP-8Aggregate-only gateA win on the majority hides a regressed minority slicePer-slice gating (OP-08)
AP-9Silent baseline auto-bumpQuality can ratchet down unnoticed if CI rewrites baselineBump baseline only in a reviewed PR (OP-06)
AP-10Disabling the gate when it flakesRemoves the only protection; flakiness is a metric/Δ bug, not a gate bugStabilize (OP-07), never disable
Boundaries (when this skill does not apply)
  • One-shot / throwaway tasks — no production surface to regress ([[dspy]] [dspy-sop-skill/SKILL.md:259]).
  • Unstable signature — gate only after the I/O contract stabilizes ([[dspy]] [dspy-sop-skill/SKILL.md:260]); otherwise you re-baseline every commit.
  • No metric possible and none willing to be built — without a metric there is nothing to gate; route to [[agentsop-metric-design]] first.
  • Public-benchmark evaluation (MMLU, HumanEval) — that is a capability benchmark, not a domain regression gate; use lm-evaluation-harness. For a domain-specific held-out set, descend to [[agentsop-domain-eval-set]].

7. 跨框架对照 (Cross-Framework Mapping)

ConceptLlamaIndexDSPypromptfooLangSmithThis skill
Eval set generationDatasetGenerator.generate_dataset_from_nodes(num=N) [llamaindex-sop-skill/SKILL.md:121]bring labelled examples; BootstrapFewShot self-generates demos (not the test set)tests: synthesis / generate from promptsDatasets created from traces / uploadsOP-01 GenerateEvalCandidates
Golden / curated setmanual review of generated QAhand-labelled trainset/devsetcurated tests YAML with assertcurated Dataset + reference outputsOP-02 CurateGoldenSet
Train/dev/test splitmanualexplicit; reversed 20/80 for prompt optimizers, held-out test for gate [dspy-sop-skill/SKILL.md:96,101]n/a (test set is the suite)dataset splitsOP-03 SplitTrainDevTest
MetricFaithfulness/Relevancy/RetrieverEvaluator(mrr,hit_rate) [llamaindex-sop-skill/SKILL.md:234]def metric(ex,pred,trace=None)->bool|float [dspy-sop-skill/SKILL.md:88]assert (equals/contains/llm-rubric/javascript)evaluator fns / LLM-as-judgeconsumed from [[agentsop-metric-design]]
Threshold / gatemanual (gate every change [llamaindex-sop-skill/SKILL.md:124])held-out beats baseline "by ≥ delta" [dspy-sop-skill/SKILL.md:101]assert pass + --fail-on thresholdsrules + alerts on eval scoresOP-04 SetRegressionThreshold
CI wiringnot built-in (DIY job around Evaluate)not built-in (DIY around dspy.Evaluate)first-class: promptfoo eval in CI, non-zero exitCI integration + regression alertsOP-05 WireCIGate
Flaky handlingrun multiple timescache=False; "stable across two runs" [dspy-sop-skill/SKILL.md:91,255]repeat + thresholdrun aggregationOP-07 StabilizeFlakyEval

Combination patterns:

  • LlamaIndex + this skill: DatasetGenerator for OP-01, the Faithfulness/Relevancy/Retriever triad as the metric, wired into a DIY CI job. The base skill names the loop ("gate every change"); this skill makes it a CI gate.
  • DSPy + this skill: DSPy's split + metric is the eval-loop substrate; this skill adds the CI enforcement DSPy leaves to you. Gate the compiled artifact on the held-out test set; bump baseline when recompiling for a new model (Case B).
  • promptfoo as the engine: promptfoo is the most CI-native option — promptfoo eval exits non-zero on failed asserts; it is the closest off-the-shelf realization of OP-05. Use it as the runner; still bring a curated set (OP-02) and a calibrated metric.
  • LangSmith for managed datasets + monitoring: managed datasets and regression alerts; pairs with this skill's curation/threshold discipline.

Opinionated default: build the eval set with the base framework's generator (OP-01), curate by hand (OP-02), keep the metric in [[agentsop-metric-design]], and run the gate with promptfoo (CI-native) or a thin script around dspy.Evaluate / LlamaIndex evaluators. The gate, the metric, and the eval set are three separable, version-controlled artifacts — never one tangled blob.


References

  • references/R1-source-evidence.md — every cited claim resolved to a source line
  • intermediate/operation_candidates.json — machine-readable operation registry

Citations: [[llamaindex]] OP-10 EvalLoop / Stage 2 [llamaindex-sop-skill/SKILL.md:114-126,232-236,348]; [[dspy]] Stage 2-3 split+metric+held-out [dspy-sop-skill/SKILL.md:85-105,137,191,254-255]; [[agentsop-metric-design]]; [[agentsop-domain-eval-set]]; external "llm regression testing CI", "promptfoo", "eval set generation".

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/agentsop-regression-gate of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-source-evidence.md

Open the folder on GitHubat commit 6ea799f

Compare with similar skills

Agentsop Regression Gate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop Regression Gate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop Regression Gate this skillagentsope/SkillAlchemy459—~6.1kAutomated safety check: PassMIT
Tulingresearch Fusion Campaign Enginemeamaturinlove221/TuringResearch_plus136—~548Automated safety check: PassCustom licence
Cc Sdd New Agentgotalab/cc-sdd3.7k—~1.1kAutomated safety check: PassMIT
DBS Business Toolkit Entrydontbesilent2025/dbskill11k—~2kAutomated safety check: PassCustom licence
Agent Sop Authorstrands-agents/agent-sop1.2k—~3.5kAutomated safety check: PassApache-2.0
Diffusion Narrative Denouncingcanwhite/Krebs1k—~831Automated safety check: PassMIT

Similar skills

  • Tulingresearch Fusion Campaign Engine

    meamaturinlove221/TuringResearch_plus

    A skill your agent uses when maintaining Campaign - Strategy - Tactic - SOP runtime behavior.

    136 GitHub stars~548 tokensUpdated 4 mo ago
    Business, Finance & HRAuto-check passed
  • Cc Sdd New Agent

    gotalab/cc-sdd

    Add or extend coding-agent support in cc-sdd by executing the SOP in docs/cc-sdd/sop-new-agent.md end-to-end.

    3.7k GitHub stars~1.1k tokensUpdated 15 days ago
    Business, Finance & HRAuto-check passed
  • DBS Business Toolkit Entry

    dontbesilent2025/dbskill

    Chinese-language entry skill for the dontbesilent business toolkit: onboards new users, orchestrates tasks across sub-skills, runs numbered prompts and lists hidden ones.

    11k GitHub stars~2k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Agent Sop Author

    strands-agents/agent-sop

    Create (or update) and validate Agent SOPs (Standard Operating Procedures) - markdown-based workflows that guide AI agents through complex, multi-step tasks with RFC 2119 constraints.

    1.2k GitHub stars~3.5k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • 基于"扩散模型叙事去噪流"的小说写作 SOP。将 AI 视为去杂质机器,通过锁定全局信号、预测叙事噪声、精准去噪、随机修正四个步骤,解决 AI 翻译腔、逻辑断层和故事平淡的问题。

    1k GitHub stars~831 tokensUpdated 1 mo ago
    Business, Finance & HRAuto-check passed
  • Polanyi Perspective

    0xenzyme/polanyi-skill

    Michael Polanyi 的思维框架。用 Polanyi 视角分析隐性知识、技能习得、经验传承、师徒制、 知识管理、学习方法、AI/工具替代边界、科学共同体与后批判哲学问题。

    137 GitHub stars~1.3k tokensUpdated 3 mo ago
    Business, Finance & HRAuto-check passed

More from agentsope/SkillAlchemy

All 45 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    459 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    459 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    459 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    459 GitHub stars~4.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    459 GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check: notes
  • Agentsop Multiscale Chunking

    agentsope/SkillAlchemy

    Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

    459 GitHub stars~4.9k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Agentsop Regression Gate

What does Agentsop Regression Gate do?

Build a held-out eval set, run it on every prompt/model change, and block regressions in CI. Agentsop Regression Gate is an agent skill from agentsope/SkillAlchemy. Build a held-out eval set, run it on every prompt/model change, and block regressions in CI.

When should I use Agentsop Regression Gate?

Agentsop Regression Gate fits situations like: tasks that involve Operations and SOPs; tasks that involve Test generation.

How do I install Agentsop Regression Gate in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a claude-code`. Or copy the skill folder (skills/agentsop-regression-gate in agentsope/SkillAlchemy) into .claude/skills/agentsop-regression-gate in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop Regression Gate in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a codex`. Or copy the skill folder (skills/agentsop-regression-gate in agentsope/SkillAlchemy) into .agents/skills/agentsop-regression-gate in your project. Codex loads it when a task matches its description.

Can I use Agentsop Regression Gate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-regression-gate, .gemini/skills/agentsop-regression-gate, .github/skills/agentsop-regression-gate and .opencode/skills/agentsop-regression-gate in your project.

What does Agentsop Regression Gate need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentsop Regression Gate is instructions for the agent only.

Does Agentsop Regression Gate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentsop Regression Gate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop Regression Gate use?

Agentsop Regression Gate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop Regression Gate use?

About 6.1k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop Regression Gate?

Skills that share tags, products or a category with Agentsop Regression Gate: Tulingresearch Fusion Campaign Engine (meamaturinlove221/TuringResearch_plus, 136 stars), Cc Sdd New Agent (gotalab/cc-sdd, 3.7k stars), DBS Business Toolkit Entry (dontbesilent2025/dbskill, 11k stars) and Agent Sop Author (strands-agents/agent-sop, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop Regression Gate?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 459 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.