Tulingresearch Fusion Campaign Engine
meamaturinlove221/TuringResearch_plus
A skill your agent uses when maintaining Campaign - Strategy - Tactic - SOP runtime behavior.
Build a held-out eval set, run it on every prompt/model change, and block regressions in CI.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-regression-gate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-regression-gate .claude/skills/agentsop-regression-gate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-regression-gate" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-regression-gate into .claude/skills/agentsop-regression-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-regression-gate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-regression-gateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-regression-gate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-regression-gate .agents/skills/agentsop-regression-gate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-regression-gate" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-regression-gate into .agents/skills/agentsop-regression-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-regression-gate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-regression-gate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-regression-gate .cursor/skills/agentsop-regression-gate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-regression-gate" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-regression-gate into .cursor/skills/agentsop-regression-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-regression-gate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-regression-gate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-regression-gate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-regression-gate .gemini/skills/agentsop-regression-gate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-regression-gate" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-regression-gate into .gemini/skills/agentsop-regression-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-regression-gate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-regression-gateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-regression-gate .github/skills/agentsop-regression-gate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-regression-gate" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-regression-gate into .github/skills/agentsop-regression-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-regression-gate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-regression-gate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-regression-gate .opencode/skills/agentsop-regression-gate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-regression-gate" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-regression-gate into .opencode/skills/agentsop-regression-gate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-regression-gate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-regression-gateBuild a held-out eval set, run it on every prompt/model change, and block regressions in CI.
Agentsop Regression Gate is an agent skill from agentsope/SkillAlchemy. Build a held-out eval set, run it on every prompt/model change, and block regressions in CI. An LM change is a code change — gate it with a test suite (eval set + metric + threshold). Cross-framework SOP not surfaced by any single base skill.
Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in Business, Finance & HR, covering Operations and SOPs and Test generation. It works with LlamaIndex. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Regression Gate loads about 6.1k tokens when it runs, and up to ~7.9k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 2,999 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 2,999 words, ~6,125 tokens.
.claude/skills/agentsop-regression-gate/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub."Every subsequent change must be gated on these numbers." — Synthesized from [[llamaindex]] Stage 2 (eval loop before optimizing) [llamaindex-sop-skill/SKILL.md:114-126]
"Compiled program beats baseline on a held-out test set (not the val set used in optimization)." — [[dspy]] Stage 3 exit criterion [dspy-sop-skill/SKILL.md:101]
This is an enhancement overlay. The regression-gate SOP exists only as fragments scattered across base skills — [[llamaindex]] OP-10 EvalLoop ("gate every change"), [[dspy]] train/dev/test split + metric — and is never assembled as a standalone cross-framework discipline. It is the discipline that turns a one-off eval into a gate: a test suite that runs in CI on every prompt/model/retriever change and fails the build on regression. It consumes a metric from [[agentsop-metric-design]] and, for domain-specific held-out sets, hands off to [[agentsop-domain-eval-set]].
Activate when any of these is true:
OP-10: "Quantitative regression test for every chunking / embedding / retriever / prompt change" [llamaindex-sop-skill/SKILL.md:235].Do NOT activate for:
[[agentsop-metric-design]] first; if the user refuses, this skill cannot help.You already gate code with unit tests in CI: a change that breaks a test fails the build. A prompt edit, a model swap, a chunk-size tweak are also changes to the system's behavior — but they slip through review because their effect is statistical, not a stack trace. The regression gate is the missing unit test for LM behavior.
The gate is exactly three artifacts plus a wiring step:
eval set metric threshold CI wiring
(held-out QA) (ex,pred)->score (fail if <X / drop>Y) (block merge)
│ │ │ │
└──────────────┴──────────────────┴────────────────────┘
REGRESSION GATEThree load-bearing principles:
The eval set is held out and frozen. It is a labelled, version-controlled fixture that the prompt/model under test has never seen. [[dspy]] is explicit: the compiled program must beat baseline on a held-out test set "not the val set used in optimization" [dspy-sop-skill/SKILL.md:101]. The split is train / dev / test; the gate runs on test only. If the eval set leaks into the prompt (few-shot demos, instructions), the gate measures memorization, not quality.
The metric comes from [[agentsop-metric-design]], not invented here. This skill does not design metrics — it consumes one. A bad metric makes the gate theatre: it will pass changes that hurt users and block changes that help them. The metric must be human-calibrated before it gates anything ([[agentsop-metric-design]] OP-M05).
The threshold is a policy, not a number you guess. Two common shapes: an absolute floor (fail if score < X) and a relative no-regression (fail if score drops > Y from the committed baseline). Relative is the regression gate proper; absolute is a quality bar. Most teams use both: a floor for "never ship below this," plus a no-regression delta for "this PR must not make it worse."
[[llamaindex]] Stage 2 is named "Build the eval loop before optimizing anything" [llamaindex-sop-skill/SKILL.md:114]. The anti-pattern it names is A3: "No eval loop; debug by anecdote" [llamaindex-sop-skill/SKILL.md:348]. The gate is the institutional form of that loop — once it exists, every change is debugged by number, not by vibe.
0. Confirm activation (§1); confirm a metric exists or invoke [[agentsop-metric-design]]
1. BUILD eval set: generate candidates -> curate to a golden set -> freeze + version
2. SPLIT: train / dev / test; the GATE runs on TEST only
3. PICK metric: consume from [[agentsop-metric-design]] (do not invent here)
4. SET threshold: absolute floor AND/OR relative no-regression delta
5. WIRE into CI: run eval on every prompt/model/retriever PR; fail on regression
6. HANDLE flakiness: pin seeds/temp, average N runs, separate flaky from real dropsTwo stages, never one. Generation gives coverage cheaply; curation gives trust.
DatasetGenerator.from_documents(docs).generate_dataset_from_nodes(num=50) produces labelled QA pairs from the documents themselves [llamaindex-sop-skill/SKILL.md:121]. promptfoo and synthetic-data generators do the same for non-RAG tasks.eval/golden_v1.jsonl), tagged with the date and the generator model. Changing it is a versioned event, not an edit.Size: [[dspy]] documents the sweet spot — "30 examples = minimum useful, 300 = recommended" [dspy-sop-skill/SKILL.md:87]. For a held-out gate, 50–200 curated domain examples is the working range; descend to [[agentsop-domain-eval-set]] for domain-specific construction.
The gate runs on test only. Keep test sealed from anything that touches the prompt:
train — feeds optimizers / few-shot demo selection.dev — tuning and threshold-setting.test — the gate. Never used to author prompts, pick demos, or tune. ([[dspy]] held-out exit criterion [dspy-sop-skill/SKILL.md:101].)[[agentsop-metric-design]])Do not invent a metric here. Consume one from [[agentsop-metric-design]]:
OP-10 [llamaindex-sop-skill/SKILL.md:234]).OP-M01/OP-M02/OP-M03).A metric that has not been human-calibrated must not gate ([[agentsop-metric-design]] OP-M05). An uncalibrated gate is worse than no gate — it gives false confidence.
| Threshold shape | Rule | Use when |
|---|---|---|
| Absolute floor | fail if score(test) < X | "never ship below this quality bar" |
| Relative no-regression | fail if baseline − score > Δ | "this PR must not make it worse" (the gate proper) |
| Per-slice floor | fail if any slice (e.g. lexical-query subset) drops > Δ | aggregate hides a regressed minority |
Set Δ above measured run-to-run noise (Stage 6), else the gate flaps. Commit the current test score as baseline.json next to the eval set; the gate compares against it.
eval CI job that runs on every PR touching prompts, model config, retriever/chunking config, or the program graph.baseline.json → exit non-zero on threshold breach → post the before/after table as a PR comment.baseline.json in the same PR (reviewed, not silent).LLM outputs are nondeterministic; a naive gate flaps and gets disabled. Mitigations:
temperature=0 and seeds where the provider supports them; disable response caching in CI ([[dspy]] AP-10: "Forgetting cache=False in stateless deploys" [dspy-sop-skill/SKILL.md:255]).[[agentsop-metric-design]], not papered over by widening Δ.DatasetGenerator.generate_dataset_from_nodes(num=50), promptfoo synthetic generation, or task-specific synthesis. Tag with generator model + date.OP-10 [llamaindex-sop-skill/SKILL.md:234]; external "eval set generation".golden_vN.jsonl).OP-M05; Dilemma 1.test from all prompt-authoring. The gate reads only test. (Note [[dspy]]'s reversed 20/80 train/val split for prompt optimizers [dspy-sop-skill/SKILL.md:96] — that is an optimizer concern; the gate still needs an untouched test slice.)test score; need a pass/fail policy.test score as baseline.json. Define absolute floor X and/or relative no-regression Δ (Δ > measured noise). Optionally per-slice floors.OP-10 regression-test framing [llamaindex-sop-skill/SKILL.md:233-235]; external "llm regression testing CI", "promptfoo".baseline.json to the new test score. Never let CI auto-bump silently.program.gpt4o.json and program.llama8b.json; A/B" [dspy-sop-skill/SKILL.md:191] (versioned-artifact discipline).[[agentsop-metric-design]]) or widen Δ — never disable the gate.cache=False AP-10 [dspy-sop-skill/SKILL.md:255]; [[agentsop-metric-design]] judge-bias hardening; [[dspy]] Stage 2 exit "stable across two runs" [dspy-sop-skill/SKILL.md:91].困境: Team needs an eval set fast. DatasetGenerator produces 200 QA pairs in minutes ([[llamaindex]] Stage 2 [llamaindex-sop-skill/SKILL.md:121]). Hand-curating 200 examples costs days of human time. Ship the generated set as the gate, or pay for curation?
约束: The generator is the same model family that powers the pipeline → its questions are answerable by exactly the kind of reasoning the pipeline already does (self-preference leakage). Generated sets skew easy and miss the long-tail failures users actually hit. But zero eval set means shipping blind (anti-pattern).
决策步骤:
OP-02): a human keeps the good items, fixes labels, drops the trivially-easy and the ambiguous, and injects known production failures and adversarial/edge cases the generator never proposes.OP-M04).结果: The golden set is the gate; the generated pool is scaffolding. Teams that gate on raw generated sets ship regressions that the easy set never exercised — the gate was green while users churned.
可提取的操作: OP-01 GenerateEvalCandidates, OP-02 CurateGoldenSet. Lesson: generation buys coverage, curation buys trust. A gate needs trust — never gate on an uncurated generated set.
困境: A no-regression gate is set at Δ = 0 (any drop fails). A genuinely good refactor — simpler prompt, 40% cheaper model — scores 0.81 vs the 0.83 baseline: a 2-point drop within run-to-run noise. The gate blocks a change that is net-positive (equal quality, far cheaper). The team starts overriding the gate, and soon ignores it entirely.
约束: Run-to-run noise on this judge-based metric is ±1.5 points (measured across 3 reruns). The 2-point "drop" is statistically indistinguishable from noise. A gate that flags noise as regression trains the team to bypass it — a bypassed gate is worse than none.
决策步骤:
OP-07): rerun the baseline config N times; compute the standard deviation. Here σ ≈ 1.5pp.[[agentsop-cost-tiered-models]]). Don't let a quality gate block a cost win that doesn't hurt quality.[[agentsop-metric-design]] (decompose, length penalty, cross-family judge), don't widen Δ to infinity.结果: Δ tuned to ~2σ passes the cheaper-equal-quality change, still catches real regressions (a 6pp drop), and the team keeps trusting the gate. A gate calibrated to noise survives; a Δ=0 gate gets disabled.
可提取的操作: OP-04 SetRegressionThreshold, OP-07 StabilizeFlakyEval. Lesson: the threshold must clear measured noise. A gate that flags noise as failure gets bypassed, and a bypassed gate protects nothing.
| # | Anti-pattern | Why it's wrong | Fix |
|---|---|---|---|
| AP-1 | No eval set; ship blind | Every prompt/model change is an uncontrolled experiment on users; "it got worse" is discovered in production | Build a held-out gate ([[llamaindex]] A3 [llamaindex-sop-skill/SKILL.md:348]) |
| AP-2 | Eval set leaks into the prompt (few-shot demos / instructions drawn from test) | The gate measures memorization, not generalization; green build, real regression | Seal test; demos come from train only (OP-03; [[dspy]] AP-9 [dspy-sop-skill/SKILL.md:254]) |
| AP-3 | Gate on a raw generated set | Inherits generator blind spots; too easy; misses real failures | Curate a golden set (OP-02; Dilemma 1) |
| AP-4 | Gate on an uncalibrated metric | A wrong metric passes harmful changes and blocks good ones — gate is theatre | Calibrate via [[agentsop-metric-design]] OP-M05 before gating |
| AP-5 | Δ = 0 / threshold below noise | Gate flaps on noise, team bypasses it | Set Δ > 2σ measured noise (OP-04, OP-07; Dilemma 2) |
| AP-6 | Run the gate on the val/dev set used for tuning | Optimistic, leaks tuning into evaluation | Gate on held-out test only ([[dspy]] [dspy-sop-skill/SKILL.md:101]) |
| AP-7 | Caching on in CI | Stale cached outputs mask the change under test | cache=False ([[dspy]] AP-10 [dspy-sop-skill/SKILL.md:255]) |
| AP-8 | Aggregate-only gate | A win on the majority hides a regressed minority slice | Per-slice gating (OP-08) |
| AP-9 | Silent baseline auto-bump | Quality can ratchet down unnoticed if CI rewrites baseline | Bump baseline only in a reviewed PR (OP-06) |
| AP-10 | Disabling the gate when it flakes | Removes the only protection; flakiness is a metric/Δ bug, not a gate bug | Stabilize (OP-07), never disable |
[[agentsop-metric-design]] first.lm-evaluation-harness. For a domain-specific held-out set, descend to [[agentsop-domain-eval-set]].| Concept | LlamaIndex | DSPy | promptfoo | LangSmith | This skill |
|---|---|---|---|---|---|
| Eval set generation | DatasetGenerator.generate_dataset_from_nodes(num=N) [llamaindex-sop-skill/SKILL.md:121] | bring labelled examples; BootstrapFewShot self-generates demos (not the test set) | tests: synthesis / generate from prompts | Datasets created from traces / uploads | OP-01 GenerateEvalCandidates |
| Golden / curated set | manual review of generated QA | hand-labelled trainset/devset | curated tests YAML with assert | curated Dataset + reference outputs | OP-02 CurateGoldenSet |
| Train/dev/test split | manual | explicit; reversed 20/80 for prompt optimizers, held-out test for gate [dspy-sop-skill/SKILL.md:96,101] | n/a (test set is the suite) | dataset splits | OP-03 SplitTrainDevTest |
| Metric | Faithfulness/Relevancy/RetrieverEvaluator(mrr,hit_rate) [llamaindex-sop-skill/SKILL.md:234] | def metric(ex,pred,trace=None)->bool|float [dspy-sop-skill/SKILL.md:88] | assert (equals/contains/llm-rubric/javascript) | evaluator fns / LLM-as-judge | consumed from [[agentsop-metric-design]] |
| Threshold / gate | manual (gate every change [llamaindex-sop-skill/SKILL.md:124]) | held-out beats baseline "by ≥ delta" [dspy-sop-skill/SKILL.md:101] | assert pass + --fail-on thresholds | rules + alerts on eval scores | OP-04 SetRegressionThreshold |
| CI wiring | not built-in (DIY job around Evaluate) | not built-in (DIY around dspy.Evaluate) | first-class: promptfoo eval in CI, non-zero exit | CI integration + regression alerts | OP-05 WireCIGate |
| Flaky handling | run multiple times | cache=False; "stable across two runs" [dspy-sop-skill/SKILL.md:91,255] | repeat + threshold | run aggregation | OP-07 StabilizeFlakyEval |
Combination patterns:
DatasetGenerator for OP-01, the Faithfulness/Relevancy/Retriever triad as the metric, wired into a DIY CI job. The base skill names the loop ("gate every change"); this skill makes it a CI gate.promptfoo eval exits non-zero on failed asserts; it is the closest off-the-shelf realization of OP-05. Use it as the runner; still bring a curated set (OP-02) and a calibrated metric.Opinionated default: build the eval set with the base framework's generator (OP-01), curate by hand (OP-02), keep the metric in [[agentsop-metric-design]], and run the gate with promptfoo (CI-native) or a thin script around dspy.Evaluate / LlamaIndex evaluators. The gate, the metric, and the eval set are three separable, version-controlled artifacts — never one tangled blob.
references/R1-source-evidence.md — every cited claim resolved to a source lineintermediate/operation_candidates.json — machine-readable operation registryCitations: [[llamaindex]] OP-10 EvalLoop / Stage 2 [llamaindex-sop-skill/SKILL.md:114-126,232-236,348]; [[dspy]] Stage 2-3 split+metric+held-out [dspy-sop-skill/SKILL.md:85-105,137,191,254-255]; [[agentsop-metric-design]]; [[agentsop-domain-eval-set]]; external "llm regression testing CI", "promptfoo", "eval set generation".
© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/agentsop-regression-gate of agentsope/SkillAlchemy.
Open the folder on GitHubat commit 6ea799f
Agentsop Regression Gate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Regression Gate this skillagentsope/SkillAlchemy | 459 | — | ~6.1k | Automated safety check: Pass | MIT | |
| Tulingresearch Fusion Campaign Enginemeamaturinlove221/TuringResearch_plus | 136 | — | ~548 | Automated safety check: Pass | Custom licence | |
| Cc Sdd New Agentgotalab/cc-sdd | 3.7k | — | ~1.1k | Automated safety check: Pass | MIT | |
| DBS Business Toolkit Entrydontbesilent2025/dbskill | 11k | — | ~2k | Automated safety check: Pass | Custom licence | |
| Agent Sop Authorstrands-agents/agent-sop | 1.2k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Diffusion Narrative Denouncingcanwhite/Krebs | 1k | — | ~831 | Automated safety check: Pass | MIT |
meamaturinlove221/TuringResearch_plus
A skill your agent uses when maintaining Campaign - Strategy - Tactic - SOP runtime behavior.
gotalab/cc-sdd
Add or extend coding-agent support in cc-sdd by executing the SOP in docs/cc-sdd/sop-new-agent.md end-to-end.
dontbesilent2025/dbskill
Chinese-language entry skill for the dontbesilent business toolkit: onboards new users, orchestrates tasks across sub-skills, runs numbered prompts and lists hidden ones.
strands-agents/agent-sop
Create (or update) and validate Agent SOPs (Standard Operating Procedures) - markdown-based workflows that guide AI agents through complex, multi-step tasks with RFC 2119 constraints.
canwhite/Krebs
基于"扩散模型叙事去噪流"的小说写作 SOP。将 AI 视为去杂质机器,通过锁定全局信号、预测叙事噪声、精准去噪、随机修正四个步骤,解决 AI 翻译腔、逻辑断层和故事平淡的问题。
0xenzyme/polanyi-skill
Michael Polanyi 的思维框架。用 Polanyi 视角分析隐性知识、技能习得、经验传承、师徒制、 知识管理、学习方法、AI/工具替代边界、科学共同体与后批判哲学问题。
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Works with
Categories
Build a held-out eval set, run it on every prompt/model change, and block regressions in CI. Agentsop Regression Gate is an agent skill from agentsope/SkillAlchemy. Build a held-out eval set, run it on every prompt/model change, and block regressions in CI.
Agentsop Regression Gate fits situations like: tasks that involve Operations and SOPs; tasks that involve Test generation.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a claude-code`. Or copy the skill folder (skills/agentsop-regression-gate in agentsope/SkillAlchemy) into .claude/skills/agentsop-regression-gate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a codex`. Or copy the skill folder (skills/agentsop-regression-gate in agentsope/SkillAlchemy) into .agents/skills/agentsop-regression-gate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-regression-gate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-regression-gate, .gemini/skills/agentsop-regression-gate, .github/skills/agentsop-regression-gate and .opencode/skills/agentsop-regression-gate in your project.
SKILL.md names no scripts, command-line tools or credentials: Agentsop Regression Gate is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Regression Gate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.1k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Regression Gate: Tulingresearch Fusion Campaign Engine (meamaturinlove221/TuringResearch_plus, 136 stars), Cc Sdd New Agent (gotalab/cc-sdd, 3.7k stars), DBS Business Toolkit Entry (dontbesilent2025/dbskill, 11k stars) and Agent Sop Author (strands-agents/agent-sop, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 459 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.