Mg CLI
modelguide/modelguide
Generate YAML configuration files and run CLI commands to onboard organizations into ModelGuide.
Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-domain-eval-set --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-domain-eval-set .claude/skills/agentsop-domain-eval-set && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-domain-eval-set" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-domain-eval-set into .claude/skills/agentsop-domain-eval-set/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-domain-eval-set", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-domain-eval-setType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-domain-eval-set --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-domain-eval-set .agents/skills/agentsop-domain-eval-set && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-domain-eval-set" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-domain-eval-set into .agents/skills/agentsop-domain-eval-set/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-domain-eval-set", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-domain-eval-set --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-domain-eval-set .cursor/skills/agentsop-domain-eval-set && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-domain-eval-set" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-domain-eval-set into .cursor/skills/agentsop-domain-eval-set/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-domain-eval-set", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-domain-eval-set--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-domain-eval-set --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-domain-eval-set .gemini/skills/agentsop-domain-eval-set && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-domain-eval-set" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-domain-eval-set into .gemini/skills/agentsop-domain-eval-set/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-domain-eval-set", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-domain-eval-setInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-domain-eval-set .github/skills/agentsop-domain-eval-set && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-domain-eval-set" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-domain-eval-set into .github/skills/agentsop-domain-eval-set/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-domain-eval-set", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-domain-eval-set --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-domain-eval-set .opencode/skills/agentsop-domain-eval-set && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-domain-eval-set" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-domain-eval-set into .opencode/skills/agentsop-domain-eval-set/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-domain-eval-set", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-domain-eval-setBuild and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic.
Agentsop Domain Eval Set is an agent skill from agentsope/SkillAlchemy. Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic. Distinct from public benchmarks (MMLU/HumanEval/GSM8K via lm-evaluation-harness) which measure GENERAL capability. Only a held-out domain set predicts whether THIS system works on YOUR data. Collect real examples, label, hold out (never train/prompt on it), size 50-200, version it, refresh on drift.
Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in AI & LLM Engineering, covering LLM evaluation and Operations and SOPs. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Domain Eval Set loads about 6.3k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 105 tokens; SKILL.md has 3,052 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 3,052 words, ~6,269 tokens.
.claude/skills/agentsop-domain-eval-set/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub."Compiled program beats baseline on a held-out test set (not the val set used in optimization)." — DSPy SOP exit criterion [dspy.ai/learn/optimization/overview/]
"Build the eval loop before optimizing anything. Every subsequent change must be gated on these numbers." — LlamaIndex SOP Stage 2
This is an ENHANCE overlay skill. It produces one artifact — a versioned,
sealed, human-labeled set of 50–200 examples drawn from your domain — that
other skills consume: [[agentsop-regression-gate]] enforces it on every PR,
[[agentsop-metric-design]] defines the scoring function applied to each example, and
[[lm-evaluation-harness]] runs the complementary public-capability axis. The
core claim: public benchmarks tell you the model is smart in general; only a
held-out domain set tells you it works on your task. The latter is the one that
predicts production.
Activate when any of these is true:
[[agentsop-regression-gate]] can do its job.Do NOT activate for:
[[lm-evaluation-harness]], not this skill.Two orthogonal axes, constantly confused:
| Axis | What it measures | Tool | Predicts production? |
|---|---|---|---|
| General capability | Reasoning, knowledge, coding in general, on shared public tasks | [[lm-evaluation-harness]] (MMLU, HumanEval, GSM8K, TruthfulQA) | No — a proxy at best |
| Domain task fit | Whether the system answers your users on your data | this skill (held-out domain set) | Yes — this is the signal |
A model can score 90% on MMLU and 40% on your insurance-claims triage. A model can score below SOTA on HumanEval and be perfect at your internal codebase's patterns. The public number and the domain number are nearly uncorrelated once you're past a basic capability floor. The public bench is a sanity check; the domain set is the decision.
Three corollaries (each maps to an SOP stage):
Real beats synthetic. The set is sampled from real domain traffic
(tickets, queries, logs, transactions), stratified, with edge cases pulled
deliberately. Auto-generated QA pairs (LlamaIndex DatasetGenerator) are a
fine bootstrap, but a model can ace generated questions and still fail real
user phrasing. Generated sets do not replace a real held-out set (§7).
Held out means SEALED. The held-out split is never shown to the optimizer, never pasted into a prompt as a few-shot demo, never used to pick chunk size or reranker, never in the fine-tune data. The moment it leaks, the number is inflated and meaningless (AP-2, OP-DE07). Per DSPy: the test set must be distinct from the val set used in optimization [dspy.ai/learn/optimization/overview/].
Small but significant. 50–200 examples. Below ~30 you are "memorizing, not training" [dspy.ai/learn/optimization/overview/] and differences are noise. The set is small enough to label by hand and large enough to detect ~5–10pp regressions and to slice by segment.
0. Confirm activation (§1) — is the question "does this work on OUR data"?
1. COLLECT — sample real domain examples; stratify; pull edge cases (OP-DE01)
2. LABEL — gold answer / reference / pass-fail; 2 annotators on subset (OP-DE02)
3. HOLD OUT — split train/dev/test; SEAL the test split (OP-DE03)
4. SIZE — land at 50-200; per-segment counts (OP-DE04)
5. VERSION — hash + date + rubric; freeze as an artifact (OP-DE05)
6. LEAK-AUDIT — diff held-out vs demos / train / fine-tune data (OP-DE07)
7. PAIR — report alongside public bench; gate on the domain set (OP-DE08)
(later) REFRESH on domain shift (OP-DE06)Pull from where the real distribution lives: support tickets, search/query logs, user transcripts, transaction records, bug reports. Stratify so the set covers the production mix — by query type (lookup / summary / compare), by segment (tenant, language, product area), by difficulty. Then deliberately over-sample edge cases and known failures — the head of the distribution is easy; the tail is where systems break.
Target a raw pool ≥ 2× the final size (you'll drop ambiguous items in labeling). Record provenance and timestamp per example (needed later for drift refresh).
Exit: a candidate pool ≥ 2× target, with provenance, spanning the real mix.
Attach ground truth per example: a gold answer, an acceptable reference
response (not "the unique correct" one for open-ended tasks — see
[[agentsop-metric-design]]), or a pass/fail label. For RAG, also label the gold
passage so RetrieverEvaluator(["mrr","hit_rate"]) can run [LlamaIndex OP-10].
Have two annotators label a subset, measure agreement, resolve disagreements, and drop genuinely ambiguous items — an example two experts can't agree on will only add noise. Record the rubric. (This is the data-side analogue of DSPy's "human-validate the metric on ≥20 spot-checks" discipline [DSPy Case C].)
Exit: labeled set with inter-annotator agreement noted, rubric recorded, ambiguous items logged as rejected.
Split into train / dev / test. The test (held-out) split is sealed:
Store it in a separate file/location with an access note. Per DSPy, the exit-gate test set must be "distinct from the val set used in optimization" [dspy.ai/learn/optimization/overview/]. The dev split is what you tune against; the test split is the one number you trust at decision time.
Exit: sealed held-out test split + train/dev splits; access policy written.
Size up (toward 200, or split into per-segment sets each ~50) when you need
per-segment confidence. LlamaIndex's DatasetGenerator default of num=50 sits
at the low end of this band — fine to bootstrap, then curate.
Freeze the set as a versioned artifact — eval_v1.jsonl plus a manifest with
a content hash, creation date, and the labeling rubric. Score every
model / prompt / retriever change against the same version; keep a results
table keyed by (eval_version, system_version); bump only on a deliberate
refresh, never silently. DSPy ships program.json as a versioned artifact
[dspy.ai/tutorials/saving/]; LlamaIndex versions indices as deployment artifacts
(SOP Stage 5) — the eval set deserves the same rigor.
Before any release, and whenever few-shot demos or fine-tune data are assembled,
diff the held-out set against (a) prompt few-shot demos, (b) fine-tune /
training data, (c) the optimizer trainset. Any overlap = contamination → the
held-out number is inflated and worthless (AP-2). Remove the overlap or rebuild
the split — the same provenance discipline as [[agentsop-metric-design]]'s calibration
receipt (OP-M10).
Run [[lm-evaluation-harness]] for the capability floor (sanity check: is the
model fundamentally competent?). Run the domain held-out set for the decision.
Report both side by side. If they disagree, the domain set wins the go/no-go.
Hand the sealed set to [[agentsop-regression-gate]] to enforce on every subsequent PR.
Domains drift: new product line, new user segment, seasonal change. When held-out scores stop tracking production complaints, refresh (OP-DE06): add fresh real examples from recent traffic, retire stale ones, re-label edge cases production surfaced, bump the version, keep the old version for back-comparison. Cadence: quarterly or on any major domain change, whichever comes first. (This mirrors LlamaIndex's live-corpus reconciliation, A10.)
Each operation: Trigger → Action → Output [Evidence]. Full Trigger/Action/
Output/Evidence form in intermediate/operation_candidates.json.
OP-DE01 SourceFromRealTraffic — No curated set, traffic available → sample real inputs (logs/tickets/queries/transactions), stratify by type/segment/ difficulty, over-sample edge cases → raw pool ≥2× target with provenance. [DSPy dev-set discipline; LlamaIndex OP-10 eval-from-corpus]
OP-DE02 LabelAndCurate — Raw pool collected → attach gold/reference/pass-fail
per item; two annotators on a subset, resolve disagreement, drop ambiguous,
record rubric; for RAG label the gold passage → curated labeled set with
agreement noted. [DSPy Case C ≥20 spot-checks; LlamaIndex RetrieverEvaluator]
OP-DE03 HoldOutDiscipline — Set about to be used → split train/dev/test; seal the test split (never to optimizer, never as few-shot demo, never to pick chunking/reranker/model, never in fine-tune data) → sealed test + train/dev. [DSPy "held-out distinct from val"; Case A step 4]
OP-DE04 SizeFor50to200 — Deciding size → target 50–200 (50 = coarse signal; 100–200 = detect ~5–10pp regressions + per-segment slices; <30 = noise) → sized set with per-segment counts. [DSPy "30 min, 200+ for MIPROv2"; LlamaIndex num=50]
OP-DE05 VersionTheSet — Set finalized → freeze as eval_v1.jsonl + manifest
(hash, date, rubric); score every change vs the same version; results keyed by
(eval_version, system_version); bump only on deliberate refresh → versioned
artifact. [DSPy program.json versioning; LlamaIndex versioned indices]
OP-DE06 RefreshOnDomainShift — Domain drifts; scores stop tracking complaints → add fresh recent-traffic examples, retire stale, re-label edge cases, bump version, keep old for comparison (quarterly or on major change) → new version + drift log. [LlamaIndex live-corpus reconciliation A10]
OP-DE07 LeakAudit — Before release / when demos or fine-tune data assembled → diff held-out vs few-shot demos, fine-tune data, optimizer trainset; any overlap = contamination → remove or rebuild → leak-audit report (0 overlap). [DSPy held-out-distinct rule; metric-design provenance OP-M10]
OP-DE08 PairWithPublicBench — Public-bench number used to justify deployment →
treat public bench as capability floor/sanity check, require the domain held-out
set as the decision gate; report both, on disagreement the domain set wins →
two-axis report gated on domain. [[[lm-evaluation-harness]] covers public, not
your domain]
困境: A team wants to ship a contract-review assistant. They have thousands of
raw contracts but only ~25 examples a lawyer has labeled with gold answers.
25 < the 50 floor and well below the 30 "memorizing, not training" line
[dspy.ai/learn/optimization/overview/]. They're tempted to (a) skip the held-out
set and ship on MMLU/legal-bench numbers, or (b) auto-generate 200 QA pairs with
LlamaIndex DatasetGenerator and call that the held-out set.
约束: Lawyer labeling time is the bottleneck (~$$/hour, scarce). Public legal benchmarks exist but don't reflect this firm's contract templates. Auto-generated questions risk testing "what the corpus says" rather than "what real reviewers ask".
决策步骤:
DatasetGenerator (LlamaIndex Stage 2) gives a cheap dev set for iteration —
but it is synthetic, so it cannot be the trusted held-out number (§7 caveat).结果: A 50-example human-labeled, sealed held-out set built from the hardest real contracts predicts production far better than 200 synthetic questions or any public legal benchmark. The synthetic set still earns its keep — as the dev set you tune against, never as the number you trust.
可提取的操作: OP-DE01, OP-DE02, OP-DE03, OP-DE04. Lesson: spend scarce
labels on a small REAL held-out set; let synthetic generation cover the dev set;
never let a public bench be the gate.
困境: A support-triage classifier shows 0.91 on eval_v1 (built 9 months ago)
and every PR passes [[agentsop-regression-gate]]. Yet production accuracy collapsed and
users are escalating. The eval set says everything is fine.
约束: eval_v1 is versioned and trusted; nobody wants to "move the goalposts".
The domain shifted — a new product line generates a third of current tickets, and
none of those ticket types existed when eval_v1 was built. Rebuilding costs
annotator time.
决策步骤:
eval_v1's segment counts. The new product line is ~33% of live
traffic and 0% of the eval set → the eval set no longer represents the
domain. The green score is measuring an obsolete distribution.eval_v2.eval_v1 for back-comparison. Re-score the
current system on eval_v2: it drops to 0.63 — now matching reality.[[agentsop-regression-gate]] on eval_v2. Add a drift check to the
refresh cadence: quarterly, compare live segment mix vs eval segment mix; if any
segment drifts >X%, trigger a refresh.结果: The "green-but-on-fire" gap was a stale held-out set, not a model
regression. A versioned refresh (eval_v2) restored the eval as a true production
predictor; the back-comparison against eval_v1 documented exactly how much the
domain moved.
可提取的操作: OP-DE06 RefreshOnDomainShift, OP-DE05 VersionTheSet. Lesson:
a held-out set is a snapshot of a moving distribution. Schedule drift checks; an
old green score can be the most dangerous number you have.
| # | Anti-pattern | Why it's wrong | Fix |
|---|---|---|---|
| AP-1 | Public bench as proxy for domain performance ("92% MMLU → ship it") | Public benches measure general capability; near-uncorrelated with task fit past a floor | Build a domain held-out set; gate on it (OP-DE08) |
| AP-2 | Eval set leaks into prompt / training / trainset | Held-out number is inflated and meaningless; you're testing on the train set | Seal it; leak-audit before release (OP-DE03, OP-DE07) |
| AP-3 | Set too small to be significant (<30 examples) | "Memorizing, not training" [dspy.ai/learn/optimization/overview/]; variance swamps signal | Target 50–200 (OP-DE04) |
| AP-4 | Synthetic-only held-out (auto-generated QA is the test set) | Tests "what the corpus says", not real user phrasing; flatters the system | Synthetic = dev set bootstrap only; real-labeled = held-out (§7, Dilemma 1) |
| AP-5 | Never refreshing as the domain drifts | Green scores on an obsolete distribution; "green but on fire" (Dilemma 2) | Schedule drift checks; refresh + version (OP-DE06) |
| AP-6 | Unversioned set silently edited | Can't compare across system versions; results table is meaningless | Hash + date + rubric; bump on deliberate refresh (OP-DE05) |
| AP-7 | No stratification / edge cases (only easy head-of-distribution) | Passes eval, fails the tail where systems actually break | Stratify by segment/type; over-sample edge cases (OP-DE01) |
| AP-8 | Tuning chunk size / reranker / model against the held-out set | That makes it a val set, not held-out; the trust is gone | Tune on dev; touch held-out only at decision time (OP-DE03) |
[[lm-evaluation-harness]] (MMLU/GSM8K/etc.),
not this skill. This skill is for your task, not the leaderboard.[[agentsop-metric-design]] to define a defensible, calibrated scoring function.When does each kind of eval set apply? They are complementary axes, not substitutes — a mature pipeline uses all three.
| Concept | Held-out domain set (this skill) | [[lm-evaluation-harness]] (public) | LlamaIndex DatasetGenerator (synthetic) |
|---|---|---|---|
| What it measures | Task fit on your data | General capability | Coverage of your corpus's content |
| Data source | Real traffic, human-labeled | Public academic datasets (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag) | LLM-generated QA from your docs |
| Size | 50–200 | thousands (fixed by benchmark) | arbitrary (default num=50) |
| Contamination risk | You control it (leak-audit) | High — public benches leak into pretraining | Low (your private corpus) but synthetic |
| Predicts production? | Yes (the decision gate) | No (capability floor / sanity check) | Partially (dev-set iteration, not the gate) |
| When to use | Go/no-go on shipping to your users; per-PR regression gate | Model selection on raw capability; academic reporting; training-progress tracking | Bootstrap a dev set fast before you've labeled real data |
| Invocation | eval_vN.jsonl + scoring fn from [[agentsop-metric-design]] | lm_eval --tasks mmlu,gsm8k,... | DatasetGenerator.from_documents(docs).generate_dataset_from_nodes(num=50) |
Decision rubric:
Q1. Are you deciding whether to SHIP / SWITCH on YOUR users' data?
YES → held-out domain set is the gate (this skill). Public bench = sanity check only.
Q2. Are you comparing raw model capability or reporting academic numbers?
YES → lm-evaluation-harness (MMLU/HumanEval/GSM8K). Not this skill.
Q3. Do you have NO real labeled data yet but a corpus exists?
YES → DatasetGenerator to bootstrap a DEV set; label real held-out as soon as traffic appears.
Q4. Is there an objective oracle (tests/schema/exact-match)?
YES → run the oracle; no curated set needed.
DEFAULT → build + version a 50-200 real held-out set; gate via [[agentsop-regression-gate]];
score via [[agentsop-metric-design]]; pair with [[lm-evaluation-harness]] for the floor.Combination patterns:
[[agentsop-regression-gate]]: this skill produces the sealed,
versioned set; regression-gate enforces it on every PR (chunking / embedding /
prompt / model change). Division of labor: produce vs enforce.[[agentsop-metric-design]]: this skill defines what's in the set;
metric-design defines how each example is scored (decomposed sub-judges,
bool-during-compile/float-during-eval, human-calibrated, length-penalized).
A set with no defensible scoring function is half a benchmark.[[lm-evaluation-harness]]: report both axes side by side
(OP-DE08). Public bench answers "is the model competent?"; the domain set
answers "does it work for us?". On disagreement, the domain set wins go/no-go.Opinionated default: build the held-out set in plain jsonl (transparent,
diffable, hashable), label it with humans on the hardest real examples, seal it,
version it, and treat the public-benchmark number as a sanity check you report
but never gate on.
references/R1-source-evidence.md — verbatim source quotes (DSPy held-out
discipline, LlamaIndex eval-loop, lm-evaluation-harness public scope)intermediate/operation_candidates.json — 8 operations in Trigger / Action /
Output / Evidence formCross-links: [[lm-evaluation-harness]] (public-benchmark axis),
[[agentsop-regression-gate]] (per-PR enforcement), [[agentsop-metric-design]] (scoring function).
Citations: [dspy.ai/learn/optimization/overview/], [dspy.ai/learn/optimization/optimizers/], [dspy.ai/learn/evaluation/metrics/], [dspy.ai/tutorials/saving/], [developers.llamaindex.ai/python/framework-api-reference/evaluation/], [llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5], ~/.claude/skills/lm-evaluation-harness/SKILL.md.
© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/agentsop-domain-eval-set of agentsope/SkillAlchemy.
Open the folder on GitHubat commit 6ea799f
Agentsop Domain Eval Set next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Domain Eval Set this skillagentsope/SkillAlchemy | 459 | — | ~6.3k | Automated safety check: Pass | MIT | |
| Mg CLImodelguide/modelguide | 108 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Harness DesignAnastasiyaW/codex-claude-code-config | 154 | — | ~764 | Automated safety check: Pass | MIT | |
| Yao Meta Skillyaojingang/yao-meta-skill | 2.7k | — | ~768 | Automated safety check: Pass | MIT | |
| Evaluating With Leakage Gatesmaziyarpanahi/openmed | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT |
modelguide/modelguide
Generate YAML configuration files and run CLI commands to onboard organizations into ModelGuide.
AnastasiyaW/codex-claude-code-config
Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…
yaojingang/yao-meta-skill
Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts.
maziyarpanahi/openmed
Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic. Agentsop Domain Eval Set is an agent skill from agentsope/SkillAlchemy. Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic.
Agentsop Domain Eval Set fits situations like: tasks that involve LLM evaluation; tasks that involve Operations and SOPs.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a claude-code`. Or copy the skill folder (skills/agentsop-domain-eval-set in agentsope/SkillAlchemy) into .claude/skills/agentsop-domain-eval-set in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a codex`. Or copy the skill folder (skills/agentsop-domain-eval-set in agentsope/SkillAlchemy) into .agents/skills/agentsop-domain-eval-set in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-domain-eval-set -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-domain-eval-set, .gemini/skills/agentsop-domain-eval-set, .github/skills/agentsop-domain-eval-set and .opencode/skills/agentsop-domain-eval-set in your project.
SKILL.md names no scripts, command-line tools or credentials: Agentsop Domain Eval Set is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Domain Eval Set is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Domain Eval Set: Mg CLI (modelguide/modelguide, 108 stars), Agent Harness Design (AnastasiyaW/codex-claude-code-config, 154 stars), Yao Meta Skill (yaojingang/yao-meta-skill, 2.7k stars) and Evaluating With Leakage Gates (maziyarpanahi/openmed, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 459 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.