Mg CLI
modelguide/modelguide
Generate YAML configuration files and run CLI commands to onboard organizations into ModelGuide.
The compile-readiness gate for prompt auto-optimization. An agent skill from agentsope/SkillAlchemy.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-compilation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-prompt-compilation .claude/skills/agentsop-prompt-compilation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-prompt-compilation" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-compilation into .claude/skills/agentsop-prompt-compilation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-compilation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-compilationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-compilation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-prompt-compilation .agents/skills/agentsop-prompt-compilation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-prompt-compilation" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-compilation into .agents/skills/agentsop-prompt-compilation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-compilation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-compilation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-prompt-compilation .cursor/skills/agentsop-prompt-compilation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-prompt-compilation" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-compilation into .cursor/skills/agentsop-prompt-compilation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-compilation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-prompt-compilation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-compilation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-prompt-compilation .gemini/skills/agentsop-prompt-compilation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-prompt-compilation" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-compilation into .gemini/skills/agentsop-prompt-compilation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-compilation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-compilationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-prompt-compilation .github/skills/agentsop-prompt-compilation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-prompt-compilation" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-compilation into .github/skills/agentsop-prompt-compilation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-compilation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-prompt-compilation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-prompt-compilation .opencode/skills/agentsop-prompt-compilation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-prompt-compilation" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-prompt-compilation into .opencode/skills/agentsop-prompt-compilation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-prompt-compilation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-prompt-compilationThe compile-readiness gate for prompt auto-optimization. An agent skill from agentsope/SkillAlchemy.
Agentsop Prompt Compilation is an agent skill from agentsope/SkillAlchemy. The compile-readiness gate for prompt auto-optimization. Decide whether you have earned the right to run an optimizer (DSPy MIPROv2 / GEPA / BootstrapFewShot) before spending compute. Two preconditions only — a real metric, and enough examples for the optimizer you picked. Garbage metric in, garbage prompt out. Pick the optimizer by data scale; GEPA inverts the scale assumption (~10 examples + textual feedback).
Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in AI & LLM Engineering, covering Operations and SOPs. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Prompt Compilation loads about 6.3k tokens when it runs, and up to ~8.3k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 2,764 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 2,764 words, ~6,316 tokens.
.claude/skills/agentsop-prompt-compilation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub."It's unproductive to launch optimization runs using a poorly designed program or a bad metric." — DSPy core team [dspy.ai/learn/optimization/overview/]
"Compile when you can measure. The optimizer maximizes your metric — garbage metric in, garbage prompt out." — this skill's operating principle (synthesized from the line above + DSPy Case C)
This is an enhancement-overlay decision skill. It answers exactly one question the broad [[dspy]] library
skill buries under API surface: have you earned the right to run an optimizer yet, and which one? It produces a
go / no-go gate plus an optimizer pick. It defers every implementation detail — Signature syntax, module choice,
compile() calls, save/deploy — to [[dspy]] and the full workflow in [[agentsop-dspy]]. It defers metric construction
to [[agentsop-metric-design]]; this skill only checks the metric exists and is validated, then uses it as the gate.
The trap it removes: people reach for MIPROv2(auto="heavy") because the API is right there, before they have a
metric worth maximizing or enough data to avoid memorization. Compilation is a hyperparameter search costing
hundreds-to-thousands of LM calls ($2–$40+, minutes-to-hours) [dspy.ai/faqs/]. Spending that on an un-validated
metric or 8 examples is pure waste.
Activate when all three of these are plausibly true (the gate then confirms them):
[[agentsop-dspy]] §1: "the team manually tunes few-shot examples; a metric exists
but isn't being used to drive prompt design."metric(example, pred) -> bool|float — or one can be written
and human-validated. Without this, do not activate; the optimizer has nothing to maximize.Concrete triggers in intent or codebase:
| Trigger | Signal |
|---|---|
| Spend intent | "auto-tune this prompt", "should I run MIPRO?", "is it worth compiling?", "GEPA vs MIPROv2", "optimize prompts for our metric" |
| API reach | MIPROv2(, BootstrapFewShot(, dspy.GEPA(, teleprompter, optimizer.compile( about to be called |
| Symptom | hand-tuned prompt stuck; few-shot examples curated by hand; metric written but only used for reporting, not optimization |
Do NOT activate when:
[[agentsop-metric-design]] first; if the user refuses any success criterion, the gate stays closed.[[agentsop-dspy]] Stage 1 and
the signature-design overlay. Promoting prose to a typed Signature and compiling it are two different gates.An optimizer (MIPROv2, GEPA, BootstrapFewShot) is a black-box search over prompt instructions + few-shot demos that
maximizes metric(pred, example). It has no taste. It will faithfully chase whatever the metric rewards, biases
and all. From [[agentsop-dspy]] Case C: "DSPy will optimize toward whatever the metric rewards. A bad metric becomes a
bad program at scale." Two corollaries make this a gate, not a step:
The metric is the precondition, not a tunable. Before any compute is spent, the metric must (a) exist and
(b) agree with human judgment on ≥20 spot-checks. An un-validated metric means the expensive search optimizes
the metric's blind spot. This is non-negotiable [dspy.ai/learn/evaluation/metrics/; Case C]. Construction is
[[agentsop-metric-design]]'s job; this skill only checks the receipt.
Data scale is a hard floor, not a preference. Below the optimizer's example floor you are not training, you are memorizing. The DSPy 20/80 train/val split exists because "prompt-based optimizers often overfit to small training sets" [dspy.ai/learn/optimization/overview/]. The floor differs per optimizer (§4.2).
┌──────────────────────────────────────────────┐
Gate 1 │ METRIC: exists? AND human-validated ≥20? │ ── No ─► STOP. Build/validate metric ([[agentsop-metric-design]]).
(measure) └──────────────────────────────────────────────┘
│ Yes
▼
┌──────────────────────────────────────────────┐
Gate 2 │ EXAMPLES: ≥ floor for the optimizer I want? │ ── No ─► Pick a lower-floor optimizer, collect data,
(data) └──────────────────────────────────────────────┘ or STOP (use LabeledFewShot as a floor).
│ Yes
▼
COMPILE (cheap probe first: auto="light")Compilation runs hundreds-to-thousands of LM calls. Reference run: ~3.2k API calls, $2–$3, 6–20 min for
auto="light"; auto="heavy" on 1000+ examples can hit tens of dollars and hours [dspy.ai/faqs/]. Cost scales with
num_trials × |trainset| × |program LM calls|. The gate's entire job is to stop you spending that on a metric or
dataset that cannot pay it back.
0. Confirm activation (§1): plateaued hand-tuning + metric + examples.
1. GATE 1 — metric exists AND validated ≥20 human spot-checks? [hard]
2. GATE 2 — example count ≥ floor for the chosen optimizer? [hard]
3. PICK — choose optimizer by data scale + feedback signal. (§4.2)
4. BUDGET — estimate cost; cap it; choose a cheap optimizer LM.
5. PROBE — run auto="light" (or smallest config) as a signal.
6. DECIDE — gain ≥ threshold → escalate; else go back, don't escalate.metric(example, pred, trace=None) -> bool|float? If not → STOP, route to [[agentsop-metric-design]].[[agentsop-dspy]] Case C; dspy.ai/learn/evaluation/metrics/].[[agentsop-metric-design]]) before compiling [arxiv.org/pdf/2506.02592].Exit: a validated metric callable + a calibration receipt. Otherwise the gate is closed; do not proceed.
Count labeled examples. The DSPy-documented thresholds: ≥30 = minimum useful, ~300 = recommended, 200+ required for MIPROv2 to avoid overfitting [dspy.ai/learn/optimization/overview/]. The floor is per optimizer:
< 10 → only LabeledFewShot(k=8) (no search, weakest) — usually means STOP, collect data.~10+ with textual feedback → GEPA is legal and sample-efficient (this inverts the usual "more data" rule).~30–50 → BootstrapFewShot / BootstrapFewShotWithRandomSearch.200+ → MIPROv2.Exit: the example count clears the floor of the optimizer you intend to run. If not, either drop to a lower-floor optimizer or stop.
Use the table in §4.2. The single most important branch: do you have textual error feedback (test diffs, schema
violations, judge rationales like "answer was verbose")? If yes, GEPA needs only ~10 examples and converges faster
because it reflects on the text of the feedback, not just a scalar [dspy.ai/tutorials/gepa_ai_program/;
arxiv.org/abs/2507.19457]. If no, fall back to the data-scale ladder.
Estimate before launching (§4.4). Cap spend. Use a cheap optimizer LM (e.g. gpt-4o-mini) even when the task LM
is expensive — community-reported parity at a fraction of cost [github.com/stanfordnlp/dspy/issues/1596]. Set
dspy.configure(track_usage=True) to log actual spend.
Run auto="light" (MIPROv2) or the smallest config first. Docs: "start with moderate values, observe behavior, and
scale up only if you see clear gains" [github.com/stanfordnlp/dspy/issues/1596]. Never start at auto="heavy".
light gives <2% lift → do not escalate. The bottleneck is the program/metric, not the optimizer. Loop
back to [[agentsop-dspy]] Stage 1 (signature ambiguous? wrong decomposition?) [dspy.ai/learn/optimization/overview/].light gives 2–10% lift → escalate to medium; go to heavy only if data ≥300 and you have a held-out
test set distinct from val.compile().compile-ready | not-ready + the failing gate.[[agentsop-dspy]] Case C, §3 three-stage gate.metric(example, pred, trace=None) -> bool|float exists. If absent → route to [[agentsop-metric-design]], gate stays closed.validated flag + receipt reference.[[agentsop-dspy]] Case C step 4; [dspy.ai/learn/evaluation/metrics/]; [arxiv.org/pdf/2506.02592].cleared | below-floor + the legal optimizer set.num_trials × |trainset| × calls; cap spend; set a cheap optimizer LM; enable track_usage=True.auto="light" (or smallest config) first. Never start at heavy.[[agentsop-dspy]] Case A.medium. heavy only if data ≥300 + held-out test set. Confirm gain on held-out test.[[agentsop-dspy]] Case A; [dspy.ai/learn/optimization/overview/].| Examples | Feedback signal | Optimizer | Why / floor |
|---|---|---|---|
<10 | any | LabeledFewShot(k=8) | No search; weakest. Usually means STOP and collect data. [dspy.ai/cheatsheet/] |
~10+ | textual (diffs, schema violations, judge rationales) | dspy.GEPA(metric=m_with_feedback) | Inverts the data-scale rule — reflection on text feedback is sample-efficient. [arxiv.org/abs/2507.19457] |
~30–50 | scalar | BootstrapFewShot / …WithRandomSearch | Self-bootstrapped demos; minimum useful regime. [dspy.ai/learn/optimization/optimizers/] |
200+ | scalar | MIPROv2(metric=m, auto="light") then escalate | Joint instruction + demo Bayesian search; 200 is the documented floor. [dspy.ai/api/optimizers/MIPROv2/] |
| any (post-MIPRO, ship smaller model) | — | chain BootstrapFinetune(student=small, teacher=optimized) | Distills prompts into weights. [dspy.ai/api/optimizers/BootstrapFinetune/] |
GATE 1 — MEASURE
[ ] metric(example, pred, trace=None) -> bool|float exists
[ ] validated vs human: >=20 spot-checks (>=30 open-ended), >=80% agreement
[ ] receipt saved {n_spot_checks, agreement, judge_model, task_model, date}
[ ] NOT a lone holistic LLM-judge (else decompose via [[agentsop-metric-design]] first)
GATE 2 — DATA
[ ] labeled examples counted: N = ____
[ ] N clears the floor of the optimizer chosen below
[ ] (prompt-based optimizers) 20/80 train/val split planned; GEPA uses standard split
PICK + BUDGET
[ ] optimizer chosen by §4.2 (GEPA if textual feedback)
[ ] cost ceiling set; cheap optimizer LM chosen; track_usage=True
[ ] plan: probe auto="light" first, escalate only on >=2% lift, confirm on held-out test| Config | Rough cost | Use when |
|---|---|---|
LabeledFewShot / BootstrapFewShot | cents–~$1 | Floor; ≤50 examples |
MIPROv2(auto="light") | ~$2–3, 6–20 min, ~3.2k calls | First probe, always |
MIPROv2(auto="medium") | single–low-tens of $ | Light showed ≥2% lift |
MIPROv2(auto="heavy") | tens of $, hours | Only if data ≥300 + held-out test + budget |
GEPA (~10+ examples) | low, sample-efficient | Textual feedback available |
Source: [dspy.ai/faqs/]; [github.com/stanfordnlp/dspy/issues/1596]; [arxiv.org/abs/2507.19457].
困境: A 3-stage RAG pipeline already hits 72% on dev after manual prompt tuning. MIPROv2(auto="heavy") would
cost ~$40 and 4 hours. Worth it? (Adapted from [[agentsop-dspy]] Case A.)
约束:
决策步骤:
auto="light" (~$2) often yields large lifts over hand-tuned baselines (paper reports 25%/65% over standard
few-shot) [arxiv.org/abs/2310.03714] — the metric was never actually driving design.light (~$2), never heavy (OP-7). It is a cheap signal for whether more compute helps.heavy. Return to program/metric: signature ambiguous? is 3-stage the right
decomposition? (OP-8) [github.com/stanfordnlp/dspy/issues/1596].medium; heavy only if data ≥300 + a held-out test set distinct from val.结果: The $40 heavy run is almost never the right first move. The $2 probe tells you whether to spend more or
to go fix the program. Often the answer is "fix the program/metric first."
可提取的操作: OP-1, OP-6 CostBudgetGuard, OP-7 CheapProbeFirst, OP-8 EscalateOrReturn. Lesson: never
open at heavy. Probe light with a cheap optimizer LM, and treat <2% as a signal to fix the program, not to add
compute.
困境: A code-fix agent has just 12 labeled examples, but each failing run produces rich textual feedback: the failing test diff, a schema-violation message, a linter error. The data-scale ladder says 12 < 30 → "STOP, collect data, you can't optimize." Is that right?
约束:
决策步骤:
feedback string. Validate it (the test is
ground truth; spot-check the feedback strings are accurate). Pass.dspy.Prediction(score=..., feedback="failing assert: expected X got Y") from the
metric; this is GEPA's superpower [dspy.ai/api/optimizers/GEPA/overview/].结果: A dataset that is far too small for MIPROv2/Bootstrap is sufficient for GEPA. The "you need 200 examples" intuition is specific to scalar-feedback optimizers; textual feedback buys an order-of-magnitude in sample efficiency. GEPA reported beating MIPROv2 by 10–13% on benchmarks while being more sample-efficient [arxiv.org/abs/2507.19457].
可提取的操作: OP-4 ExampleFloorCheck, OP-5 OptimizerByDataScale. Lesson: the example floor is per-optimizer.
If you can express why an output failed as text, GEPA inverts the data-scale assumption — ~10 examples suffice.
Do not reflexively gate out small datasets that carry rich feedback.
| # | Anti-pattern | Why it's wrong | Fix |
|---|---|---|---|
| AP-1 | Compiling without a metric | The optimizer has nothing to maximize; DSPy degenerates to verbose prompting | Refuse the compile; build a metric (OP-2, [[agentsop-metric-design]]) |
| AP-2 | Compiling against an un-validated single LLM-judge | The search faithfully chases the judge's length / self-preference / position bias [arxiv.org/pdf/2506.02592] | Validate ≥20 spot-checks; decompose the judge (OP-3) |
| AP-3 | Compiling on <10 examples (scalar feedback) | Below the floor you memorize, not train; "prompt-based optimizers overfit small sets" [dspy.ai/learn/optimization/overview/] | Collect data, or use LabeledFewShot as a floor — or GEPA if textual feedback (OP-4) |
| AP-4 | Starting at auto="heavy" | Tens of $ / hours before you know if compute even helps | Probe auto="light" first (OP-7) |
| AP-5 | Escalating after a <2% probe lift | The bottleneck is the program/metric; more compute won't fix it | Return to [[agentsop-dspy]] Stage 1 (OP-8) |
| AP-6 | Using the expensive task LM as the optimizer LM | Multiplies trial-scale cost for no documented gain | Cheap optimizer LM (gpt-4o-mini) (OP-6) |
| AP-7 | Applying the scalar data-floor to a feedback-rich task | Wrongly gates out GEPA-legal small datasets | Check for textual feedback before counting examples (OP-5) |
| AP-8 | Reporting the gain on the val set used in optimization | Overfit signal; not a real held-out improvement | Confirm on a held-out test set distinct from val (OP-8) |
| AP-9 | Compiling while the Signature/I-O contract is still churning | You pay compile cost for prompts you'll throw away | Stabilize the contract first (signature-design / [[agentsop-dspy]] §1) |
[[agentsop-metric-design]], then scientific-critical-thinking).signature-design / [[agentsop-dspy]] Stage 1), a different decision from compile.compile() args, save/deploy: that is [[dspy]] and
[[agentsop-dspy]], not this overlay.[[agentsop-metric-design]].
This skill only checks the receipt exists.The readiness gate (metric + data floor) is framework-agnostic; the operators differ.
| Concept | DSPy optimizers | Manual few-shot tuning | OpenAI fine-tuning |
|---|---|---|---|
| What is optimized | prompt instructions + demos (and optionally weights via BootstrapFinetune) | the prompt string, by hand | model weights |
| Metric required? | Yes — metric(ex, pred) -> bool|float drives the search | implicit / eyeballed (the failure mode) | a held-out eval set + loss; eval suite recommended |
| Example floor | ~10 (GEPA+feedback) / ~30 (Bootstrap) / 200+ (MIPROv2) | none, but no guarantees | OpenAI guidance: ~50–100+ examples minimum, more is better |
| Cost shape | `num_trials × | trainset | × calls` ($2–$40+) [dspy.ai/faqs/] |
| Readiness gate (this skill) | both gates apply directly | Gate 1 is exactly what's missing — you're "tuning" with no validated metric | both gates apply; "metric validated" = your eval set is trustworthy |
| When to prefer | metric + ≥10 examples + want a portable, recompilable artifact | one-shot, unstable contract, or audit-mandated verbatim prompts | task is stable, latency/cost matters at high volume, prompt optimization plateaued |
Decision summary: If you have a validated metric and examples clearing a floor, compile (DSPy). If you
have a metric but it's never been validated, you are doing manual tuning dressed up — close Gate 1 first. If
prompt optimization has plateaued and you have hundreds of stable examples and volume justifies it, consider
fine-tuning (or BootstrapFinetune to distill an already-compiled program into a smaller model).
Combination patterns:
[[agentsop-metric-design]]: this skill's Gate 1 is satisfied by a [[agentsop-metric-design]]
calibration receipt. No receipt → gate closed.[[agentsop-dspy]]: this overlay is the sharpened version of [[agentsop-dspy]] §3 Stage 2→3
transition (the "are you allowed to compile yet?" boundary). After the gate opens, hand off to [[agentsop-dspy]] for
the full compile/save/deploy workflow.[[dspy]]: [[dspy]] provides the optimizer APIs; this skill decides whether and
which.references/R1-source-evidence.md — every claim traced to the local dspy-sop skill and the upstream DSPy docs it cites; overlap check vs [[dspy]] and [[agentsop-metric-design]].intermediate/operation_candidates.json — the 8 operations + tables in machine-readable form.Citations: [dspy.ai/learn/optimization/overview/], [dspy.ai/learn/optimization/optimizers/],
[dspy.ai/api/optimizers/MIPROv2/], [dspy.ai/api/optimizers/GEPA/overview/], [dspy.ai/api/optimizers/BootstrapFinetune/],
[dspy.ai/learn/evaluation/metrics/], [dspy.ai/faqs/], [dspy.ai/cheatsheet/], [dspy.ai/tutorials/gepa_ai_program/],
[arxiv.org/abs/2310.03714], [arxiv.org/abs/2507.19457], [arxiv.org/pdf/2506.02592],
[github.com/stanfordnlp/dspy/issues/1596]. Primary source:
/Users/5imp1ex/Desktop/Skill-Workplace/output/dspy-sop-skill/SKILL.md.
© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/agentsop-prompt-compilation of agentsope/SkillAlchemy.
Open the folder on GitHubat commit 6ea799f
Agentsop Prompt Compilation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Prompt Compilation this skillagentsope/SkillAlchemy | 459 | — | ~6.3k | Automated safety check: Pass | MIT | |
| Mg CLImodelguide/modelguide | 108 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Creatoraiskillstore/marketplace | 430 | 1 repos | ~4.8k | Automated safety check: Pass | None | |
| Agent Harness DesignAnastasiyaW/codex-claude-code-config | 154 | — | ~764 | Automated safety check: Pass | MIT | |
| Yao Meta Skillyaojingang/yao-meta-skill | 2.7k | — | ~768 | Automated safety check: Pass | MIT | |
| Agent Creatorjdforsythe/forge | 151 | — | ~4.5k | Automated safety check: Pass | MIT |
modelguide/modelguide
Generate YAML configuration files and run CLI commands to onboard organizations into ModelGuide.
aiskillstore/marketplace
Creates specialized AI agents with optimized system prompts using the official 4-phase SOP methodology from Desktop .claude-flow, combined with evidence-based prompting techniques and Claude Agent…
AnastasiyaW/codex-claude-code-config
Designing agent harnesses and tool systems — risk taxonomy for tools, permission decisions, draft/commit pattern, structured tool results, agent budgets (10 types), context trust labels against…
yaojingang/yao-meta-skill
Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts.
jdforsythe/forge
Creates structured agent definitions using the 7-component format grounded in persona science (the alignment-accuracy tradeoff), vocabulary routing, and the MAST failure taxonomy + Forge watchlist.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when producing a VSS analysis report — Mode A per-clip VLM, Mode B incident-range via video-analytics, Mode C SOP compliance via the SOP tools.
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Categories
The compile-readiness gate for prompt auto-optimization. An agent skill from agentsope/SkillAlchemy. Agentsop Prompt Compilation is an agent skill from agentsope/SkillAlchemy. The compile-readiness gate for prompt auto-optimization.
Agentsop Prompt Compilation fits situations like: tasks that involve Operations and SOPs.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a claude-code`. Or copy the skill folder (skills/agentsop-prompt-compilation in agentsope/SkillAlchemy) into .claude/skills/agentsop-prompt-compilation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a codex`. Or copy the skill folder (skills/agentsop-prompt-compilation in agentsope/SkillAlchemy) into .agents/skills/agentsop-prompt-compilation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-prompt-compilation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-prompt-compilation, .gemini/skills/agentsop-prompt-compilation, .github/skills/agentsop-prompt-compilation and .opencode/skills/agentsop-prompt-compilation in your project.
SKILL.md names no scripts, command-line tools or credentials: Agentsop Prompt Compilation is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Prompt Compilation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Prompt Compilation: Mg CLI (modelguide/modelguide, 108 stars), Agent Creator (aiskillstore/marketplace, 430 stars), Agent Harness Design (AnastasiyaW/codex-claude-code-config, 154 stars) and Yao Meta Skill (yaojingang/yao-meta-skill, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 459 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.