Agent skill

Agentsop Per Model Artifacts

by agentsope in agentsope/SkillAlchemy

Lifecycle SOP for per-model prompt artifacts — the compiled prompts, instructions, few-shot demos, edit-format pins, and embedding-bound indices that change behavior when the underlying LM, dataset…

MITAuto-check passedAI & LLM Engineering

Install Agentsop Per Model Artifacts

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-per-model-artifacts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-per-model-artifacts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-per-model-artifacts .claude/skills/agentsop-per-model-artifacts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-per-model-artifacts
GitHub stars
459
Token cost
~8k tokens
SKILL.md length
3,741 words
Files
5 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Lifecycle SOP for per-model prompt artifacts — the compiled prompts, instructions, few-shot demos, edit-format pins, and embedding-bound indices that change behavior when the underlying LM, dataset…

  • Works in 7 steps: 何时激活 (When to activate) → 核心心智模型 (Core mental model) → SOP 工作流 (SOP workflow) → …
  • Tasks that involve Prompt engineering
  • SKILL.md covers 1. 何时激活 (When to activate), 2. 核心心智模型 (Core mental model), 3. SOP 工作流 (SOP workflow) and 4. 操作模型 (Trigger / Action /…, plus 3 more sections
  • Calls git

What it does

Agentsop Per Model Artifacts is an agent skill from agentsope/SkillAlchemy. Lifecycle SOP for per-model prompt artifacts — the compiled prompts, instructions, few-shot demos, edit-format pins, and embedding-bound indices that change behavior when the underlying LM, dataset, or framework version changes. Activate when adopting compiled prompts (DSPy, GEPA, BootstrapFewShot output), when supporting multiple LMs in production, when a provider deprecates a model snapshot, or when a framework deprecates a config surface (LlamaIndex ServiceContext → Settings, Aider edit-format defaults). Do…

Its SKILL.md is about 8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).

It sits in AI & LLM Engineering, covering Prompt engineering, Operations and SOPs and Embeddings. It works with LlamaIndex and OpenAI. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Tasks that involve Prompt engineering
  • Tasks that involve Operations and SOPs
  • Tasks that involve Embeddings

Example prompts

  • “/agentsop-per-model-artifacts”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 何时激活 (When to activate)
  2. 核心心智模型 (Core mental model)
  3. SOP 工作流 (SOP workflow)
  4. 操作模型 (Trigger / Action / Output / Evidence)
  5. 困境决策案例 (Dilemma cases)
  6. 反模式与边界 (Anti-patterns & boundaries)
  7. 跨框架对照 (Cross-framework comparison)

What it can do on your machine

Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop Per Model Artifacts loads about 8k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 206 tokens; SKILL.md has 3,741 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~206
When it runs · the whole SKILL.md, loaded when a task matches
~8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 3,741 words, ~7,967 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-per-model-artifacts/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
agentsop-per-model-artifacts
description
Lifecycle SOP for **per-model prompt artifacts** — the compiled prompts, instructions, few-shot demos, edit-format pins, and embedding-bound indices that change behavior when the underlying LM, dataset, or framework version changes. Activate when adopting compiled prompts (DSPy, GEPA, BootstrapFewShot output), when supporting multiple LMs in production, when a provider deprecates a model snapshot, or when a framework deprecates a config surface (LlamaIndex `ServiceContext` → `Settings`, Aider edit-format defaults). Do NOT activate for one-off raw prompt edits or for truly model-agnostic system prompts that have been swap-tested. Search keywords: prompt portability, model swap, recompile prompt, model deprecation, prompt per model, prompt breaks on new model, version compiled prompts.
version
0.1.0

Per-Model Prompt Artifacts — SOP

"Prompts are effectively the weights of an LLM application." — DSPy core philosophy [arxiv.org/abs/2310.03714]

"Treat the compiled program as a (program × LM) pair. Changing the LM invalidates the artifact — recompile." — dspy-sop SKILL, Dilemma Case B


1. 何时激活 (When to activate)

Activate this skill when any of the following appears in the user's intent, codebase, or workflow:

TriggerSignal
Adopting compiled promptscompiled.save("v1.json"), dspy.load_program(), BootstrapFewShot, MIPROv2, GEPA, LangChain Hub hub.push/pull, prompt files checked into prompts/ or artifacts/
Multi-LM productionThe same program runs against ≥2 of: gpt-4o-*, gpt-4o-mini-*, gpt-4.1-*, claude-3-5-sonnet-*, claude-3-7-sonnet-*, claude-3-opus-*, Llama-3-*, Llama-3.1-*, DeepSeek-V3, gemini-2.5-pro
Provider deprecationOpenAI/Anthropic deprecation notice mentioning a pinned snapshot; alias rollover (gpt-4o → new dated snapshot); silent model behavior drift reports
Framework deprecationLlamaIndex ServiceContext → Settings; LangChain LLMChain → LCEL; DSPy major version bump; aider edit-format default change
Symptomsprompts/system_prompt.txt (no model in path), generic alias pins (model="gpt-4o"), missing parent_artifact lineage, hand-edited compiled JSON, no held-out test set re-runs
Cross-skill bridgesDSPy compile produced output → ship via this SOP. LlamaIndex index baked with embed model → tag artifact per this SOP. Aider edit-format pin → per-model config artifact per Recipe 6 (R2)

Do NOT activate when:

  • Raw, one-off prompt edits with no compile step and no production deploy.
  • A genuinely model-agnostic system prompt that has been swap-tested across ≥3 LMs with <2-point dev metric drift.
  • Prompts that must remain verbatim human-authored for compliance — versioning still matters, but the optimizer/recompile loop does not apply.
  • The whole pipeline lives behind a vendor's managed prompt (e.g. OpenAI's Prompt Library) where the vendor owns the artifact.

2. 核心心智模型 (Core mental model)

The artifact is a triple, not a string
                ┌────────── ARTIFACT ──────────┐
                │                              │
   program  ×   LM (snapshot)  ×   dataset (hash)   →   compiled.json + metadata
   (code)       (provider+date)    (canonical hash)
                │                              │
                └─ change any axis → new artifact ─┘

Three load-bearing claims:

  1. A compiled prompt is not portable. The DSPy doctrine (sibling skill, Case B): swap the LM → recompile. Reusing GPT-4o-compiled program.json on Llama-3-8B typically loses 15–30 points (R1 §5). The asymmetry is real and documented.

  2. The snapshot, not the alias, is the identity. gpt-4o is a moving alias; gpt-4o-2024-08-06 is an immutable identifier. Pinning to the alias means your behavior changes silently when OpenAI rolls the alias forward. Anthropic does not even roll aliases — Sonnet 3.5 (claude-3-5-sonnet-20240620) and Sonnet 3.5 v2 (claude-3-5-sonnet-20241022) are different models with different behavior; the org that pinned the wrong one in 2024-10 ate a regression in 2025-10 when the older one was retired (R1 §4).

  3. The dataset is part of the identity. Compiled prompts overfit their training distribution. Two artifacts compiled from the "same dataset" that turn out to differ by 50 examples produce silently different programs. A sha256 over canonicalized dataset bytes, embedded in the artifact path, makes this impossible to confuse (OP-9).

The PyTorch checkpoint analogy
PyTorch checkpointPer-model prompt artifact
.pt weights fileprogram.json (compiled prompt)
Architecture (forward pass)DSPy Signature + Module graph (source-of-truth Python)
Dataset versiondataset_sha256_8 in path
Hyperparametersoptimizer config (MIPROv2 auto="light", seed, demos)
Framework versiondspy_version, python_version in metadata
Eval score on valdev_score in REGISTRY
Final test scoretest_score in REGISTRY
Model registry (MLflow)artifacts/REGISTRY.jsonl or MLflow (Recipe 5)
Promotion to prodgit tag prompt/<program>/<snapshot>/v<n>

The same engineering discipline applies. Anyone who would not check a .pt into prod without a checkpoint registry should not check a compiled prompt into prod without an artifact registry.

Why "alone is not enough"

A bare compiled.save("v1.json") produces a JSON file that, viewed in isolation, looks model-agnostic. The instructions and demos are text — they could be for any model. But they were selected by an optimizer running calls against a specific LM. The model's distributional response to the bootstrapping queries shaped which demos got kept. The artifact is implicitly LM-conditioned without saying so. The SOP exists to make this conditioning explicit and auditable.


3. SOP 工作流 (SOP workflow)

A five-stage lifecycle. Each stage has an exit criterion.

Stage 1 — Initial compile
  1. Source-of-truth Python module (src/<program>/program.py) defines the DSPy module graph.
  2. Pin the dated snapshot of the LM (not the alias): dspy.LM("openai/gpt-4o-2024-08-06").
  3. Pin the dataset: canonicalize → sha256 → use first 8 hex chars (a8f1c2e9) as the dataset tag.
  4. Run the optimizer (start with MIPROv2(auto="light") per dspy-sop). Record cost_usd, wall_seconds.
  5. Evaluate on a held-out test set (not the optimizer's val split).

Exit: Held-out test score recorded. No artifact saved yet.

Stage 2 — Version (write artifact + registry)
  1. Compute the path: artifacts/<program>/<provider>/<model-snapshot>/<dataset>-<sha8>/v<n>.json where n is the next integer after existing files in that directory.
  2. compiled.save(path) for the JSON; optionally also compiled.save(path_dir, save_program=True) for whole-program reproducibility.
  3. Append one line to artifacts/REGISTRY.jsonl with the schema in R2 Recipe 1.
  4. Commit both the artifact AND the REGISTRY line in one PR (gated by Stage 3).

Exit: Artifact + REGISTRY entry on disk, both staged for PR.

Stage 3 — Register (regression gate)
  1. CI loads the new artifact, identifies its parent_artifact (previous artifact for same program × model × dataset), re-runs held-out test set, compares.
  2. Fail PR if new_test_score - parent_test_score < -GATE (default GATE = 0.02, 2 points on 0-1 scale).
  3. Block merge unless gate passes or a human attaches the artifact-override label (with justification in PR body).

Exit: PR merged to main, REGISTRY canonical.

Stage 4 — Swap-test (on model swap or new snapshot)
  1. New LM snapshot announced (e.g. gpt-4o-2024-11-20 replacing 2024-08-06). Or considering family swap (GPT-4o → Sonnet 3.5).
  2. Load existing artifact unchanged. Configure DSPy to the new LM. Run held-out test set.
  3. Compute Δ-score vs original.
  • |Δ| ≤ 2 points: tag artifact transferable[old→new] in REGISTRY. May ship without recompile.
  • |Δ| > 2 points: tag recompile_required[old→new]. Proceed to Stage 5.
  1. Emit swap-report-<old>-to-<new>.md artifact regardless of outcome (audit trail).

Exit: Swap report committed, recompile decision recorded.

Stage 5 — Recompile or accept
  1. If recompile_required: re-run the same optimizer config against the new LM. Produces a new artifact at the new model's path. Increment v<n>.
  2. If transferable: keep using the old artifact; mark in REGISTRY which model snapshots it's certified-transferable to.
  3. Keep BOTH old and new artifacts checked in. Production deploy script reads the git tag, not HEAD.
  4. Loop back to Stage 2 for the new artifact.

Exit: Production deploys an artifact whose REGISTRY entry shows passing test score against the model it will actually run on.

Quarterly maintenance loop
  • Run deprecation scan (R2 Recipe 4). Open issues for artifacts within 90 days of model retirement.
  • Run canary re-eval: re-run held-out test set against the pinned snapshot to detect silent provider-side drift. Threshold same as regression gate.

4. 操作模型 (Trigger / Action / Output / Evidence)

4.1 Storage operations
TriggerActionOutputEvidence
Compile finishes successfullyPath <program>/<provider>/<snapshot>/<dataset-sha8>/v<n>.json; append REGISTRY line with parent_artifact pointerGrep-able path + lineageR2 Recipe 1; OP-1, OP-9
About to pin a modelUse dated snapshot never an alias: gpt-4o-2024-08-06 not gpt-4o; claude-3-5-sonnet-20241022 not claude-3-5-sonnetMetadata header inside artifact + REGISTRYR1 §4 (alias rollover); OP-2
Multiple compiles share a program codeAll keyed by program field in REGISTRY (string), commit field links to source-of-truth Python at compile timeReproducible: git checkout <commit> restores the program code that produced itOP-3
Promoting to productiongit tag prompt/<program>/<snapshot>/v<n>. Deploy reads the tag, not main HEADAtomic rollback via git checkout <tag>MLflow Stage analogy (R1 §7); OP-10
Need whole-program portabilitycompiled.save(dir, save_program=True) for pickled program + state; ship bothSelf-contained; survives source-tree reorgs (until Python/DSPy version mismatch)DSPy save_program docs [dspy.ai/tutorials/saving/]
4.2 Swap-test operations
TriggerActionOutputEvidence
Provider releases new snapshotOP-4 swap-test old artifact against new model on held-out test setΔ-score; recompile_required booleanR2 Recipe 2
Family swap candidate (GPT-4o → Sonnet)Swap-test BOTH directions (artifact-from-A run on B; artifact-from-B run on A); record asymmetryTwo Δ-scores; choose recompile directionR1 §5 (down-transfer 15-30pt drop)
Aider edit-format default changeRe-run aider edit benchmark on fixed task subset for affected modeledit-format-bench delta; pin if changedaider SKILL line 302 (udiff 20→61%)
LlamaIndex Settings.embed_model change candidateRefuse to swap without rebuilding index. Re-embed against full corpus; tag new index artifactNew indices/<name>/openai__<embed-model>__<dim>/<corpus-sha8>/LlamaIndex SKILL anti-pattern A2; R2 Recipe 7
4.3 Gate operations
TriggerActionOutputEvidence
PR touches artifacts/**/v*.jsonCI regression-gate: re-run held-out test against new artifact and parent artifactCI status artifact-regression: pass|fail|overriddenR2 Recipe 3; OP-5
PR tries to add file under artifacts/ without REGISTRY appendPre-commit hook failsLocal hook error; can't pushOP-8
PR pins a snapshot ≤90 days from deprecationPre-commit hook warns; CI fails unless override labelSurfaces deprecation debt at compile timeOP-7; R2 Recipe 4
Hand edit to compiled JSONPre-commit hook fails: artifact files are read-only except via compile.pyForces recompile pathOP-8; dspy-sop anti-pattern #3
4.4 Lineage operations
TriggerActionOutputEvidence
Looking up "which artifact is in prod for this program?"git ls-remote --tags origin 'prompt/<program>/*'Tag list with snapshot + versionOP-10
Reproducing a past compileREGISTRY entry has commit, dspy_version, optimizer_config, dataset_sha256_8, seed; replay against original LM snapshotBit-identical or near-identical artifactR2 Recipe 1; OP-3
Auditing "what changed when score regressed"Diff REGISTRY entries; if dataset_sha256_8 changed → dataset drift; if optimizer_config changed → config drift; if dspy_version changed → framework drift; if commit changed without code changes → environment driftRoot-cause class identifiedOP-3

5. 困境决策案例 (Dilemma cases)

Case A — "GPT-4o deprecation: artifact loses 8 points on Sonnet 3.7"

困境 (Dilemma): Production runs a DSPy-compiled program against gpt-4o-2024-05-13 at 82% test score. OpenAI announces deprecation in 60 days. The team wants to move to Sonnet 3.7 (claude-3-7-sonnet-20250219) for unrelated cost/latency reasons. Swap-test (OP-4) shows the unchanged artifact scores 74% on Sonnet 3.7 — Δ = −8 points. Recompile, accept, or hybrid?

约束 (Constraints):

  • Recompile cost on Sonnet 3.7 (200-example trainset, MIPROv2 light): ~$4 and 20 minutes.
  • Customer SLA is "≥80% on the public eval set".
  • Sonnet 3.7 has extended thinking budget — different output distribution from GPT-4o.
  • Old artifact's instructions reference "be terse" — Sonnet 3.7 with thinking ignores terseness hints.
  • Team has 60 days, not 60 minutes.

决策步骤 (Decision steps):

  1. Do not accept the −8 swap as-is. The artifact is below SLA on the target model. Per dspy-sop Case B doctrine, "Changing the LM invalidates the artifact — recompile" is the default.
  2. Recompile against Sonnet 3.7 with the same dataset. Run MIPROv2(auto="light") first (R2 cost guardrail). If light gives ≥80%, stop. If <80%, escalate to medium only after Stage 1 review of the program/metric.
  3. Tune the optimizer for Sonnet 3.7 specifics. Set max_bootstrapped_demos lower (Sonnet 3.7's thinking emits richer reasoning per-demo, so fewer richer demos > more thin demos — R1 §1).
  4. Keep BOTH artifacts in artifacts/<program>/. The GPT-4o-2024-05-13 artifact stays available for canary-rollback during the cutover window.
  5. Git tag the new artifact as prompt/<program>/claude-3-7-sonnet-20250219/v1 only after the regression gate passes against the Sonnet 3.7 held-out test set.
  6. Emit swap-report-gpt-4o-2024-05-13-to-claude-3-7-sonnet-20250219.md showing: transfer Δ −8, recompiled Δ +3 vs original. Audit trail.

结果 (Outcome): Recompiled Sonnet 3.7 artifact lands at 83% on the same held-out set. Cost: $4 + reviewer time. The +1 point over the GPT-4o original is bonus; the value was avoiding the −8 cliff.

可提取的操作 (Extractable operation): A negative swap-test Δ exceeding the regression gate is a recompile signal, not a "ship anyway" signal. The artifact path makes both old and new available simultaneously — there is no migration risk to keeping both.


Case B — "Silent 4o-patch: same snapshot string, different behavior"

困境: Quarterly canary re-runs the held-out test set against the production artifact (gpt-4o-2024-08-06, no model change). Score dropped from 78% to 73% over three months. No code changed. No artifact changed. No registry entry changed. What happened?

约束:

  • The pinned snapshot string did not change.
  • OpenAI's policy is "snapshots are stable" but in practice server-side mitigations (safety, hallucination patches, throughput) ship without bumping the snapshot ID.
  • Customer noticed before the team did. SLA at risk.
  • Reverting the model is not possible — the canary IS the latest behavior on the same snapshot.

决策步骤:

  1. Verify the canary against a stored sample-output hash. Per OP-2, REGISTRY entries log a sample-output hash from compile time. Re-run against the same 20 sample inputs; compare current outputs to stored hashes. Divergence → confirms behavior shifted on the same snapshot.
  2. Run the regression gate locally with current canary outputs as the "new artifact" and compile-time outputs as "parent". This isolates whether the metric also dropped or just the outputs changed (semantically-equivalent rewrites would shift hashes but not score).
  3. Recompile against the same snapshot. Even though the snapshot string is unchanged, recompiling lets the optimizer re-bootstrap demos against the new server-side behavior. Cost is the same $2–4 as initial light compile.
  4. Tag the new artifact as v<n+1> under the same model directory: artifacts/<program>/openai/gpt-4o-2024-08-06/<dataset-sha>/v<n+1>.json. The compiled_at field in REGISTRY makes the temporal lineage clear even though the snapshot string is unchanged.
  5. Add the canary cadence to weekly for 30 days post-incident; revert to quarterly when stable.
  6. File issue with OpenAI (if support contract): provide REGISTRY entries showing same snapshot string, same code commit, same dataset hash, drift > regression gate.

结果: Recompile recovers score to 79% (+1 over baseline, because the optimizer found demos that work with the new server behavior). Lesson: snapshot pins protect against MOST drift but not all. Canary catches what pins miss.

可提取的操作: The snapshot string is necessary but not sufficient. A periodic canary against held-out test data is the only true contract with the model provider. Store sample-output hashes per artifact so silent drift is detectable, not just inferable.


Show full SKILL.md (1,585 more words)Show less
Case C — "LlamaIndex ServiceContext deprecation: which artifacts to migrate"

困境: Team upgrades from LlamaIndex 0.9.x to 0.11.x. ServiceContext(llm=..., embed_model=...) is deprecated in favor of Settings.llm = ... / Settings.embed_model = ... (LlamaIndex SKILL §4.4). There are 14 indices in production, each baked with various embed models including text-embedding-ada-002 (deprecated) and text-embedding-3-small (current). What needs to be touched?

约束:

  • Code-level migration: ServiceContext → Settings is mechanical, ~1 hour.
  • Index-level migration: re-embedding the largest corpus (4M docs, text-embedding-3-small) is ~$200 and 6 hours.
  • Indices on text-embedding-ada-002 MUST be re-embedded — that model is going away (R1 §3).
  • Mixing embedding spaces silently is the worst-case failure (LlamaIndex SKILL anti-pattern A2).

决策步骤:

  1. List every index and its manifest (per R2 Recipe 7). If any index lacks a manifest.json with embed_model and dim, treat as unknown — must re-embed defensively.
  2. Bucket indices by embed model:
    • On deprecated embed (text-embedding-ada-002): must re-embed. New artifact path indices/<name>/openai__text-embedding-3-small__1536/<new-corpus-sha8>/.
    • On current embed: code-only migration (ServiceContext → Settings). Index artifact unchanged. Update manifest with new llama_index_version.
  3. Refuse to start the new app if Settings.embed_model at runtime doesn't match a loaded index's manifest.embed_model (R2 Recipe 7 loader guard). This converts a silent data-corruption bug into a startup error.
  4. For each re-embedded index, write a NEW manifest with a NEW corpus_sha8 even if the source corpus is unchanged — different model = different artifact, period.
  5. Keep the old index files for 30 days under indices/<name>/openai__text-embedding-ada-002__1536/_archived/. Roll back if production answers degrade. (Anti-pattern: deleting the old index "to save space" before validating the new one.)
  6. Tag the new app version only after at least one rep query per re-embedded index passes a regression test (top-3 results overlap with golden set ≥80%).

结果: 9 of 14 indices needed code-only migration; 5 needed re-embed. Total cost ~$300 and 1 engineer-day. The discipline of per-index manifests caught two indices that had been silently broken (ada-002 query vector against 3-small index from a botched earlier migration).

可提取的操作: Framework deprecations cascade into per-artifact decisions. The "easy" code migration is rarely the full migration — interrogate every artifact's framework binding (embed model, edit format, tokenizer version) before declaring a deprecation handled.


6. 反模式与边界 (Anti-patterns & boundaries)

Anti-patterns
  1. Assuming portability across models. A prompt.txt or compiled.json with no model in its path or metadata is implicitly "for any model." Empirically, prompts transfer down (big → small) badly and sideways unpredictably (R1 §5). The artifact must encode its LM.

  2. No model in the artifact name/path. artifacts/v3.json is unauditable. artifacts/rag_synth/openai/gpt-4o-2024-08-06/devset-v3-a8f1c2e9/v3.json is. Long paths are good; they replace metadata files no one reads.

  3. Pinning to an alias instead of a snapshot. gpt-4o, claude-3-5-sonnet, llama-3-8b-instruct — all moving targets. Pin gpt-4o-2024-08-06, claude-3-5-sonnet-20241022, and an HF revision SHA respectively.

  4. No regression gate on artifact PRs. Every artifact replacement is implicitly a behavior change. CI must re-run held-out test and compare to parent artifact. Without this, regressions land silently and surface in production.

  5. Hand-editing compiled JSON. "Just one little instruction tweak" invalidates the metric-optimality the artifact carries. Either re-compile or don't touch. Pre-commit hook should enforce.

  6. Skipping the parent_artifact pointer. Without lineage, you cannot answer "what was the last known-good artifact" during an incident. Every REGISTRY entry must reference its predecessor (or null for the very first).

  7. Single artifact, multiple LMs. Production hits two LMs (e.g. failover routing GPT-4o ↔ Sonnet 3.5). Each LM needs its OWN artifact, with its own swap-tested score. A "shared" artifact biased toward whichever LM compiled it.

  8. No dataset hash. Two "v3 devset" files that differ by 50 examples produce different artifacts that look identical in the registry. dataset_sha256_8 in the path makes drift visible.

  9. Forgetting framework version. A DSPy 2.4 artifact loaded by DSPy 2.6 may silently behave differently due to module-level changes. dspy_version in REGISTRY enables targeted rollback.

  10. Deleting old artifacts on new compile. Disk is cheap, history is precious. Keep old artifacts checked in. Production deploys read git tags, not HEAD — the old artifact stays reachable for rollback.

  11. Treating LangChain Hub as the lifecycle. Hub is a viewer + diff tool. It doesn't bind to model, dataset, or run a regression gate. Pair with REGISTRY.jsonl or MLflow (R1 §8).

  12. Promoting an artifact without swap-test on the actual production LM. "It scored well in compile" — but compile was against the same LM that production uses, right? Not always: the optimizer LM can differ from the task LM (dspy-sop §4.4). Swap-test against the actual production LM before tagging.

Boundaries (when this SOP is overkill)
  • One artifact, one LM, never changes. A research demo or proof-of-concept can use compiled.save("v1.json") flat. Adopt this SOP when adding the second artifact or moving to a managed environment.
  • Pure inference-time prompts with no compile loop. A system prompt manually authored and unchanging is just a config file — version it normally, but the regression gate and swap-test machinery are not load-bearing until the prompt is metric-optimized.
  • Provider with built-in artifact registry. OpenAI's Prompt Library, Anthropic's Workbench saved prompts, vendor-managed RAG — let the vendor own the lifecycle. This SOP is for org-controlled artifacts.
  • Framework changes more often than the LM. If you re-write the program every week, the compile artifact has no shelf life and the registry overhead exceeds the audit value. Stabilize the program first.

7. 跨框架对照 (Cross-framework comparison)

How four ecosystems handle the per-model artifact problem.

DSPy — save_program is the artifact
  • What it stores: instructions per predictor, bootstrapped demos, signature shapes (for whole-program save).
  • What it does NOT store: the LM identity, the dataset, the metric, the run cost. These must be added by the SOP (REGISTRY.jsonl).
  • Two save modes:
    • compiled.save("v1.json") — state JSON, needs source Python at load. Use for code-review-friendly diffs.
    • compiled.save("v1/", save_program=True) — pickled program + state. Use for portability across source-tree refactors.
  • Reload: program_cls().load("v1.json") requires the Python class import; dspy.load("v1/") for whole-program.
  • Versioning practice: DSPy docs leave it to the user. The dspy-sop SKILL Case B says "Treat the compiled program as a (program × LM) pair" but doesn't prescribe directory layout — this SOP fills that gap.
LangChain Hub — versioned text templates
  • What it stores: prompt template text, partial variables, model name as a hint (not enforced).
  • What it does NOT store: evaluation scores, dataset binding, recompile lineage.
  • API:
    python
    from langchain import hub
    hub.push("user/rag-qa", prompt, parent_commit_hash=last)
    hub.pull("user/rag-qa:v3")
  • Versioning practice: Hub gives you a UI diff between versions but not a regression gate or model-pinning enforcement. Use Hub as the viewer; keep your local REGISTRY.jsonl as the gate.
  • Limitation: Hub is org-shared; per-model variants pollute the namespace unless you name them rag-qa-gpt4o / rag-qa-sonnet35.
Manual jsonl + git — lightweight, fully local
  • What it stores: whatever you put in REGISTRY.jsonl. Recipe 1 schema (R2) covers model, dataset, metric, cost, lineage.
  • Strengths: zero infra, fully grep-able, git-blame-able. Works for solo developers and small teams.
  • Limitations: no UI, no multi-tenant access control, no eval-vs-baseline comparison out of the box (must script). Scales to ~50 artifacts before the jsonl gets unwieldy.
  • Decision rule: start here. Migrate to MLflow when you exceed 50 artifacts OR have >5 contributing developers OR need a stage-promotion UI.
MLflow Model Registry — org-scale
  • What it stores: model artifact (pickled DSPy program via mlflow.dspy.log_model), full param/metric history per run, signature, version with explicit Stage labels (None / Staging / Production / Archived).
  • What it does NOT store natively: dataset hash binding (must add as a param), provider-snapshot-deprecation calendar (add via tags).
  • API:
    python
    with mlflow.start_run():
        mlflow.log_params({"model": MODEL, "dataset_sha8": SHA8})
        mlflow.log_metrics({"dev_score": d, "test_score": t})
        mlflow.dspy.log_model(compiled, "program",
                              registered_model_name=f"rag_synth-{MODEL}")
    client.transition_model_version_stage(
        name=f"rag_synth-{MODEL}", version=v, stage="Production")
  • Strengths: UI, RBAC, multi-team. Promotion via stage labels replaces git tags.
  • Limitations: infra overhead (Postgres + artifact store). Overkill for solo.
Aider model defaults — config artifacts, not data artifacts
  • What it stores: .aider.conf.yml per-model edit format, params (R2 Recipe 6).
  • Versioning practice: This SOP applied to config: pre-commit hook + bench delta gate on changes.
  • Lesson: "artifact" generalizes beyond compiled JSON. Aider's edit-format choice IS a per-model artifact even though it's a string in YAML. Treat it like one.
LlamaIndex index files — embedding-bound artifacts
  • What it stores: vector store, doc store, index store — all bound to one specific embedding model + dimension.
  • Versioning practice: R2 Recipe 7. Per-index manifest.json with embed_model, dim, corpus_sha8. Loader refuses to query if Settings.embed_model at runtime mismatches manifest.
  • Lesson: embedding model is part of index identity, identical in structure to LM being part of compiled-prompt identity. Same SOP applies.
Summary table
LayerNative artifact mechanismGap this SOP fills
DSPysave_program (JSON or pickled dir)Adds LM/dataset/metric binding, registry, gate
LangChain HubVersioned text templates with diff UIAdds model-pinning enforcement, eval gate
Manual jsonl + gitNone — this SOP IS the recipeProvides Recipe 1 schema and gate scripts
MLflow Model RegistryStages, signatures, run historyAdds dataset_sha8 convention, deprecation scan
AiderPer-model edit-format defaults in codeAdds change-gate (re-run bench on edit-format change)
LlamaIndexSettings + index filesAdds manifest-based loader guard, embed-model in path

Quick-reference appendix

Filesystem layout (canonical)
artifacts/
├── REGISTRY.jsonl
└── <program>/
    └── <provider>/
        └── <model-snapshot>/        # e.g. gpt-4o-2024-08-06
            └── <dataset>-<sha8>/    # e.g. devset-v3-a8f1c2e9
                ├── v1.json
                ├── v2.json
                └── v3.json
REGISTRY.jsonl one-liner
jsonl
{"path":"...","sha256":"...","program":"...","provider":"...","model":"...","dataset":"...","dataset_sha256_8":"...","optimizer":"...","optimizer_config":{...},"metric":"...","dev_score":0.0,"test_score":0.0,"cost_usd":0.0,"wall_seconds":0,"dspy_version":"...","compiled_at":"...","commit":"...","parent_artifact":null}
Pre-commit hooks (required)
  • no-edit-artifact: rejects any change under artifacts/** that doesn't also append to REGISTRY.jsonl.
  • no-alias-pin: scans Python source for model="gpt-4o", model="claude-3-5-sonnet" etc. with no date suffix; fails.
  • no-deprecated-pin: scans for snapshots within 90 days of deprecation per scripts/deprecation_table.py; fails without override label.
CI gates (required)
  • artifact-regression: re-run held-out test on new artifact + parent; fail if Δ < −2 points.
  • swap-test-on-snapshot-change: when a new provider snapshot is announced (via webhook or scheduled scan), open a PR running OP-4 against affected artifacts.
Decision tree
New compile? ──► Stage 1-3 (compile, version, gate)
   │
New snapshot from provider? ──► Stage 4 swap-test
   │   ├── |Δ| ≤ 2 pts ──► tag transferable; ship
   │   └── |Δ| > 2 pts ──► Stage 5 recompile
   │
Framework deprecation? ──► R2 Recipe 7 pattern: list artifacts, bucket by binding, migrate per-artifact
   │
Canary score drift on stable snapshot? ──► Case B path: recompile against same snapshot
   │
Family swap (GPT-4o ↔ Sonnet)? ──► Stage 5 recompile, always; swap-test gives the magnitude, not the decision
Key references
  • DSPy compiled-prompt model dependency: /Users/5imp1ex/Desktop/Skill-Workplace/output/dspy-sop-skill/SKILL.md Dilemma Case B.
  • Aider per-model edit format defaults: /Users/5imp1ex/Desktop/Skill-Workplace/output/aider-sop-skill/SKILL.md §"edit format" table.
  • LlamaIndex ServiceContext → Settings deprecation: /Users/5imp1ex/Desktop/Skill-Workplace/output/llamaindex-sop-skill/SKILL.md §4.4 / anti-pattern A4.
  • DSPy save_program API: dspy.ai/tutorials/saving/.
  • MLflow Model Registry: mlflow.org/docs/latest/model-registry.html.
  • OpenAI deprecation policy: platform.openai.com/docs/deprecations.
  • Anthropic model deprecations: docs.anthropic.com/en/docs/about-claude/model-deprecations.
  • This SOP's reference files: references/R1-source-evidence.md, references/R2-versioning-recipes.md.

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/agentsop-per-model-artifacts of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-source-evidence.md
  • references/R2-versioning-recipes.md

Open the folder on GitHubat commit 6ea799f

Compare with similar skills

Agentsop Per Model Artifacts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop Per Model Artifacts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop Per Model Artifacts this skillagentsope/SkillAlchemy459—~8kAutomated safety check: PassMIT
Spring AI Integrationrrezartprebreza/spring-boot-skills2981 repos~2.1kAutomated safety check: PassMIT
Spring AI Integrationrrezartprebreza/spring-boot-skills2981 repos~2.5kAutomated safety check: PassMIT
AI SDKvercel-labs/ai-facts16821 repos~1.2kAutomated safety check: PassNone
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0

Similar skills

  • Spring AI Integration

    rrezartprebreza/spring-boot-skills

    A skill your agent uses when integrating LLMs, chat clients, embeddings, RAG pipelines, or AI agents into Spring Boot.

    298 GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Spring AI Integration

    rrezartprebreza/spring-boot-skills

    A skill your agent uses when integrating LLMs, chat clients, embeddings, RAG pipelines, or AI agents into Spring Boot.

    298 GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • AI SDK

    vercel-labs/ai-facts

    Official

    Answer questions about the AI SDK and help build AI-powered features.

    168 GitHub starsUsed in 21 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed

More from agentsope/SkillAlchemy

All 45 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    459 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    459 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    459 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    459 GitHub stars~4.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    459 GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check: notes
  • Agentsop Multiscale Chunking

    agentsope/SkillAlchemy

    Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

    459 GitHub stars~4.9k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Agentsop Per Model Artifacts

What does Agentsop Per Model Artifacts do?

Lifecycle SOP for per-model prompt artifacts — the compiled prompts, instructions, few-shot demos, edit-format pins, and embedding-bound indices that change behavior when the underlying LM, dataset…. Agentsop Per Model Artifacts is an agent skill from agentsope/SkillAlchemy. Lifecycle SOP for per-model prompt artifacts — the compiled prompts, instructions, few-shot demos, edit-format pins, and embedding-bound indices that change behavior when the underlying LM, dataset, or framework version changes.

When should I use Agentsop Per Model Artifacts?

Agentsop Per Model Artifacts fits situations like: tasks that involve Prompt engineering; tasks that involve Operations and SOPs; tasks that involve Embeddings.

How do I install Agentsop Per Model Artifacts in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-per-model-artifacts -a claude-code`. Or copy the skill folder (skills/agentsop-per-model-artifacts in agentsope/SkillAlchemy) into .claude/skills/agentsop-per-model-artifacts in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop Per Model Artifacts in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-per-model-artifacts -a codex`. Or copy the skill folder (skills/agentsop-per-model-artifacts in agentsope/SkillAlchemy) into .agents/skills/agentsop-per-model-artifacts in your project. Codex loads it when a task matches its description.

Can I use Agentsop Per Model Artifacts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-per-model-artifacts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-per-model-artifacts, .gemini/skills/agentsop-per-model-artifacts, .github/skills/agentsop-per-model-artifacts and .opencode/skills/agentsop-per-model-artifacts in your project.

What does Agentsop Per Model Artifacts need to run?

Going by SKILL.md and its folder, Agentsop Per Model Artifacts needs the command-line tools its instructions call (git). Our summary lists: Python 3.

Does Agentsop Per Model Artifacts access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Agentsop Per Model Artifacts safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop Per Model Artifacts use?

Agentsop Per Model Artifacts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop Per Model Artifacts use?

About 8k tokens (SKILL.md is roughly 32k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop Per Model Artifacts?

Skills that share tags, products or a category with Agentsop Per Model Artifacts: Spring AI Integration (rrezartprebreza/spring-boot-skills, 298 stars), Spring AI Integration (rrezartprebreza/spring-boot-skills, 298 stars), AI SDK (vercel-labs/ai-facts, 168 stars) and Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop Per Model Artifacts?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 459 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.