Build ML Pipeline
probabl-ai/skills
Declare the pipeline from data source to predictor as a skrub DataOps graph.
Enhancement overlay — version the WHOLE deployable LLM-app artifact as one bundle: prompts + compiled programs + model snapshot pins + retrieval config + eval-set version, versioned together so a…
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-llm-artifact-versioning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-llm-artifact-versioning .claude/skills/agentsop-llm-artifact-versioning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-llm-artifact-versioning" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-llm-artifact-versioning into .claude/skills/agentsop-llm-artifact-versioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-llm-artifact-versioning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-llm-artifact-versioningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-llm-artifact-versioning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-llm-artifact-versioning .agents/skills/agentsop-llm-artifact-versioning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-llm-artifact-versioning" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-llm-artifact-versioning into .agents/skills/agentsop-llm-artifact-versioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-llm-artifact-versioning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-llm-artifact-versioning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-llm-artifact-versioning .cursor/skills/agentsop-llm-artifact-versioning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-llm-artifact-versioning" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-llm-artifact-versioning into .cursor/skills/agentsop-llm-artifact-versioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-llm-artifact-versioning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-llm-artifact-versioning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-llm-artifact-versioning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-llm-artifact-versioning .gemini/skills/agentsop-llm-artifact-versioning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-llm-artifact-versioning" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-llm-artifact-versioning into .gemini/skills/agentsop-llm-artifact-versioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-llm-artifact-versioning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-llm-artifact-versioningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-llm-artifact-versioning .github/skills/agentsop-llm-artifact-versioning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-llm-artifact-versioning" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-llm-artifact-versioning into .github/skills/agentsop-llm-artifact-versioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-llm-artifact-versioning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-llm-artifact-versioning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-llm-artifact-versioning .opencode/skills/agentsop-llm-artifact-versioning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-llm-artifact-versioning" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-llm-artifact-versioning into .opencode/skills/agentsop-llm-artifact-versioning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-llm-artifact-versioning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-llm-artifact-versioningEnhancement overlay — version the WHOLE deployable LLM-app artifact as one bundle: prompts + compiled programs + model snapshot pins + retrieval config + eval-set version, versioned together so a…
Agentsop LLM Artifact Versioning is an agent skill from agentsope/SkillAlchemy. Enhancement overlay — version the WHOLE deployable LLM-app artifact as one bundle: prompts + compiled programs + model snapshot pins + retrieval config + eval-set version, versioned together so a deploy is reproducible and rollback is atomic. Activate when preparing to deploy an LLM app, when asking "what exactly is running in prod right now?", when a deploy must be reproducible months later, or when an incident needs a clean rollback. The core reframe: an LLM app artifact is NOT an ML model — it is a manifest…
Its SKILL.md is about 7.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in DevOps & Cloud, covering Machine learning and Operations and SOPs. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop LLM Artifact Versioning loads about 7.1k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 261 tokens; SKILL.md has 3,344 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 3,344 words, ~7,080 tokens.
.claude/skills/agentsop-llm-artifact-versioning/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub."Prompts are effectively the weights of an LLM application." — DSPy core philosophy [arxiv.org/abs/2310.03714] (R1 §1)
"Treat the compiled program as a (program × LM) pair. Changing the LM invalidates the artifact — recompile." — dspy-sop SKILL, Dilemma Case B (R1 §2)
This is an enhancement overlay, not a framework SOP. It sits on top of whatever stack you use (DSPy, LangChain, raw API) and adds one discipline: define, pin, and version the entire deployable bundle as a unit. It is the broad sibling of [[agentsop-per-model-artifacts]] — that skill versions one compiled prompt; this one versions everything that ships together.
Activate when any of these appears in the user's intent, codebase, or workflow:
| Trigger | Signal |
|---|---|
| Preparing to deploy | "ship this to prod", a Dockerfile/deploy.yaml/serving entrypoint wrapping an LLM app, a release checklist |
| "What is running in prod?" | Nobody can name the exact prompt text + model snapshot + retriever config currently serving traffic |
| Reproducibility need | "reproduce the deploy from last quarter", an audit, a regulator asking what produced an output |
| Rollback need | Incident: prod behavior changed and the team needs the last known-good combination of components back |
| Drift symptoms | Score moved, no code change merged; or "we updated the prompt but forgot which model it was tuned for" |
| Multi-component apps | RAG + reranker + synthesizer + judge, each naming its own model/config, none bundled |
| Cross-skill bridges | DSPy save_program produced a compiled program → it is one component of the bundle; pin the rest. Per-prompt lifecycle handled by [[agentsop-per-model-artifacts]] → wrap as a bundle component here. |
Do NOT activate when:
v1.json is fine.The single most common mistake is reasoning about an LLM app the way you reason about a trained model. They are not the same shape.
| ML model | LLM app artifact |
|---|---|
One weights file (.pt, .safetensors) | A manifest over many parts |
| Identity = file hash | Identity = hash of the whole bundle |
| Mutates only on retrain | Each part mutates independently and silently |
| Versioned by a model registry | Versioned by a bundle manifest + tag |
The deployable artifact is:
┌──────────────── DEPLOYABLE BUNDLE (one tag) ────────────────┐
│ │
prompts compiled model pins retrieval eval-set
(text + programs (snapshot id config version
hashes) (program.json) per call site) (index ptr, (sha256)
│ embed model, │
│ top_k, reranker) │
│ │
└── version them TOGETHER, or you can't reproduce a deploy ──┘The load-bearing claim: the deployable artifact = prompts + compiled programs + model pins
Each part can change without touching the others (R1 §3, §4, §6):
gpt-4o → new snapshot) — prompt unchanged,
behavior changed.If these are versioned separately, "the deploy" is not a thing you can name. If they are versioned as one bundle with one id, the deploy is reproducible and rollback is atomic (R1 §8).
MLflow's Model Registry gives the right primitives: versioning, stage transitions
(Staging/Production/Archived), reproduce-from-config, compare-versions
[~/.claude/skills/mlflow/SKILL.md]. Borrow those primitives. But the analogy breaks on
heterogeneity: a registry versions one model + signature + run; the LLM bundle is a
composite of many models, prompts, configs, and an eval set. You can log the manifest into
MLflow as one "model", but the registry was built for the single-weights case. The overlay
exists to make the composite explicit.
Five steps. Each has an exit criterion. The output is one versioned, reproducible bundle.
List every component that affects behavior at runtime. Nothing implicit:
program.json or save_program dir + sha).Exit: a written component inventory — no "and whatever the dashboard says" gaps.
latest, at every call site (OP-2,
R1 §3). Cross-link [[agentsop-per-model-artifacts]] for per-prompt swap-test detail.save_program=True for portability (R1 §2).Exit: every component has an immutable identifier. Grep for aliases finds none.
Write a single manifest.<version>.json (OP-1) referencing every pinned component, and assign
one monotonic bundle id (OP-5). Components keep their own internal versions; the bundle has
one id production deploys atomically.
Exit: deploy/manifest.v<n>.json exists and a git tag deploy/v<n> points to it.
Record eval_set_sha + scores in the manifest (OP-4). A bundle without a pinned eval set has
no reproducible score. Cross-link [[agentsop-regression-gate]]: the gate re-runs this exact eval set
on the new bundle vs the parent bundle at PR time, and fails the PR on regression.
Exit: manifest shows eval_set_sha and the scores measured on it; the regression gate is
wired to that eval-set version.
Production deploy reads the bundle tag, not HEAD (OP-6, R1 §8). Keep prior bundles tagged and reachable. Rollback = re-point the deploy to the previous tag, which restores prompts + model + config + index pointer together — never a partial revert.
Exit: one-command rollback (deploy deploy/v<n-1>) restores a known-good combination.
| Trigger | Preparing to deploy; "what exactly is running in prod?" |
| Action | Write one deploy/manifest.<version>.json enumerating every behavior-affecting component with an immutable id (prompt hashes, compiled-program path+sha, model snapshots per call site, retrieval config, eval_set_sha, framework versions). The manifest IS the deployable unit. |
| Output | One grep-able answer to "what is in prod" — the bundle id resolves all components. |
| Evidence | R1 §1 (LLM app ≠ one weights file); R1 §7 (registry analogy + limit); [[agentsop-per-model-artifacts]] extends the (program × LM × dataset) triple to the full bundle. |
| Trigger | Any call site naming a model; CI sees an alias or latest. |
| Action | Pin dated snapshots everywhere (gpt-4o-2024-08-06, claude-3-7-sonnet-20250219); record each call site in the manifest. See [[agentsop-per-model-artifacts]] OP-2 for the per-prompt detail. |
| Output | manifest.models[] with provider/snapshot per call site; zero aliases. |
| Evidence | R1 §3: "the snapshot, not the alias, is the identity"; Anthropic does not roll aliases, OpenAI does — both bite silently. |
| Trigger | A behavior knob lives in a dashboard, env var, or notebook cell. |
| Action | Move every knob (top_k, chunk size, reranker on/off, temperature, system-prompt path, index pointer) into version-controlled config; manifest references config_sha. |
| Output | config/ under git; nothing behavioral outside VCS. |
| Evidence | R1 §6: Aider edit-format and LlamaIndex Settings are config artifacts too — "'artifact' generalizes beyond compiled JSON." |
| Trigger | Bundling a version; about to tag a release. |
| Action | Record eval_set_sha (sha256 of canonical eval bytes) + the scores measured on it. The regression gate ([[agentsop-regression-gate]]) re-runs THIS eval set on new bundle vs parent. |
| Output | manifest eval_set_sha + scores; reproducible "this bundle scored X on eval-set Y." |
| Evidence | R1 §5: a score is meaningful only relative to a fixed eval set; dspy-sop dev-set sizing; per-model-artifacts dataset_sha256_8. |
| Trigger | All components pinned; ready to ship. |
| Action | Assign one monotonic bundle id (git tag deploy/v<n> → manifest.v<n>.json). Components keep internal versions; the bundle deploys atomically. |
| Output | deploy/v7 tag resolving the whole manifest. |
| Evidence | R1 §8; per-model-artifacts OP-10 (per-prompt tag) elevated to a whole-bundle tag; MLflow stage-promotion analogy (R1 §7). |
| Trigger | Prod incident; need last known-good behavior. |
| Action | Deploy reads the bundle tag, not HEAD. Rollback = re-point to the previous tag (git checkout deploy/v6), restoring prompts + model + config + index pointer together. |
| Output | One-command atomic rollback; no "prompt reverted but model didn't." |
| Evidence | R1 §8: deploy reads the tag; the bundle tag makes the unit the whole combination, not one part. |
| Trigger | Score regressed; need a root-cause class. |
| Action | Diff manifest vN vs vN-1 field-by-field: prompt_sha → prompt drift; model snapshot → model drift; config_sha → config drift; eval_set_sha → apples-vs-oranges; index pointer → retrieval drift. |
| Output | Root-cause class from a manifest diff, before re-running anything. |
| Evidence | per-model-artifacts §4.4 lineage ops (classify dataset/config/framework/environment drift). |
| Trigger | App uses RAG; index built with a specific embedding model. |
| Action | Pin the index pointer + embed model + dim in the manifest; loader refuses to start if runtime embed model ≠ manifest embed model. |
| Output | manifest.retrieval = {index_path, embed_model, dim, corpus_sha8, top_k, reranker}. |
| Evidence | R1 §6; per-model-artifacts Case C: embedding model is part of index identity; mixing spaces is the worst-case silent failure. |
困境 (Dilemma): A RAG app has shipped on gpt-4o-2024-08-06 for the synthesizer. No code,
prompt, config, or model pin has changed in three months — every component's id in the
manifest is identical. Yet the quarterly canary on the pinned eval set drops from 78% to 73%.
Support escalations are rising. What changed, and what is the fix when nothing in the bundle
moved?
约束 (Constraints):
决策步骤 (Decision steps):
eval_set_sha, the
73% and the old 78% were measured on the same eval set. Rule out apples-vs-oranges first
(OP-7) — a changed eval_set_sha would be the more boring explanation.deploy/v<n+1> even though the model snapshot string is unchanged. The
built_at field carries the temporal lineage the snapshot string cannot.结果 (Outcome): Re-bundling recovers to ~79%. The pin protected against alias rollover and most drift; the eval-set-linked canary caught what the pin missed. Crucially, the bundle id incremented even though no human-authored component "changed" — the behavior changed, so the bundle changed.
可提取的操作 (Extractable operation): A snapshot pin is necessary but not sufficient. Tie every bundle to a pinned eval set and canary against it; when behavior drifts on a stable snapshot, mint a new bundle version rather than mutating the live one.
困境: Under deadline pressure, an engineer edits the synthesizer system prompt directly in the running config to fix one bad answer, redeploys, and moves on. No new bundle, no eval run, no tag. A week later a different regression appears in prod, and a model-snapshot rollover also landed in the same window. Now: which prompt is live? Was it ever evaluated? Can we roll back to before the prompt edit without also reverting unrelated changes?
约束:
决策步骤:
deploy/v<n> (OP-6, R1 §8). Because the bundle pins prompt and model
and config together, this restores a known-good combination — it does not produce the
untested (old-prompt × new-model) Frankenstein that a prompt-only revert would.deploy/v<n+1>.结果: Rollback restores service immediately because the bundle is atomic. The fix re-lands behind the gate. The lesson the team internalizes: a prompt change shipped without an eval-linked bundle is invisible to rollback and reproduction — it is exactly the failure mode this overlay exists to prevent.
可提取的操作: No behavior-affecting change ships outside a versioned, eval-linked bundle. Roll back by bundle tag (atomic) — never by reverting one component, which yields untested combinations.
"latest" (or any alias) model pin. gpt-4o, claude-3-5-sonnet, latest are moving
targets that change behavior with no version bump (R1 §3). Pin dated snapshots at every
call site (OP-2).eval_set_sha.save_program captures one component, not the bundle (R1 §2).v1.json is enough. Adopt this overlay at the second component or the first real deploy.How four mechanisms handle (or fail to handle) the whole-bundle versioning problem. None of them natively version the composite; each covers a slice.
save_program — one component, not the bundleprogram.json (or a save_program=True dir) is one component the
manifest references. The overlay pins the rest. Per-prompt lifecycle of this component →
[[agentsop-per-model-artifacts]].hub.pull("user/rag-qa:v3") resolves a prompt version, not a deploy.~/.claude/skills/mlflow/SKILL.md, mlflow.org/docs/latest/model-registry.html].deploy/manifest.<version>.json — prompt hashes, compiled
program shas, model snapshots per call site, retrieval config, eval_set_sha, framework
versions.deploy/v<n>)
gives atomic deploy + rollback (R1 §8); the regression gate ([[agentsop-regression-gate]]) wires to
the pinned eval set in CI.| Mechanism | What it versions | Bundle gap this overlay fills |
|---|---|---|
DSPy save_program | One compiled prompt | Model pins, retrieval config, eval-set, bundling |
| LangChain Hub | Prompt templates (diff UI) | Model pin enforcement, eval linkage, gate, bundling |
| MLflow Registry | One model + signature + run | Composite of many components; eval-set hash; one bundle tag |
| git + manifest | Whatever you list | Nothing — this overlay IS the recipe (manifest + tag + gate) |
{
"bundle_version": "v7", "built_at": "...", "commit": "<git-sha>", "parent_bundle": "v6",
"prompts": [{"name": "synth_system", "path": "prompts/synth_system.txt", "sha256": "..."}],
"compiled_programs": [{"name": "rag_synth", "path": "artifacts/rag_synth/.../v3.json", "sha256": "..."}],
"models": [
{"call_site": "router", "provider": "openai", "snapshot": "gpt-4o-mini-2024-07-18"},
{"call_site": "synthesizer", "provider": "openai", "snapshot": "gpt-4o-2024-08-06"},
{"call_site": "judge", "provider": "anthropic", "snapshot": "claude-3-7-sonnet-20250219"}
],
"retrieval": {"index_path": "indices/kb/.../<corpus-sha8>/", "embed_model": "text-embedding-3-small",
"dim": 1536, "top_k": 8, "reranker": "bge-reranker-v2", "corpus_sha8": "a8f1c2e9"},
"eval_set_sha": "9c3f...", "scores": {"dev": 0.81, "test": 0.79},
"frameworks": {"dspy": "2.6.1", "python": "3.11.8"}
}Preparing to deploy / "what's in prod?" ─► Step 1-3 (enumerate, pin, bundle)
│
Score regressed, nothing committed changed? ─► Case A: confirm eval_set unchanged,
│ check sample-hashes, re-bundle, gate
Prompt edited live without eval? ──────────► Case B: roll back by BUNDLE tag (atomic),
│ re-land fix behind the gate
Need rollback? ─────────────────────────────► OP-6: deploy previous bundle tag
│
Only ONE compiled prompt to version? ───────► use [[agentsop-per-model-artifacts]] instead
Need the CI comparison itself? ─────────────► use [[agentsop-regression-gate]]no-alias-pin: fail if any manifest model id lacks a date suffix or is latest (OP-2).manifest-matches-live: fail deploy if live component hashes ≠ tagged bundle (Case B).eval-set-pinned: fail bundle if eval_set_sha missing or scores recorded without it (OP-4).regression-gate: re-run pinned eval set on new bundle vs parent — see [[agentsop-regression-gate]].dspy-sop-skill/SKILL.md (R1 §1, §2).d-per-model-artifacts-skill/SKILL.md (R1 §3–§6).~/.claude/skills/mlflow/SKILL.md (R1 §7).references/R1-source-evidence.md.© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/agentsop-llm-artifact-versioning of agentsope/SkillAlchemy.
Open the folder on GitHubat commit d0f0355
Agentsop LLM Artifact Versioning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop LLM Artifact Versioning this skillagentsope/SkillAlchemy | 436 | — | ~7.1k | Automated safety check: Pass | MIT | |
| Build ML Pipelineprobabl-ai/skills | 138 | — | ~4.4k | Automated safety check: Pass | BSD-3-Clause | |
| ML Pipeline Workflowwshobson/agents | 40k | 12 repos | ~1.8k | Automated safety check: Pass | MIT | |
| Editomegaml/omegaml | 108 | — | ~206 | Automated safety check: Pass | Apache-2.0 | |
| ML Pipeline ExpertJeffallan/claude-skills | 12k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Machine Learning Ops ML Pipelineaiskillstore/marketplace | 433 | 7 repos | ~2.6k | Automated safety check: Pass | None |
probabl-ai/skills
Declare the pipeline from data source to predictor as a skrub DataOps graph.
wshobson/agents
Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.
omegaml/omegaml
how to use the edit command properly
Jeffallan/claude-skills
Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.
aiskillstore/marketplace
Design and implement a complete ML pipeline for: $ARGUMENTS. An agent skill from aiskillstore/marketplace.
bybren-llc/safe-agentic-workflow
Deployment workflows, pre-deploy validation, smoke testing, and rollback procedures. Use when deploying to staging or production, running smoke tests…
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Categories
Enhancement overlay — version the WHOLE deployable LLM-app artifact as one bundle: prompts + compiled programs + model snapshot pins + retrieval config + eval-set version, versioned together so a…. Agentsop LLM Artifact Versioning is an agent skill from agentsope/SkillAlchemy. Enhancement overlay — version the WHOLE deployable LLM-app artifact as one bundle: prompts + compiled programs + model snapshot pins + retrieval config + eval-set version, versioned together so a deploy is reproducible and rollback is atomic.
Agentsop LLM Artifact Versioning fits situations like: tasks that involve Machine learning; tasks that involve Operations and SOPs.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a claude-code`. Or copy the skill folder (skills/agentsop-llm-artifact-versioning in agentsope/SkillAlchemy) into .claude/skills/agentsop-llm-artifact-versioning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a codex`. Or copy the skill folder (skills/agentsop-llm-artifact-versioning in agentsope/SkillAlchemy) into .agents/skills/agentsop-llm-artifact-versioning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-llm-artifact-versioning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-llm-artifact-versioning, .gemini/skills/agentsop-llm-artifact-versioning, .github/skills/agentsop-llm-artifact-versioning and .opencode/skills/agentsop-llm-artifact-versioning in your project.
Going by SKILL.md and its folder, Agentsop LLM Artifact Versioning needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop LLM Artifact Versioning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.1k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop LLM Artifact Versioning: Build ML Pipeline (probabl-ai/skills, 138 stars), ML Pipeline Workflow (wshobson/agents, 40k stars), Edit (omegaml/omegaml, 108 stars) and ML Pipeline Expert (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.