MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion.
$ npx skills add wshobson/agents --skill checkpoint-promotion -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents checkpoint-promotion --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/checkpoint-promotion .claude/skills/checkpoint-promotion && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "checkpoint-promotion" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/checkpoint-promotion into .claude/skills/checkpoint-promotion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "checkpoint-promotion", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/checkpoint-promotionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill checkpoint-promotion -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents checkpoint-promotion --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/llm-finetuning/skills/checkpoint-promotion .agents/skills/checkpoint-promotion && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "checkpoint-promotion" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/checkpoint-promotion into .agents/skills/checkpoint-promotion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "checkpoint-promotion", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill checkpoint-promotion -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents checkpoint-promotion --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/llm-finetuning/skills/checkpoint-promotion .cursor/skills/checkpoint-promotion && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "checkpoint-promotion" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/checkpoint-promotion into .cursor/skills/checkpoint-promotion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "checkpoint-promotion", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/llm-finetuning/skills/checkpoint-promotion--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill checkpoint-promotion -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents checkpoint-promotion --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/llm-finetuning/skills/checkpoint-promotion .gemini/skills/checkpoint-promotion && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "checkpoint-promotion" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/checkpoint-promotion into .gemini/skills/checkpoint-promotion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "checkpoint-promotion", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents checkpoint-promotionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill checkpoint-promotion -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/llm-finetuning/skills/checkpoint-promotion .github/skills/checkpoint-promotion && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "checkpoint-promotion" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/checkpoint-promotion into .github/skills/checkpoint-promotion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "checkpoint-promotion", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill checkpoint-promotion -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents checkpoint-promotion --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/llm-finetuning/skills/checkpoint-promotion .opencode/skills/checkpoint-promotion && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "checkpoint-promotion" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/checkpoint-promotion into .opencode/skills/checkpoint-promotion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "checkpoint-promotion", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
checkpoint-promotionGate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion.
Checkpoint Promotion is an agent skill from wshobson/agents. Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/gate-templates.md`).
It sits in Agent Workflows. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Checkpoint Promotion loads about 2k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 1,024 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,024 words, ~2,012 tokens.
.claude/skills/checkpoint-promotion/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.The Phase 5 gate for the whole
plugin: a checkpoint that trains
cleanly and beats its task metric
still doesn't ship without
clearing all four stages below.
eval-harness-first built the
suite re-run here — this skill is
where that suite's baseline
decides something.
Input: a trained checkpoint,
eval/baseline-<model>.json from
eval-harness-first, and the
frozen eval/drift-suite.yaml.
Output format:
promotion-report.md — the
four-stage evidence plus a
terminal PROMOTE or REJECT
verdict that /finetune Phase 5
and /promote-checkpoint consume
directly.
Each stage gates the next — a failure at stage 2 means stage 3 doesn't run. Stages 2 and 3 share one expensive inference pass, so running them concurrently and applying gate order at verdict time is licensed on a deterministic arena (nothing saved by serializing); a judge-based arena should still wait for stage 2 first — that's where the real savings are.
trace-to-training-data's
Hygiene section exists to
prevent), and scan for label
noise. A checkpoint trained on
leaked goldens invalidates
every later stage.eval-harness-first's
eval/drift-suite.yaml —
MMLU/GSM8K/IFEval plus 200–500
domain-adjacent items — against
the checkpoint and diff against
baseline-<model>.json per
benchmark against the Drift
Budget table below.references/gate-templates.md
when every grader in the
harness is deterministic (no
LLM-judge; position
randomization N/A there).
A holdout win that
loses the live arena does not
ship — stage-2 numbers and
stage-3 judgments must agree; a
win on frozen goldens and a
loss in paired comparison is a
real signal, not a discrepancy
to explain away.| Drift (pts) | Verdict |
|---|---|
| ≤1 | Noise — proceed |
| 2–5 | Rerun with seed variation before deciding |
| >5 | HARD FAIL — no exception for task gains |
The >5pt row governs regardless of the others: a checkpoint that gained 8 points on the target task and lost 6 points of general capability still fails here — task improvement never buys back a drift-budget breach.
Item count derives from the
budget, not convenience: the
strict n for a half-width under
half the 5pt hard-fail threshold
is ~1,300 at typical accuracy
(p≈0.7); n=200 is a pragmatic
floor (±6pt half-width at that
same p, n=50 ±13pt) — report the
half-width with every verdict,
and treat a margin smaller than
it as REJECT (uncertain), not
PASS/HARD FAIL. Full math and a
5-run cautionary example:
references/gate-templates.md.
RERUN is not a verdict. A
2–5pt drift only ever produces a
PROMOTE or REJECT after the
seed-variation rerun completes —
PROMOTE requires landing back
at ≤1pt (noise); any rerun still
1pt — 2–5pt band or >5pt breach alike — resolves stage 2 to a hard
REJECT. No report may reach the Verdict section with stage 2 still showingRERUN.
Unmanaged LoRA fine-tuning loses real general capability, and stage 2 is what catches it:
If a checkpoint hits the >5pt
hard fail in stage 2, work this
escalation ladder in order — the
one canonical order this skill
and references/gate-templates.md
both point to:
lora-qlora-recipes and
preference-optimization tune
for the training run, applied
here in reverse.This order is a default, not a
law: remediation guidance from
a single before/after run pair
is a hypothesis — label it
low-confidence once any lever
produces a reversal, and prefer
a seed-variation repeat over
trusting the next rung blindly.
A lever that clears the drift
breach but drops a
success-criterion metric below
target is a two-sided tradeoff
for a human, not a reason to
keep descending the ladder. Full
reasoning and the 5-run
trajectory behind both caveats:
references/gate-templates.md.
Disclose drift-suite instruction reuse. A replay row copying the drift harness's exact instruction phrasing (not just disjoint source items) makes that benchmark's post-replay score an upper bound — flag it instruction-familiar, or re-probe with a paraphrase, before treating a near-budget pass as clean.
promotion-report.md covers all
four stages as sections and
must end with a terminal
verdict: PROMOTE or REJECT,
the evidence that produced it,
and exactly one top remediation
when the verdict is REJECT.
Template: references/gate-templates.md.
The terminal contract other
skills parse:
## Verdict
REJECT
Evidence: domain-adjacent drift
suite dropped 6.2pt (threshold:
>5pt hard fail) despite +8pt on
the target task.
Top remediation: swap the
replay-mix fraction from 10%
toward 20%, holding step count
constant.REJECT hands
the remediation back to a human
decision at
finetuning-method-selection or
the relevant training skill.eval-harness-first — owns the
drift suite and baseline this
skill re-runs and diffs
against; no baseline-<model>.json
means nothing to gate against.quantized-export — the only
valid next step after a
PROMOTE verdict.preference-optimization and
lora-qlora-recipes — own the
LR and rank levers in the
Catastrophic Forgetting
escalation path; this skill
diagnoses the breach, those
skills own the config that
caused it.dataset-curation — owns the
replay-mix construction recipe
the escalation ladder's first
rung applies.Complete promotion-report.md
template with all four stages,
the drift-suite scoring table,
the paired-arena protocol (item
count, position randomization,
win-rate threshold), and a
replay-mix configuration example:
references/gate-templates.md.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/llm-finetuning/skills/checkpoint-promotion of wshobson/agents.
Open the folder on GitHubat commit 46891e7
Checkpoint Promotion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Checkpoint Promotion this skillwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| MCP Server Builderanthropics/skills | 180k | 63 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Hook Development for Claude Code Pluginsanthropics/claude-plugins-official | 38k | 10 repos | ~4.1k | Automated safety check: Notes | Apache-2.0 | |
| Using Superpowersfarm-fe/farm | 5.6k | 35 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Executing Plans Inlineobra/superpowers | 297k | 2 repos | ~5.1k | Automated safety check: Pass | MIT | |
| Skill CreatorAzure/azqr | 796 | 89 repos | ~8.2k | Automated safety check: Pass | Apache-2.0 |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
anthropics/claude-plugins-official
Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.
farm-fe/farm
A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions
obra/superpowers
Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.
Azure/azqr
Create new skills, modify and improve existing skills, and measure skill performance.
anthropics/claude-plugins-official
Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
Categories
Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Checkpoint Promotion is an agent skill from wshobson/agents. Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion.
Checkpoint Promotion fits situations like: agent Workflows work in your project.
Run `npx skills add wshobson/agents --skill checkpoint-promotion -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/checkpoint-promotion in wshobson/agents) into .claude/skills/checkpoint-promotion in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill checkpoint-promotion -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/checkpoint-promotion in wshobson/agents) into .agents/skills/checkpoint-promotion in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill checkpoint-promotion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/checkpoint-promotion, .gemini/skills/checkpoint-promotion, .github/skills/checkpoint-promotion and .opencode/skills/checkpoint-promotion in your project.
SKILL.md names no scripts, command-line tools or credentials: Checkpoint Promotion is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Checkpoint Promotion is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Checkpoint Promotion: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.