Agent skill

Checkpoint Promotion

by wshobson in wshobson/agents

Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion.

MITAuto-check passedAgent Workflows

Install Checkpoint Promotion

skills CLI
$ npx skills add wshobson/agents --skill checkpoint-promotion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents checkpoint-promotion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/checkpoint-promotion .claude/skills/checkpoint-promotion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
checkpoint-promotion
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
1,024 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion.

  • Works in 4 steps: Data-quality gate. Before → **Held-out + frozen → Paired arena vs. base. → …
  • Agent Workflows work in your project
  • SKILL.md covers The Four-Stage Gate, Catastrophic Forgetting, The Verdict and Related Skills
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Checkpoint Promotion is an agent skill from wshobson/agents. Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/gate-templates.md`).

It sits in Agent Workflows. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “/checkpoint-promotion”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Data-quality gate. Before
  2. **Held-out + frozen
  3. Paired arena vs. base.
  4. Canary. 5–10% stratified

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Checkpoint Promotion loads about 2k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 1,024 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,024 words, ~2,012 tokens.

Download SKILL.mdSave it as .claude/skills/checkpoint-promotion/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
checkpoint-promotion
description
Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.

Checkpoint Promotion

The Phase 5 gate for the whole plugin: a checkpoint that trains cleanly and beats its task metric still doesn't ship without clearing all four stages below. eval-harness-first built the suite re-run here — this skill is where that suite's baseline decides something.

Input: a trained checkpoint, eval/baseline-<model>.json from eval-harness-first, and the frozen eval/drift-suite.yaml. Output format: promotion-report.md — the four-stage evidence plus a terminal PROMOTE or REJECT verdict that /finetune Phase 5 and /promote-checkpoint consume directly.

The Four-Stage Gate

Each stage gates the next — a failure at stage 2 means stage 3 doesn't run. Stages 2 and 3 share one expensive inference pass, so running them concurrently and applying gate order at verdict time is licensed on a deterministic arena (nothing saved by serializing); a judge-based arena should still wait for stage 2 first — that's where the real savings are.

  1. Data-quality gate. Before any eval touches the checkpoint: dedup the training set, check for eval-goldens leakage (the exact failure trace-to-training-data's Hygiene section exists to prevent), and scan for label noise. A checkpoint trained on leaked goldens invalidates every later stage.
  2. Held-out + frozen capability-drift suite. Re-run eval-harness-first's eval/drift-suite.yaml — MMLU/GSM8K/IFEval plus 200–500 domain-adjacent items — against the checkpoint and diff against baseline-<model>.json per benchmark against the Drift Budget table below.
  3. Paired arena vs. base. Position-randomized judge, checkpoint vs. base model, same prompts — or the deterministic paired-comparison variant in references/gate-templates.md when every grader in the harness is deterministic (no LLM-judge; position randomization N/A there). A holdout win that loses the live arena does not ship — stage-2 numbers and stage-3 judgments must agree; a win on frozen goldens and a loss in paired comparison is a real signal, not a discrepancy to explain away.
  4. Canary. 5–10% stratified rollout with auto-rollback for any checkpoint reaching production traffic. Local-only users stop at stage 3 — skipping stage 4 for a local deployment is the correct stopping point, not a shortcut.
Drift Budget
Drift (pts)Verdict
≤1Noise — proceed
2–5Rerun with seed variation before deciding
>5HARD FAIL — no exception for task gains

The >5pt row governs regardless of the others: a checkpoint that gained 8 points on the target task and lost 6 points of general capability still fails here — task improvement never buys back a drift-budget breach.

Item count derives from the budget, not convenience: the strict n for a half-width under half the 5pt hard-fail threshold is ~1,300 at typical accuracy (p≈0.7); n=200 is a pragmatic floor (±6pt half-width at that same p, n=50 ±13pt) — report the half-width with every verdict, and treat a margin smaller than it as REJECT (uncertain), not PASS/HARD FAIL. Full math and a 5-run cautionary example: references/gate-templates.md.

RERUN is not a verdict. A 2–5pt drift only ever produces a PROMOTE or REJECT after the seed-variation rerun completes — PROMOTE requires landing back at ≤1pt (noise); any rerun still

1pt — 2–5pt band or >5pt breach alike — resolves stage 2 to a hard REJECT. No report may reach the Verdict section with stage 2 still showing RERUN.

Show full SKILL.md (526 more words)Show less

Catastrophic Forgetting

Unmanaged LoRA fine-tuning loses real general capability, and stage 2 is what catches it:

  • ~43% knowledge loss unmanaged — no replay, no regularization.
  • ~10% with basic management — some replay or a conservative LR.
  • ~3% with replay + EWC — the disciplined case.
  • 10–30% general-data replay mix is the standard mitigation — blend general- domain data into training rather than target-task data alone.

If a checkpoint hits the >5pt hard fail in stage 2, work this escalation ladder in order — the one canonical order this skill and references/gate-templates.md both point to:

  1. Adjust the replay-mix fraction — swap rows, don't add them (adding confounds fraction with total optimizer steps). Dose is not monotonic at small-run scale (<~100 steps) — re-check drift after any swap.
  2. Lower the learning rate.
  3. Fewer epochs.
  4. A smaller LoRA rank — the same rank/LR levers lora-qlora-recipes and preference-optimization tune for the training run, applied here in reverse.

This order is a default, not a law: remediation guidance from a single before/after run pair is a hypothesis — label it low-confidence once any lever produces a reversal, and prefer a seed-variation repeat over trusting the next rung blindly. A lever that clears the drift breach but drops a success-criterion metric below target is a two-sided tradeoff for a human, not a reason to keep descending the ladder. Full reasoning and the 5-run trajectory behind both caveats: references/gate-templates.md.

Disclose drift-suite instruction reuse. A replay row copying the drift harness's exact instruction phrasing (not just disjoint source items) makes that benchmark's post-replay score an upper bound — flag it instruction-familiar, or re-probe with a paraphrase, before treating a near-budget pass as clean.

The Verdict

promotion-report.md covers all four stages as sections and must end with a terminal verdict: PROMOTE or REJECT, the evidence that produced it, and exactly one top remediation when the verdict is REJECT. Template: references/gate-templates.md. The terminal contract other skills parse:

## Verdict

REJECT

Evidence: domain-adjacent drift
suite dropped 6.2pt (threshold:
>5pt hard fail) despite +8pt on
the target task.

Top remediation: swap the
replay-mix fraction from 10%
toward 20%, holding step count
constant.
  • REJECT is a result, not an error. A checkpoint that fails stage 2's drift budget or stage 3's arena comparison did its job. Don't treat a REJECT as a failed run needing a rerun of this skill; it's the correct output of a working gate.
  • One remediation, not a menu. Evidence sections may list everything observed; the verdict section names the single highest-leverage fix per the escalation ladder above. A report that hedges across three possible fixes hasn't done the prioritization this skill exists to do.
  • No auto-retraining. This skill produces a verdict and a report, not a re-triggered training run. A REJECT hands the remediation back to a human decision at finetuning-method-selection or the relevant training skill.
  • eval-harness-first — owns the drift suite and baseline this skill re-runs and diffs against; no baseline-<model>.json means nothing to gate against.
  • quantized-export — the only valid next step after a PROMOTE verdict.
  • preference-optimization and lora-qlora-recipes — own the LR and rank levers in the Catastrophic Forgetting escalation path; this skill diagnoses the breach, those skills own the config that caused it.
  • dataset-curation — owns the replay-mix construction recipe the escalation ladder's first rung applies.

Complete promotion-report.md template with all four stages, the drift-suite scoring table, the paired-arena protocol (item count, position randomization, win-rate threshold), and a replay-mix configuration example: references/gate-templates.md.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/llm-finetuning/skills/checkpoint-promotion of wshobson/agents.

  • SKILL.md
  • references/gate-templates.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Checkpoint Promotion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Checkpoint Promotion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Checkpoint Promotion this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k10 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k35 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers297k2 repos~5.1kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79689 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 10 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 35 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    297k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    796 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 5 days ago
    Auto-check passed

Categories

Questions about Checkpoint Promotion

What does Checkpoint Promotion do?

Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Checkpoint Promotion is an agent skill from wshobson/agents. Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion.

When should I use Checkpoint Promotion?

Checkpoint Promotion fits situations like: agent Workflows work in your project.

How do I install Checkpoint Promotion in Claude Code?

Run `npx skills add wshobson/agents --skill checkpoint-promotion -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/checkpoint-promotion in wshobson/agents) into .claude/skills/checkpoint-promotion in your project. Claude Code loads it when a task matches its description.

How do I install Checkpoint Promotion in Codex?

Run `npx skills add wshobson/agents --skill checkpoint-promotion -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/checkpoint-promotion in wshobson/agents) into .agents/skills/checkpoint-promotion in your project. Codex loads it when a task matches its description.

Can I use Checkpoint Promotion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill checkpoint-promotion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/checkpoint-promotion, .gemini/skills/checkpoint-promotion, .github/skills/checkpoint-promotion and .opencode/skills/checkpoint-promotion in your project.

What does Checkpoint Promotion need to run?

SKILL.md names no scripts, command-line tools or credentials: Checkpoint Promotion is instructions for the agent only.

Does Checkpoint Promotion access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Checkpoint Promotion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Checkpoint Promotion use?

Checkpoint Promotion is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Checkpoint Promotion use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Checkpoint Promotion?

Skills that share tags, products or a category with Checkpoint Promotion: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Checkpoint Promotion?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.