Agent skill

Deploy Hero Change

by marin-community in marin-community/marin

Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial…

Apache-2.0Auto-check passedDevOps & Cloud

Install Deploy Hero Change

skills CLI
$ npx skills add marin-community/marin --skill deploy-hero-change -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install marin-community/marin deploy-hero-change --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deploy-hero-change .claude/skills/deploy-hero-change && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deploy-hero-change
GitHub stars
3.9k
Token cost
~1.9k tokens
SKILL.md length
1,068 words
Files
2 (incl. references)
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial…

  • Works in 5 steps: Agree the plan → Land the launch record → Rehearse preflight and rollback → …
  • Asks to deploy a change to the hero
  • SKILL.md covers 1. Agree the plan, 2. Land the launch record, 3. Rehearse preflight and… and 4. Execute the trial, plus 1 more section
  • Calls python

What it does

Deploy Hero Change is an agent skill from marin-community/marin. Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial window, and roll back if the gate fails; use only when the user asks to deploy a change to the hero.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/silent-hangs.md`).

It sits in DevOps & Cloud. The repository describes itself as: Open-source framework for the research and development of foundation models. The licence is Apache-2.0.

When your agent uses it

  • Asks to deploy a change to the hero

Example prompts

  • “/deploy-hero-change”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Agree the plan
  2. Land the launch record
  3. Rehearse preflight and rollback
  4. Execute the trial
  5. Decide

What it can do on your machine

Read from SKILL.md and the folder at commit 61bb85c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deploy Hero Change loads about 1.9k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 1,068 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from marin-community/marin at commit 61bb85c, republished under its Apache-2.0 licence (© marin-community). 1,068 words, ~1,917 tokens.

Download SKILL.mdSave it as .claude/skills/deploy-hero-change/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
deploy-hero-change
description
Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial window, and roll back if the gate fails; use only when the user asks to deploy a change to the hero.

Deploy a change to the live hero run

Use manage-hero-run for the run record, DRI, retention, and babysitting. This checklist governs a code cutover and its trial. Single-rack validation does not exercise cross-domain collectives or checkpoint recovery at production scale.

1. Agree the plan

Record the following with the user before changing the live run:

  • A complete permanent handoff checkpoint step-N. Use the newest scheduled permanent checkpoint (every 6000 steps, about 26 hours at the hero's pace), or request one from the old run's training-control endpoint with the request-permanent-checkpoint header value. Only a run whose code includes that action can take the request; otherwise schedule the cutover shortly after a scheduled permanent checkpoint. Do not hand off from a temporary checkpoint: it expires three days after it is written.
  • A new run ID and checkpoint tree, the unchanged W&B entity/project, and the verified parent history boundary. Follow the launcher procedure.
  • A matched control window, 200 steps by default. The old run must record those steps after the handoff checkpoint before it is stopped.
  • Explicit gates for loss, throughput/MFU, token drops, router metrics, and the intended effect of the change. Specify tolerances and expected directions; require no unexpected crash, retry, watchdog, or alert.
  • Coverage gaps: evals, the child's own save/resume, and retries may fall outside the trial. Include these limitations in the decision.
  • A schedule that leaves time for rollback during working hours, a submission owner, and the status issue/communication thread. Report kill, launch, first steps, go/no-go, and rollback transitions.

2. Land the launch record

Land the code change and finalized trigger_hero.sh values on main before cutover. Use a pristine checkout at the verified SHA. The launcher records the new run, retained checkpoint, and W&B fork boundary; its commands and recovery semantics are documented in the linked launcher procedure.

Inventory downstream reports and trackers that pin the run ID or affected metric keys. Update their selection as part of accepting the child.

3. Rehearse preflight and rollback

Use small scripts whose queries fail closed on errors or unrecognized output. Dry-run every guard against the live cluster and obtain an independent review.

Verify:

  • Deploy checkout is clean at fetched main; rollback checkout is clean at the old run's recorded SHA. Preserve its original launch command.
  • Handoff metadata.json exists, records the expected step, and has the intended retention. Confirm layout and checkpoint lineage.
  • Child checkpoint trees are empty for initial cutover. During later recovery, preserve them and verify the newest complete child checkpoint instead.
  • No competing coordinator, gang, or hero pods exist apart from the old run. Iris, Kubernetes, object-store, and W&B credentials work.
  • Launch refuses a live parent or child coordinator and a dirty checkout. One operator owns submission; capture its output and verify exactly one child coordinator afterward. Resolve an uncertain submission before retrying.
  • Rollback cancels the child coordinator, confirms it is terminal, and launches the old revision and run ID as IRIS_USER=marin. Verify its intended resume checkpoint first. Never create another W&B fork for rollback.
  • Object-storage headroom, checked again right before requesting the handoff. Temporary checkpoints, which the hero writes hourly, go to hero-checkpoints in US-EAST-08A (100 TiB quota); permanent checkpoints go to marin-us-east-02a in US-EAST-02A, whose quota all Marin work shares. At quota CoreWeave suspends writes for the whole zone, and the hero's next save hangs without an error (#8506, 2026-09-23). Require five checkpoints of free space in 08A (about 21 TB at 4.29 TB each): the old run's newest temporary checkpoint, a restore-smoke copy, the child's temporary checkpoint plus the next one being written (the older is pruned only after the newer commits), and one spare. Require two checkpoints plus a day of recent growth in 02A for the handoff and the child's next permanent checkpoint. Read usage and quota from the Finelog storage.usage namespace; the collector runs every few hours, so if the newest collected_at is more than an hour old, get current values with python -m scripts.ops.storage.coreweave_usage --dry-run.

Distinguish a successful query with no matching jobs from a failed query. Parse Iris CSV headers and CRLF correctly; do not interpret grep -c exit status as a query result. Match pods by task identity because Kubernetes names are sanitized and truncated. Confirm restore from Loading checkpoint from and entry into the training loop.

Show full SKILL.md (362 more words)Show less

4. Execute the trial

  1. Pass preflight. Verify the full matched control window is recorded and create the W&B fork once using the launcher procedure.
  2. Cancel the old coordinator and confirm it is terminal and its training tasks have stopped before launching the child from the recorded SHA as IRIS_USER=marin.
  3. Allow for restore, compilation, and data prefetch. Read startup and step watchdog limits from the resolved launch configuration; do not treat normal compilation as a hang.
  4. Before the first child step, publish a comparison report for the 200 updates starting at N (steps N through N+199), plus wider context. Include loss, cross entropy, MFU, step time/tokens per second, drops, routing entropy, router losses, gradient norm, and peak memory. Share the URL.
  5. Monitor fresh Finelog rows and execution identity, job state, task attempt IDs, step age, watchdogs, and errors. W&B inherited history and API lag do not prove new progress. Compare gate metrics against the old run at matching steps; emit updates on change.
  6. Join the two runs by step and report mean/max loss deltas and all gate metrics. Judge restore correctness against the paired control or thresholds agreed before launch.

5. Decide

The trial gate overrides manage-hero-run's ordinary retry policy. On a failed gate, unexplained hang, or retry loop, cancel the child coordinator and roll back without waiting for further attempts. Retries may restore either the handoff or a newer child checkpoint; neither substitutes for a successful trial.

Go: leave the child running, publish the comparison and coverage gaps, update the status issue and downstream reporting, and resume ordinary recovery policy. The old run's leftover temporary checkpoints expire on their own within three days, or 14 if its launcher predates the three-day TTL. Delete any restore-smoke copy once it has served its purpose, deleting metadata.json first so a partial directory never looks complete.

After rollback, compare the old run's first replayed steps with its earlier trajectory. This checks trajectory consistency, not bitwise determinism. Its W&B history may not advance until it passes the old counter, so verify progress in Finelog. Update the status issue and file the failure with supporting evidence.

For collective stalls, use silent-hang diagnostics.

© marin-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .agents/skills/deploy-hero-change of marin-community/marin.

  • SKILL.md
  • references/silent-hangs.md

Open the folder on GitHubat commit 61bb85c

Compare with similar skills

Deploy Hero Change next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deploy Hero Change compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deploy Hero Change this skillmarin-community/marin3.9k—~1.9kAutomated safety check: PassApache-2.0
Monitor CInrwl/nx29k6 repos~4.7kAutomated safety check: PassMIT
Terraform and OpenTofu Guideagentscope-ai/QwenPaw36k6 repos~4.2kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence
Openclaw Live Updateropenclaw/openclaw392k—~3.7kAutomated safety check: PassMIT

Similar skills

  • Monitor CI

    nrwl/nx

    Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.

    29k GitHub starsUsed in 6 repos~4.7k tokens
    DevOps & CloudAuto-check passed
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    36k GitHub starsUsed in 6 repos~4.2k tokens
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed

More from marin-community/marin

All 41 skills in this repo
  • Noslop

    marin-community/marin

    Deslop, simplify, or review low-value tests and prose only when explicitly requested for a branch or diff.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Use Iris

    marin-community/marin

    Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs.

    3.9k GitHub stars~745 tokensUpdated today
    Auto-check passed
  • Launch Rl

    marin-community/marin

    Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.

    3.9k GitHub stars~894 tokensUpdated today
    Auto-check passed
  • Marina Applet

    marin-community/marin

    Build, validate, publish, update, inspect, query, roll back, or archive a dynamic Marina applet.

    3.9k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Query Finelog

    marin-community/marin

    Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Trace Pulumi Diff

    marin-community/marin

    Run a read-only preview for a specified Marin infra/pulumi stack and trace each pending resource change to merged pull requests since its latest successful update when that update records a clean…

    3.9k GitHub stars~663 tokensUpdated today
    Auto-check passed

Categories

Questions about Deploy Hero Change

What does Deploy Hero Change do?

Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial…. Deploy Hero Change is an agent skill from marin-community/marin. Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial window, and roll back if the gate fails; use only when the user asks to deploy a change to the hero.

When should I use Deploy Hero Change?

Deploy Hero Change fits situations like: asks to deploy a change to the hero.

How do I install Deploy Hero Change in Claude Code?

Run `npx skills add marin-community/marin --skill deploy-hero-change -a claude-code`. Or copy the skill folder (.agents/skills/deploy-hero-change in marin-community/marin) into .claude/skills/deploy-hero-change in your project. Claude Code loads it when a task matches its description.

How do I install Deploy Hero Change in Codex?

Run `npx skills add marin-community/marin --skill deploy-hero-change -a codex`. Or copy the skill folder (.agents/skills/deploy-hero-change in marin-community/marin) into .agents/skills/deploy-hero-change in your project. Codex loads it when a task matches its description.

Can I use Deploy Hero Change in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add marin-community/marin --skill deploy-hero-change -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deploy-hero-change, .gemini/skills/deploy-hero-change, .github/skills/deploy-hero-change and .opencode/skills/deploy-hero-change in your project.

What does Deploy Hero Change need to run?

Going by SKILL.md and its folder, Deploy Hero Change needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Deploy Hero Change access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Deploy Hero Change safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deploy Hero Change use?

Deploy Hero Change is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deploy Hero Change use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 411 tokens, read only when the agent opens those files.

What are the alternatives to Deploy Hero Change?

Skills that share tags, products or a category with Deploy Hero Change: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 36k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deploy Hero Change?

marin-community (a GitHub organization) maintains it in marin-community/marin, which has 3,920 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 9, 2026.

Source: marin-community/marin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.