Monitor CI
nrwl/nx
Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.
Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial…
$ npx skills add marin-community/marin --skill deploy-hero-change -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install marin-community/marin deploy-hero-change --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deploy-hero-change .claude/skills/deploy-hero-change && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "deploy-hero-change" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/deploy-hero-change into .claude/skills/deploy-hero-change/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-hero-change", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/marin-community/marin/tree/main/.agents/skills/deploy-hero-changeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add marin-community/marin --skill deploy-hero-change -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install marin-community/marin deploy-hero-change --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/deploy-hero-change .agents/skills/deploy-hero-change && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "deploy-hero-change" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/deploy-hero-change into .agents/skills/deploy-hero-change/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-hero-change", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add marin-community/marin --skill deploy-hero-change -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install marin-community/marin deploy-hero-change --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/deploy-hero-change .cursor/skills/deploy-hero-change && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "deploy-hero-change" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/deploy-hero-change into .cursor/skills/deploy-hero-change/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-hero-change", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/marin-community/marin.git --path .agents/skills/deploy-hero-change--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add marin-community/marin --skill deploy-hero-change -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install marin-community/marin deploy-hero-change --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/deploy-hero-change .gemini/skills/deploy-hero-change && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "deploy-hero-change" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/deploy-hero-change into .gemini/skills/deploy-hero-change/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-hero-change", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install marin-community/marin deploy-hero-changeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add marin-community/marin --skill deploy-hero-change -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/deploy-hero-change .github/skills/deploy-hero-change && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "deploy-hero-change" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/deploy-hero-change into .github/skills/deploy-hero-change/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-hero-change", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add marin-community/marin --skill deploy-hero-change -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install marin-community/marin deploy-hero-change --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/deploy-hero-change .opencode/skills/deploy-hero-change && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "deploy-hero-change" agent skill from https://github.com/marin-community/marin/tree/main/.agents/skills/deploy-hero-change into .opencode/skills/deploy-hero-change/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deploy-hero-change", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
deploy-hero-changeDeploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial…
Deploy Hero Change is an agent skill from marin-community/marin. Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial window, and roll back if the gate fails; use only when the user asks to deploy a change to the hero.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/silent-hangs.md`).
It sits in DevOps & Cloud. The repository describes itself as: Open-source framework for the research and development of foundation models. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 61bb85c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Deploy Hero Change loads about 1.9k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 1,068 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from marin-community/marin at commit 61bb85c, republished under its Apache-2.0 licence (© marin-community). 1,068 words, ~1,917 tokens.
.claude/skills/deploy-hero-change/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use manage-hero-run for the run record, DRI, retention, and babysitting. This
checklist governs a code cutover and its trial. Single-rack validation does not
exercise cross-domain collectives or checkpoint recovery at production scale.
Record the following with the user before changing the live run:
step-N. Use the newest scheduled
permanent checkpoint (every 6000 steps, about 26 hours at the hero's pace), or
request one from the old run's training-control endpoint with the
request-permanent-checkpoint header value. Only a run whose code includes
that action can take the request; otherwise schedule the cutover shortly after
a scheduled permanent checkpoint. Do not hand off from a temporary checkpoint:
it expires three days after it is written.Land the code change and finalized trigger_hero.sh values on main before
cutover. Use a pristine checkout at the verified SHA. The launcher records the
new run, retained checkpoint, and W&B fork boundary; its commands and recovery
semantics are documented in the linked launcher procedure.
Inventory downstream reports and trackers that pin the run ID or affected metric keys. Update their selection as part of accepting the child.
Use small scripts whose queries fail closed on errors or unrecognized output. Dry-run every guard against the live cluster and obtain an independent review.
Verify:
metadata.json exists, records the expected step, and has the intended
retention. Confirm layout and checkpoint lineage.IRIS_USER=marin. Verify its intended resume
checkpoint first. Never create another W&B fork for rollback.hero-checkpoints
in US-EAST-08A (100 TiB quota); permanent checkpoints go to
marin-us-east-02a in US-EAST-02A, whose quota all Marin work shares. At
quota CoreWeave suspends writes for the whole zone, and the hero's next save
hangs without an error (#8506, 2026-09-23).
Require five checkpoints of free space in 08A (about 21 TB at 4.29 TB each):
the old run's newest temporary checkpoint, a restore-smoke copy, the child's
temporary checkpoint plus the next one being written (the older is pruned only
after the newer commits), and one spare. Require two checkpoints plus a day of
recent growth in 02A for the handoff and the child's next permanent
checkpoint. Read usage and quota from the Finelog storage.usage namespace;
the collector runs every few hours, so if the newest collected_at is more
than an hour old, get current values with
python -m scripts.ops.storage.coreweave_usage --dry-run.Distinguish a successful query with no matching jobs from a failed query. Parse
Iris CSV headers and CRLF correctly; do not interpret grep -c exit status as a
query result. Match pods by task identity because Kubernetes names are sanitized
and truncated. Confirm restore from Loading checkpoint from and entry into the
training loop.
IRIS_USER=marin.N (steps N through N+199), plus wider context. Include loss,
cross entropy, MFU, step time/tokens per second, drops, routing entropy, router
losses, gradient norm, and peak memory. Share the URL.The trial gate overrides manage-hero-run's ordinary retry policy. On a failed
gate, unexplained hang, or retry loop, cancel the child coordinator and roll back
without waiting for further attempts. Retries may restore either the handoff or
a newer child checkpoint; neither substitutes for a successful trial.
Go: leave the child running, publish the comparison and coverage gaps, update the
status issue and downstream reporting, and resume ordinary recovery policy.
The old run's leftover temporary checkpoints expire on their own within three
days, or 14 if its launcher predates the three-day TTL. Delete any restore-smoke
copy once it has served its purpose, deleting metadata.json first so a partial
directory never looks complete.
After rollback, compare the old run's first replayed steps with its earlier trajectory. This checks trajectory consistency, not bitwise determinism. Its W&B history may not advance until it passes the old counter, so verify progress in Finelog. Update the status issue and file the failure with supporting evidence.
For collective stalls, use silent-hang diagnostics.
© marin-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in .agents/skills/deploy-hero-change of marin-community/marin.
Open the folder on GitHubat commit 61bb85c
Deploy Hero Change next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Deploy Hero Change this skillmarin-community/marin | 3.9k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Monitor CInrwl/nx | 29k | 6 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Terraform and OpenTofu Guideagentscope-ai/QwenPaw | 36k | 6 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Vercel Optimize Auditvercel-labs/agent-skills | 32k | 8 repos | ~4.3k | Automated safety check: Pass | None | |
| Analyze GitHub Action Logswithastro/astro | 63k | 1 repos | ~1.3k | Automated safety check: Pass | Custom licence | |
| Openclaw Live Updateropenclaw/openclaw | 392k | — | ~3.7k | Automated safety check: Pass | MIT |
nrwl/nx
Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.
agentscope-ai/QwenPaw
Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.
vercel-labs/agent-skills
Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.
withastro/astro
Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.
openclaw/openclaw
Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.
netdata/netdata
Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.
marin-community/marin
Deslop, simplify, or review low-value tests and prose only when explicitly requested for a branch or diff.
marin-community/marin
Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs.
marin-community/marin
Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.
marin-community/marin
Build, validate, publish, update, inspect, query, roll back, or archive a dynamic Marina applet.
marin-community/marin
Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.
marin-community/marin
Run a read-only preview for a specified Marin infra/pulumi stack and trace each pending resource change to merged pull requests since its latest successful update when that update records a clean…
Categories
Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial…. Deploy Hero Change is an agent skill from marin-community/marin. Deploy a significant code change (backend, kernel, optimizer, data path) to the live hero run: relaunch it under a new run id from a permanent checkpoint, compare against the old run over a trial window, and roll back if the gate fails; use only when the user asks to deploy a change to the hero.
Deploy Hero Change fits situations like: asks to deploy a change to the hero.
Run `npx skills add marin-community/marin --skill deploy-hero-change -a claude-code`. Or copy the skill folder (.agents/skills/deploy-hero-change in marin-community/marin) into .claude/skills/deploy-hero-change in your project. Claude Code loads it when a task matches its description.
Run `npx skills add marin-community/marin --skill deploy-hero-change -a codex`. Or copy the skill folder (.agents/skills/deploy-hero-change in marin-community/marin) into .agents/skills/deploy-hero-change in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add marin-community/marin --skill deploy-hero-change -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deploy-hero-change, .gemini/skills/deploy-hero-change, .github/skills/deploy-hero-change and .opencode/skills/deploy-hero-change in your project.
Going by SKILL.md and its folder, Deploy Hero Change needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Deploy Hero Change is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 411 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Deploy Hero Change: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 36k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
marin-community (a GitHub organization) maintains it in marin-community/marin, which has 3,920 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 9, 2026.
Source: marin-community/marin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.