Agent skill

Babysit Zephyr

by marin-community in marin-community/marin

Launch or continuously monitor a specified Zephyr pipeline on Iris only when asked to babysit it; restart only after explicit approval.

Apache-2.0Auto-check passed

Install Babysit Zephyr

skills CLI
$ npx skills add marin-community/marin --skill babysit-zephyr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install marin-community/marin babysit-zephyr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/babysit-zephyr .claude/skills/babysit-zephyr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
babysit-zephyr
GitHub stars
3.9k
Token cost
~1.1k tokens
SKILL.md length
493 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Launch or continuously monitor a specified Zephyr pipeline on Iris only when asked to babysit it; restart only after explicit approval.

  • Works in 3 steps: Smoke check (first 2-5 minutes): Confirm… → Steady-state monitoring: Check stage… → Failure detection: If workers get KILLED…
  • SKILL.md covers Zephyr Job Structure, Iris Config, Starting a Job and Stopping a Job, plus 4 more sections
  • Calls uv

What it does

Babysit Zephyr is an agent skill from marin-community/marin. Launch or continuously monitor a specified Zephyr pipeline on Iris only when asked to babysit it; restart only after explicit approval.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Open-source framework for the research and development of foundation models. The licence is Apache-2.0.

Example prompts

  • “/babysit-zephyr”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Smoke check (first 2-5 minutes): Confirm coordinator and workers child jobs appear and reach RUNNING. Check coordinator logs for early…
  2. Steady-state monitoring: Check stage progress via coordinator logs. Confirm (a) shards complete within the current stage, and (b) stages…
  3. Failure detection: If workers get KILLED or the coordinator goes zombie, the StepRunner may retry automatically (new child jobs with a…

What it can do on your machine

Read from SKILL.md and the folder at commit 2ea1c1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Babysit Zephyr loads about 1.1k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 493 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from marin-community/marin at commit 2ea1c1d, republished under its Apache-2.0 licence (© marin-community). 493 words, ~1,093 tokens.

Download SKILL.mdSave it as .claude/skills/babysit-zephyr/SKILL.md (or your agent's skills folder).
name
babysit-zephyr
description
Launch or continuously monitor a specified Zephyr pipeline on Iris only when asked to babysit it; restart only after explicit approval.

Babysit a Zephyr job

Zephyr Job Structure

A job has a one-task *-coord coordinator and a *-workers pool. Sequential pipelines produce different p<N> child names; retries produce different hashes with the same p0. Child names follow <hash>-p<pipeline>-a<attempt>-{coord,workers}. Old coordinators may linger.

Iris Config

Resolve the requested cluster to its file under lib/iris/config/; substitute that file for <CONFIG> below. See lib/zephyr/OPS.md for the dashboard and coordinator query reference.

Starting a Job

Get the run command from the user:

bash
uv run iris --config <CONFIG> job run --region <REGION> --no-wait -- python <SCRIPT>

The entrypoint container defaults to 1GB memory. For long-running pipelines that accumulate state (GCS clients, logging), increase with --memory:

bash
uv run iris --config <CONFIG> job run --region <REGION> --memory 5GB --no-wait -- python <SCRIPT>

The command prints a job ID on success. Note it for monitoring.

Stopping a Job

Always ask the user before stopping. Stopping kills all child jobs (coordinators, workers).

bash
uv run iris --config <CONFIG> job cancel <JOB_ID>

Monitoring

Health Checks

Check child job states via the Iris CLI (returns per-task state and resourceUsage):

bash
# diskMb is updated every ~60s. On K8s it is always 0 (workdir lives inside the pod).
uv run iris --config <CONFIG> rpc controller list-tasks --job-id <JOB_ID>

A healthy zephyr job has:

  • Coordinator: RUNNING, 1 task running
  • Workers: RUNNING, tasks ramping up toward target count
Stage Progress

The coordinator logs stage, completed, in-flight, queued, and worker counts:

bash
uv run iris --config <CONFIG> rpc controller get-task-logs \
  --id <COORD_JOB_ID> --max-total-lines 5000 --attempt-id -1 --tail

Large pools can flood the log with pull_task, Started operation, report_result, registered, and tasks completed; filter those entries.

Coordinator Thread Dump

When logs are flooded, a thread dump tells you if the coordinator is alive and working:

bash
uv run iris --config <CONFIG> rpc controller profile-task \
  --json '{"target":"<COORD_JOB_ID>/0","durationSeconds":1,"profileType":{"threads":{}}}'

Key patterns:

  • actor-method_0 in _wait_for_stage → pipeline active, waiting for current stage to complete
  • _coordinator_loop thread present → heartbeat/dispatch loop running
  • All threads in _worker (thread pool idle) → pipeline exited, coordinator is a zombie
Show full SKILL.md (254 more words)Show less

Monitoring Lifecycle

After submitting, monitor in escalating stages:

  1. Smoke check (first 2-5 minutes): Confirm coordinator and workers child jobs appear and reach RUNNING. Check coordinator logs for early errors. Failure here is likely a code bug, config issue, or bundle fetch timeout.

  2. Steady-state monitoring: Check stage progress via coordinator logs. Confirm (a) shards complete within the current stage, and (b) stages advance. Calibrate check-in interval so you see at least one stage transition between checks — every few minutes for many short stages, every 15-30 minutes for few long stages.

  3. Failure detection: If workers get KILLED or the coordinator goes zombie, the StepRunner may retry automatically (new child jobs with a different hash). Check the latest attempt. Stale coordinators from previous attempts may accumulate (#3705). If retries keep failing, escalate to debug.

"Terminated by user" is misleading: This does not necessarily mean a human killed the job. The system uses this message for various internal termination reasons. Always check the actual logs at each level (parent job, coordinator, workers) to find the real cause.

Restarting After Failure

  1. Ask the user if it's okay to stop and restart.
  2. Stop the job.
  3. Get the run command (or reuse the previous one).
  4. Submit and resume monitoring.

When to Escalate

Escalate to debug when:

  • A stage is stuck (no shard progress for an extended period)
  • Stragglers are holding up a stage (few in-flight, 0 queued, most workers idle)
  • Workers are failing repeatedly with the same error
  • Controller issues (e.g., RPCs timing out)

© marin-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/babysit-zephyr of marin-community/marin.

Open the folder on GitHubat commit 2ea1c1d

Compare with similar skills

Babysit Zephyr next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Babysit Zephyr compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Babysit Zephyr this skillmarin-community/marin3.9k—~1.1kAutomated safety check: PassApache-2.0
Launch Monitoraaron-he-zhu/aaron-marketing-skills2.9k—~3.3kAutomated safety check: PassApache-2.0
Shipping and Launch Checklistaddyosmani/agent-skills105k1 repos~2.8kAutomated safety check: PassMIT
Continuetelegramdesktop/tdesktop33k2 repos~9.4kAutomated safety check: PassGPL-3.0
Babysitsimstudioai/sim30k—~2.9kAutomated safety check: PassApache-2.0
Qdrant Monitoringgithub/awesome-copilot40k1 repos~276Automated safety check: PassMIT

Similar skills

  • Launch Monitor

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "monitor my launch", "track our Product Hunt / Hacker News ranking", or "watch the launch window"; runs the T-0 to T+30 window watch — pre-launch…

    2.9k GitHub stars~3.3k tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Shipping and Launch Checklist

    addyosmani/agent-skills

    Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.

    105k GitHub starsUsed in 1 repo~2.8k tokens
    DevOps & CloudAuto-check passed
  • Continue

    telegramdesktop/tdesktop

    Continue autonomous Telegram Desktop development from the shared ai-tdesktop repository.

    33k GitHub starsUsed in 2 repos~9.4k tokens
    Productivity & AutomationAuto-check passed
  • Babysit

    simstudioai/sim

    Drive a PR to a clean review (Greptile 5/5, zero open threads) — ships if needed, keeps it mergeable against staging, re-triggers both Greptile and cubic, fixes real findings, replies to and…

    30k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Qdrant Monitoring

    github/awesome-copilot

    Official

    Guides Qdrant monitoring and observability setup. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~276 tokens
    DevOps & CloudAuto-check passed
  • [DEPRECATED - use continuous-learning-v2] Legacy v1 stop-hook skill extractor.

    276k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed

More from marin-community/marin

All 41 skills in this repo
  • Noslop

    marin-community/marin

    Deslop, simplify, or review low-value tests and prose only when explicitly requested for a branch or diff.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Use Iris

    marin-community/marin

    Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs.

    3.9k GitHub stars~745 tokensUpdated today
    Auto-check passed
  • Launch Rl

    marin-community/marin

    Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.

    3.9k GitHub stars~894 tokensUpdated today
    Auto-check passed
  • Marina Applet

    marin-community/marin

    Build, validate, publish, update, inspect, query, roll back, or archive a dynamic Marina applet.

    3.9k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Query Finelog

    marin-community/marin

    Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Trace Pulumi Diff

    marin-community/marin

    Run a read-only preview for a specified Marin infra/pulumi stack and trace each pending resource change to merged pull requests since its latest successful update when that update records a clean…

    3.9k GitHub stars~663 tokensUpdated today
    Auto-check passed

Questions about Babysit Zephyr

What does Babysit Zephyr do?

Launch or continuously monitor a specified Zephyr pipeline on Iris only when asked to babysit it; restart only after explicit approval. Babysit Zephyr is an agent skill from marin-community/marin. Launch or continuously monitor a specified Zephyr pipeline on Iris only when asked to babysit it; restart only after explicit approval.

How do I install Babysit Zephyr in Claude Code?

Run `npx skills add marin-community/marin --skill babysit-zephyr -a claude-code`. Or copy the skill folder (.agents/skills/babysit-zephyr in marin-community/marin) into .claude/skills/babysit-zephyr in your project. Claude Code loads it when a task matches its description.

How do I install Babysit Zephyr in Codex?

Run `npx skills add marin-community/marin --skill babysit-zephyr -a codex`. Or copy the skill folder (.agents/skills/babysit-zephyr in marin-community/marin) into .agents/skills/babysit-zephyr in your project. Codex loads it when a task matches its description.

Can I use Babysit Zephyr in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add marin-community/marin --skill babysit-zephyr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/babysit-zephyr, .gemini/skills/babysit-zephyr, .github/skills/babysit-zephyr and .opencode/skills/babysit-zephyr in your project.

What does Babysit Zephyr need to run?

Going by SKILL.md and its folder, Babysit Zephyr needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Babysit Zephyr access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Babysit Zephyr safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Babysit Zephyr use?

Babysit Zephyr is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Babysit Zephyr use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Babysit Zephyr?

Skills that share tags, products or a category with Babysit Zephyr: Launch Monitor (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Shipping and Launch Checklist (addyosmani/agent-skills, 105k stars), Continue (telegramdesktop/tdesktop, 33k stars) and Babysit (simstudioai/sim, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Babysit Zephyr?

marin-community (a GitHub organization) maintains it in marin-community/marin, which has 3,925 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 11, 2026.

Source: marin-community/marin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.