Agent skill

Monitor Restore Iris

by open-thoughts in open-thoughts/OpenThoughts-Agent

Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost.

Apache-2.0Auto-check passedProductivity & Automation

Install Monitor Restore Iris

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore-iris -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent monitor-restore-iris --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/monitor-restore-iris .claude/skills/monitor-restore-iris && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
monitor-restore-iris
GitHub stars
301
Token cost
~2.5k tokens
SKILL.md length
488 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost.

  • Works in 3 steps: Check if it already exists — call… → If absent, call CronCreate with → Tell the user the new job id + the two…
  • Asks to restore/check the iris monitoring cron
  • SKILL.md covers When to run, Steps, Notes and Canonical cron prompt (copy…
  • Calls python

What it does

Monitor Restore Iris is an agent skill from open-thoughts/OpenThoughts-Agent. Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost. Primarily the marin TPU datagen/eval jobs ("iris" = the marin TPU cluster); also queries CoreWeave (cw-us-east-02a) GPU-RL as monitor-only. The cron is session-only and recurring crons auto-expire after 7 days, so it's routinely lost on a session restart. Use at the start of a new session, after a restart, or when the user asks to restore/check the iris monitoring cron. The sweep…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Scheduled and recurring tasks. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asks to restore/check the iris monitoring cron
  • Tasks that involve Scheduled and recurring tasks

Example prompts

  • “/monitor-restore-iris”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Check if it already exists — call CronList. If a recurring job whose prompt mentions "status check on ALL Iris jobs for user…
  2. If absent, call CronCreate with
  3. Tell the user the new job id + the two caveats: session-only (dies when this Claude session exits — re-run this skill next session) and…

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Monitor Restore Iris loads about 2.5k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 488 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 488 words, ~2,501 tokens.

Download SKILL.mdSave it as .claude/skills/monitor-restore-iris/SKILL.md (or your agent's skills folder).
name
monitor-restore-iris
description
Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost. Primarily the marin TPU datagen/eval jobs ("iris" = the marin TPU cluster); also queries CoreWeave (cw-us-east-02a) GPU-RL as monitor-only. The cron is session-only and recurring crons auto-expire after 7 days, so it's routinely lost on a session restart. Use at the start of a new session, after a restart, or when the user asks to restore/check the iris monitoring cron. The sweep PROCEDURE the cron runs lives in monitor-cron-sweep-iris.

monitor-restore-iris

📍 Iris orientation — read first. Read the Iris tools catalog (.agents/ops/iris/ops.md) and the Iris ops directory (.agents/ops/iris/ — ops.md for CoreWeave GPU, ops.md for TPU marin) for binding access/preamble/gotchas and the helper-script inventory.

The recurring cron watching all benjaminfeuer Iris jobs is session-only and recurring crons auto-expire after 7 days — routinely lost on a session restart. This skill is the durable source of truth for re-creating it: the canonical cron prompt below is what gets (re-)installed — copy it verbatim into CronCreate. The per-tick sweep methodology is monitor-cron-sweep-iris. (The separate broader tri-cluster monitor — Leonardo + CoreWeave + TACC — is monitor-restore / monitor-cron-sweep.)

When to run

  • Start of a new session where Iris jobs are in flight.
  • The user says the monitor/cron is gone, down, or "not firing."
  • After ~7 days (expiry).

Steps

  1. Check if it already exists — call CronList. If a recurring job whose prompt mentions "status check on ALL Iris jobs for user benjaminfeuer" is present, do nothing (a duplicate causes redundant SQL/tunnel load). If a stale datagen-only variant exists (prompt mentions only qwen3.5-122b-32k-%), CronDelete it and recreate with the all-jobs prompt below.
  2. If absent, call CronCreate with:
    • cron: 23 */3 * * * (every 3 h at :23 — off the :00/:30 marks)
    • recurring: true
    • prompt: the exact text in the fenced block below.
  3. Tell the user the new job id + the two caveats: session-only (dies when this Claude session exits — re-run this skill next session) and 7-day auto-expiry.
Show full SKILL.md (253 more words)Show less

Notes

  • durable: true is NOT honored in this harness (still creates a session-only job) — this skill IS the persistence layer.
  • The cron only fires while the REPL is idle (not mid-task). If it reliably misses, fallback is the user pasting the prompt manually or an external launchd monitor (out of scope).
  • It tracks ALL /benjaminfeuer/% jobs but the autonomous write actions (auto-rescue, keep-2-in-flight) are datagen-only; eval jobs are monitor-only (self-sync to Supabase+HF). See datagen-launch-iris (launch/refill), datagen-job-cleanup (canonical idempotent post-run cleanup for a TERMINAL datagen arm), and eval-agentic-launch-iris.
  • Two clusters. The cron queries both the marin TPU cluster and the cw-us-east-02a CoreWeave GPU cluster. The marin .venv iris carries the [controller] deps so it drives CoreWeave too — but the CoreWeave query MUST be prefixed KUBECONFIG=~/.kube/coreweave-iris-gpu, else iris falls back to the shell-default kubeconfig (~/.kube/lambdaconfig) and errors with Invalid kube-config file … Expected object with name. GPU-RL jobs on CoreWeave are monitor-only (no rescue, no keep-2); pods GC on terminal, so logs come from the persistent finelog server. Other CoreWeave GPU configs (coreweave* = US-WEST-04A, CI/smoke) are NOT in scope unless the user runs jobs there.
  • The methodology each step encodes (how to run the analyzer, classify, rescue, refill) is monitor-cron-sweep-iris — read it when actually executing a tick; this skill is just the (re)install wrapper + the canonical prompt.

Canonical cron prompt (copy verbatim into CronCreate)

Every-3-hours status check on ALL Iris jobs for user benjaminfeuer (datagen + eval + GPU-RL + anything else), across BOTH the marin TPU cluster and the CoreWeave GPU cluster.

**⚠ NO EXPERIMENT-SPECIFICS IN THIS PROMPT (they go stale): the per-campaign values — in-flight TARGET, refill cluster/grouping/order, harvest gates, repo/image patterns, and current bugs — live in the EXPERIMENT TRACKERS under `~/Documents/experiments/active/` (and the `*-launch` / `*-cleanup` / `analyze-*` skills). READ the relevant tracker each tick and drive off IT; never rely on a number hardcoded here.**

**⛔ DATAGEN IS OUT OF SCOPE: a DIFFERENT agent manages ALL datagen (`tracegen-iris-%` / `qwen3.5-122b-%`). Do NOT analyze, rescue, keep-N, or take any action on datagen jobs. This monitor covers EVAL (§3B) + Levanter TRAINING (§3C) + CoreWeave GPU-RL (§3D) only.**

1. Active jobs (query BOTH clusters):
   1a. marin (TPU):
       /Users/benjaminfeuer/Documents/marin/.venv/bin/iris --cluster=marin query "SELECT job_id, state FROM jobs WHERE state IN (1,2,3) AND job_id LIKE '/benjaminfeuer/%' ORDER BY job_id DESC LIMIT 20" -f csv
   1b. cw-us-east-02a (CoreWeave GPU) — KUBECONFIG prefix REQUIRED (else iris uses the wrong shell-default kubeconfig):
       KUBECONFIG=~/.kube/coreweave-iris-gpu /Users/benjaminfeuer/Documents/marin/.venv/bin/iris --cluster=cw-us-east-02a query "SELECT job_id, state FROM jobs WHERE state IN (1,2,3) AND job_id LIKE '/benjaminfeuer/%' ORDER BY job_id DESC LIMIT 20" -f csv
   For EACH cluster also query state IN (4,5,6) LIMIT 8 to catch jobs that went terminal since the last tick. If the cw query errors (cluster down / creds), report that and continue with marin.

2. For each ACTIVE marin (TPU) datagen/eval job, run the harbor analyzer (does NOT apply to CoreWeave GPU-RL — handle per class D). Use the analyze-job-history-iris skill:
   /Users/benjaminfeuer/miniconda3/envs/otagent/bin/python /Users/benjaminfeuer/Documents/OpenThoughts-Agent/scripts/iris/analyze_iris_harbor_job.py <job_id> --output /tmp/$(basename <job_id>)_history.md --resync
   Report from the .json sidecar: runtime_h, iris_preemption_count, cycles total/served, samples (serving_summary.gen_tps.n), gen tok/s mean/peak, Running mean/peak, non_empty/total trials = rate, t_first_serve, top harbor_exception_stats. ALSO report mean reward + completed/total tasks from the harbor progress line (NOT in the sidecar): /Users/benjaminfeuer/Documents/marin/.venv/bin/iris --cluster=marin job logs <job_id> --max-lines 8000 | grep -aoE '[0-9]+/[0-9]+ Mean: [-0-9.]+' | tail -1

3. Print `## Iris jobs status — <ISO UTC>`: one line per job (name + state + CLOSED/PARTIAL/OPEN/DEAD), a compact metrics block, and a survival check (past cold compile? throughput sane? traces/results landing on HF?). Classify each job by job_id prefix and apply the right treatment:

   A. **Datagen** (`qwen3.5-122b-%` / `tracegen-iris-%`): **⛔ OUT OF SCOPE — a DIFFERENT agent manages ALL datagen.** Do NOT query, analyze (§2), rescue (§4), keep-N (§5), or take ANY action on datagen jobs. If one appears in the state query, note its existence in ONE line at most and move on. §4 + §5 are BOTH retired for this monitor.

   B. **Eval** (`eval-%`): auto-sync to Supabase + HF on completion (`--upload_to_database`); build sandboxes at runtime (MAIN Daytona org). **ALWAYS report the leading metric (`<done>/<total> Mean: <X>`) per in-flight eval** (from `iris … job logs <job_id>`; not in the analyzer sidecar) + productive rate + exceptions; on terminal, whether results landed. A **one-off** eval is monitor-only (no rescue/relaunch). **⚠ EXCEPTION — an eval CAMPAIGN with a tracker in `active/` (e.g. `~/Documents/experiments/active/flawed_summ_evals/reeval_tracker.md`) DOES run an active harvest+refill loop: drive it PER THAT TRACKER each tick — its in-flight TARGET, refill cluster/grouping/order, harvest gate + discriminator, and gotchas ALL live in the tracker's TOP BLOCK (never hardcode them here). Route harvest via the `eval-agentic-cleanup` skill, refill via the `eval-*-launch` skill.**

   C. **Other** job types (e.g. Levanter training `iris-run-…` — health via `analyze-training-run-iris`; source of truth = its `active/` experiment dir): report state + a one-line health read; take no autonomous write action.

   D. **GPU-RL** (CoreWeave `cw-us-east-02a`, e.g. `rl-iris-%` / `rl-%` — MarinSkyRL GRPO on whole H100x8 nodes, possibly gang-scheduled multi-node `replicas>1`): **monitor-only — NO rescue, NO keep-2-in-flight, NO auto-relaunch.** The harbor analyzer in §2 does NOT apply (no harbor trial sidecars). For each in-flight GPU-RL job report state + the latest RL progress by reading the persistent finelog (pods GC on terminal): `KUBECONFIG=~/.kube/coreweave-iris-gpu /Users/benjaminfeuer/Documents/marin/.venv/bin/iris --cluster=cw-us-east-02a job logs <job_id> --max-lines 100000 --no-tail` then grep `WANDB_MIRROR kind=train step=` for the latest `trainer/global_step`, `loss/avg_raw_reward`, and `generate/num_failed_trajectories`/`generate/errors`. For multi-node confirm `All N Ray node(s) joined`. On a terminal job report exit state (4=SUCCEEDED). NEVER kill/relaunch GPU-RL jobs.

6. NEVER kill/restart/bounce a RUNNING job or the cluster without express user permission. GPU-RL and all other RUNNING jobs stay strictly no-touch (flag for the user, never kill). If a job is stuck PENDING (no capacity), report it and surface the unpinned-relaunch option — do not kill a running/placed job unprompted. (Datagen zombie-kill+rescue authority, when datagen WAS in scope: state 3 + harbor progress frozen ≥3h + task log ONLY `[fd-monitor]` heartbeats in that window with no recent healthy vLLM engine marker. But datagen is now §A out of scope — a different agent owns it.)

If you change the cadence or scope, update BOTH the cron/prompt above and the live job (delete + recreate), and keep monitor-cron-sweep-iris (the procedure) in sync — so this skill stays the canonical copy.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/monitor-restore-iris of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Monitor Restore Iris next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Monitor Restore Iris compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Monitor Restore Iris this skillopen-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.0
ScheduleTinyAGI/tinyagi3.6k—~1.4kAutomated safety check: PassMIT
Send User MessageTinyAGI/tinyagi3.6k—~829Automated safety check: PassMIT
Cron Opsczl9707/build-your-own-openclaw1.9k—~593Automated safety check: PassMIT
X Bookmarkssharbelxyz/x-bookmarks289—~2kAutomated safety check: NotesNone
Jobsphysiclaw/PhysiClaw3861 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Schedule

    TinyAGI/tinyagi

    Create, list, and delete scheduled tasks (recurring or one-time) that send messages to agents.

    3.6k GitHub stars~1.4k tokensUpdated 6 mo ago
    Productivity & AutomationAuto-check passed
  • Send User Message

    TinyAGI/tinyagi

    Send a proactive message to a paired user via their channel (Discord, Telegram, or WhatsApp).

    3.6k GitHub stars~829 tokensUpdated 6 mo ago
    Productivity & AutomationAuto-check passed
  • Cron Ops

    czl9707/build-your-own-openclaw

    Create, list, and delete scheduled cron jobs. An agent skill from czl9707/build-your-own-openclaw.

    1.9k GitHub stars~593 tokensUpdated 3 mo ago
    Productivity & AutomationAuto-check passed
  • X Bookmarks

    sharbelxyz/x-bookmarks

    Fetch, summarize, and manage X/Twitter bookmarks via bird CLI or X API v2.

    289 GitHub stars~2k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check: notes
  • Jobs

    physiclaw/PhysiClaw

    A skill your agent uses when the task involves scheduling future work — any "remind me at …", "every weekday …", "check again in 30 min", or closing a fired cron job.

    386 GitHub starsUsed in 1 repo~1.1k tokens
    Productivity & AutomationAuto-check passed
  • Wp Wpcli And Ops

    Automattic/agent-skills

    A skill your agent uses when working with WP-CLI (wp) for WordPress operations: safe search-replace, db export/import, plugin/theme/user/content management, cron, cache flushing, multisite, and…

    211 GitHub starsUsed in 2 repos~988 tokens
    Productivity & AutomationAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 9 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed

Questions about Monitor Restore Iris

What does Monitor Restore Iris do?

Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost. Monitor Restore Iris is an agent skill from open-thoughts/OpenThoughts-Agent. Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost.

When should I use Monitor Restore Iris?

Monitor Restore Iris fits situations like: asks to restore/check the iris monitoring cron; tasks that involve Scheduled and recurring tasks.

How do I install Monitor Restore Iris in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore-iris -a claude-code`. Or copy the skill folder (.agents/skills/monitor-restore-iris in open-thoughts/OpenThoughts-Agent) into .claude/skills/monitor-restore-iris in your project. Claude Code loads it when a task matches its description.

How do I install Monitor Restore Iris in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore-iris -a codex`. Or copy the skill folder (.agents/skills/monitor-restore-iris in open-thoughts/OpenThoughts-Agent) into .agents/skills/monitor-restore-iris in your project. Codex loads it when a task matches its description.

Can I use Monitor Restore Iris in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore-iris -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitor-restore-iris, .gemini/skills/monitor-restore-iris, .github/skills/monitor-restore-iris and .opencode/skills/monitor-restore-iris in your project.

What does Monitor Restore Iris need to run?

Going by SKILL.md and its folder, Monitor Restore Iris needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Monitor Restore Iris access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Monitor Restore Iris safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Monitor Restore Iris use?

Monitor Restore Iris is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Monitor Restore Iris use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Monitor Restore Iris?

Skills that share tags, products or a category with Monitor Restore Iris: Schedule (TinyAGI/tinyagi, 3.6k stars), Send User Message (TinyAGI/tinyagi, 3.6k stars), Cron Ops (czl9707/build-your-own-openclaw, 1.9k stars) and X Bookmarks (sharbelxyz/x-bookmarks, 289 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Monitor Restore Iris?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.