Ito Training
affaan-m/ECC
Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest.
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-training-run-iris --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-training-run-iris .claude/skills/analyze-training-run-iris && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyze-training-run-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-training-run-iris into .claude/skills/analyze-training-run-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-training-run-iris", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-training-run-irisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-training-run-iris --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/analyze-training-run-iris .agents/skills/analyze-training-run-iris && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyze-training-run-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-training-run-iris into .agents/skills/analyze-training-run-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-training-run-iris", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-training-run-iris --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/analyze-training-run-iris .cursor/skills/analyze-training-run-iris && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyze-training-run-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-training-run-iris into .cursor/skills/analyze-training-run-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-training-run-iris", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/analyze-training-run-iris--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-training-run-iris --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/analyze-training-run-iris .gemini/skills/analyze-training-run-iris && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyze-training-run-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-training-run-iris into .gemini/skills/analyze-training-run-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-training-run-iris", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-training-run-irisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/analyze-training-run-iris .github/skills/analyze-training-run-iris && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyze-training-run-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-training-run-iris into .github/skills/analyze-training-run-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-training-run-iris", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent analyze-training-run-iris --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/analyze-training-run-iris .opencode/skills/analyze-training-run-iris && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyze-training-run-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/analyze-training-run-iris into .opencode/skills/analyze-training-run-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyze-training-run-iris", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyze-training-run-irisDetailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
Analyze Training Run Iris is an agent skill from open-thoughts/OpenThoughts-Agent. Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g. the delphi midtraining runs) — step progress vs target, loss/throughput, preemption + MAJOR step-gap detection, and checkpoint cadence. Use for an executor coordinator (<run-coord) plus its nested <run-coord/checkpoints-<step-<hash training child, which the harbor analyzer (analyze-job-history-iris) does NOT cover (training has no harbor trial sidecars, same as GPU-RL). Reads W&B per-step history + iris job summary + GCS…
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gsutilpythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
WANDB_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyze Training Run Iris loads about 2k tokens when it runs. Until then it costs about 148 tokens; SKILL.md has 689 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 689 words, ~2,027 tokens.
.claude/skills/analyze-training-run-iris/SKILL.md (or your agent's skills folder).📍 Iris orientation — read first. Before acting on anything in this skill, read the Iris tools catalog (
.agents/ops/iris/ops.md) and the Iris ops directory (.agents/ops/iris/— the CoreWeave GPU particulars inops.md, the TPUmarinparticulars inops.md). They carry the binding access/preamble/gotchas and the helper-script inventory the steps below rely on.
A Levanter training run launched through the marin executor surfaces as TWO Iris jobs:
/<user>/<run>-coord — the executor_main DAG-walker (it submits the
training job and then blocks); and/<user>/<run>-coord/checkpoints-<step>-<hash> — the multi-task v5p
job where the actual training steps happen (e.g. 8 tasks for a v5p-64).Health = the CHILD's step progress + the run's preemption/gap history. The harbor analyzer
(analyze-job-history-iris) does NOT apply — a training job has no harbor trial sidecars (just like
GPU-RL). Use the three sources below instead. W&B is primary for step/loss/throughput; iris job summary
is primary for preemptions/liveness; GCS is primary for checkpoint cadence.
iris job summary: preemptions, liveness, per-task state (always available)IRIS=/Users/benjaminfeuer/Documents/marin/.venv/bin/iris
$IRIS --cluster=marin job summary <child_job_id>Report preemptions=N failures=N, tasks running/completed (all N tasks should be running together —
a v5p job is gang-scheduled), the longest task DURATION, and PEAK MEM. preemptions>0 is expected on a
preemptible v5p — each one means iris restarted the slice and Levanter resumed from the last checkpoint
(every preemption costs a wall-clock gap: re-place + reload weights + XLA recompile). failures>0, a
shrinking task count, or a crash-restart loop is a red flag → read the child logs for the error.
The run logs to W&B project delphi-midtraining (entity nyu-dice-lab); the run name is the
GCS output-path hash — the last path segment of gs://marin-us-east5/checkpoints/<run>-<hash> (e.g.
delphi-1e23-p33m67-k0p20-lr0.67-b6607e → run delphi-1e23-p33m67-k0p20-lr0.67-b6607e). Per-step history
is NOT mirrored by mum — query the W&B API directly (needs WANDB_API_KEY from $DC_AGENT_SECRET_ENV; use the
otagent python which has wandb):
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python - <<'PY'
import wandb
ENTITY, PROJECT, RUN = "nyu-dice-lab", "delphi-midtraining", "<run-hash>" # <-- the ...-b6607e hash
api = wandb.Api()
r = api.run(f"{ENTITY}/{PROJECT}/{RUN}")
total = r.config.get("trainer", {}).get("num_train_steps") or r.config.get("num_train_steps")
h = r.history(keys=["_step", "_timestamp", "_runtime", "train/loss"], pandas=True)
if h is None or len(h) == 0:
print("state:", r.state, "-> pre-first-step (still setup/HF-download/XLA-compile); no training step yet")
else:
h = h.dropna(subset=["_step"]).sort_values("_step")
cur = int(h["_step"].iloc[-1])
ts = h["_timestamp"].to_numpy()
deltas = [b - a for a, b in zip(ts[:-1], ts[1:])]
med = sorted(deltas)[len(deltas)//2] if deltas else 0.0
thr = max(300.0, 20.0 * med) # same MAJOR-GAP rule as compute_time.md
gaps = [d for d in deltas if d > thr]
toks = 0
bs = (r.config.get("trainer", {}) or {}).get("train_batch_size")
sl = r.config.get("train_seq_len") or (r.config.get("model", {}) or {}).get("max_seq_len")
if bs and sl and med: toks = bs * sl / med
loss = h["train/loss"].dropna()
loss = float(loss.iloc[-1]) if len(loss) else None
eta_h = ((total - cur) * med / 3600.0) if (total and med) else None
print(f"state : {r.state}")
print(f"step : {cur}/{total} ({round(100*cur/total,1) if total else '?'}%)")
print(f"train/loss : {loss}")
print(f"median step dt : {round(med,2)} s -> ~{round(toks):,} tok/s" if med else "median step dt: n/a")
print(f"MAJOR gaps : count={len(gaps)} total={round(sum(gaps)/3600,2)} h (threshold {round(thr)} s)")
print(f"ETA to {total} : ~{round(eta_h,1)} h of compute (excludes future preemption gaps)" if eta_h else "ETA: n/a")
PY_timestamp deltas exceeding max(300 s, 20 × median step interval) —
preemption / idle gaps (the metric the user wants: "note major gaps"). Cross-check count against
iris job summary preemptions — they should be the same order (a gap with NO matching preemption is a
silent stall worth flagging). Report gap count + total hours.batch × seq / median_dt. Compare across ticks; a sudden drop = contention or a bad slice.gsutil ls gs://marin-us-east5/checkpoints/<run>-<hash>/ | grep -E 'step-[0-9]+' | tail -5
gsutil ls -l gs://marin-us-east5/checkpoints/<run>-<hash>/step-<latest>/ 2>/dev/null | tail -2 # timestampReport the latest persisted step + its timestamp (the checkpointer saves on an interval — e.g.
save_interval 10m, keep every 1500). The latest checkpoint lagging a bit behind the W&B step is fine
(async save). But no step-* checkpoint long after training started is a red flag — under preemption
the run would lose all un-checkpointed progress. Only .executor_info / .executor_status* present (no
step-*) = still pre-first-checkpoint (early bring-up).
iris job logs <child> early on shows [iris setup] step N/M lines — those are uv-sync SETUP steps,
NOT training steps. Do NOT grep step N from the logs for progress. Use the W&B _step (Source 2)
as the authoritative training-step counter; only fall back to Levanter's own in-log training-step line if
W&B is unreachable.
<run> state=running step=<cur>/<total> (X%) loss=<L> ~<T>tok/s preempts=<P> gaps=<G>/<H>h ckpt=step-<C>
plus a one-line health read: past setup/compile? step rate sane vs the prior tick? loss finite and trending
down (not NaN/spiking)? preemptions resuming cleanly (checkpoint advancing)? ETA to the K-budget target.
The W&B pull is fast (seconds), unlike the harbor analyzer — you usually do NOT need a subagent. If a run
is huge or you are sweeping several, the analyze-job-history-iris foreground-and-wait discipline still
applies to any slow gsutil/log reads, but the W&B query itself is quick.
analyze-job-history-iris.© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/analyze-training-run-iris of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Analyze Training Run Iris next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyze Training Run Iris this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Ito Trainingaffaan-m/ECC | 276k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Skin Health Analyzersickn33/agentic-awesome-skills | 47k | 2 repos | ~293 | Automated safety check: Pass | MIT | |
| Ray Train Distributed TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.7k | Automated safety check: Pass | MIT | |
| Codebase Health Dashboardgarrytan/gstack | 136k | — | ~11k | Automated safety check: Notes | MIT | |
| Fal Trainnexu-io/open-design | 100k | — | ~293 | Automated safety check: Pass | Apache-2.0 |
affaan-m/ECC
Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest.
sickn33/agentic-awesome-skills
Analyze skin health data, identify skin problem patterns, assess skin health status.
Orchestra-Research/AI-Research-SKILLs
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
garrytan/gstack
Runs a project's own type checker, linter, test runner, dead-code detector and shell linter, combines them into a weighted 0-10 score and tracks the trend.
nexu-io/open-design
Train custom AI models (LoRA) on fal.ai for personalized image generation tailored to a brand, character, or style.
agenticnotetaking/arscontexta
Run condition-based vault health diagnostics. An agent skill from agenticnotetaking/arscontexta.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
open-thoughts/OpenThoughts-Agent
Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g. Analyze Training Run Iris is an agent skill from open-thoughts/OpenThoughts-Agent.g.
Analyze Training Run Iris fits situations like: an executor coordinator (<run-coord) plus its nested <run-coord/checkpoints-<step-<hash training child; which the harbor analyzer (analyze-job-history-iris) does NOT cover (training has no harbor trial sidecars; same as GPU-RL).
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a claude-code`. Or copy the skill folder (.agents/skills/analyze-training-run-iris in open-thoughts/OpenThoughts-Agent) into .claude/skills/analyze-training-run-iris in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a codex`. Or copy the skill folder (.agents/skills/analyze-training-run-iris in open-thoughts/OpenThoughts-Agent) into .agents/skills/analyze-training-run-iris in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill analyze-training-run-iris -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-training-run-iris, .gemini/skills/analyze-training-run-iris, .github/skills/analyze-training-run-iris and .opencode/skills/analyze-training-run-iris in your project.
Going by SKILL.md and its folder, Analyze Training Run Iris needs the command-line tools its instructions call (gsutil and python) and credentials named WANDB_API_KEY. Our summary lists: Python 3; A credential in WANDB_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyze Training Run Iris is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Analyze Training Run Iris: Ito Training (affaan-m/ECC, 276k stars), Skin Health Analyzer (sickn33/agentic-awesome-skills, 47k stars), Ray Train Distributed Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Codebase Health Dashboard (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.