Standards
ItamarZand88/CLI-Anything-WEB
Runs Phase 4 review/publish/verify for a cli-web- CLI: implementation review by 3 parallel agents, the tiered quality checklist (Tier 1 critical fail-fast, then comprehensive), pip install + smoke…
Deep single-RL-job health probe → a KILL / NO-KILL / ERROR recommendation for the supervisor.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent rl-job-health-deep-dive --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/rl-job-health-deep-dive .claude/skills/rl-job-health-deep-dive && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rl-job-health-deep-dive" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/rl-job-health-deep-dive into .claude/skills/rl-job-health-deep-dive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rl-job-health-deep-dive", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/rl-job-health-deep-diveType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent rl-job-health-deep-dive --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/rl-job-health-deep-dive .agents/skills/rl-job-health-deep-dive && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rl-job-health-deep-dive" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/rl-job-health-deep-dive into .agents/skills/rl-job-health-deep-dive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rl-job-health-deep-dive", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent rl-job-health-deep-dive --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/rl-job-health-deep-dive .cursor/skills/rl-job-health-deep-dive && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rl-job-health-deep-dive" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/rl-job-health-deep-dive into .cursor/skills/rl-job-health-deep-dive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rl-job-health-deep-dive", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/rl-job-health-deep-dive--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent rl-job-health-deep-dive --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/rl-job-health-deep-dive .gemini/skills/rl-job-health-deep-dive && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rl-job-health-deep-dive" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/rl-job-health-deep-dive into .gemini/skills/rl-job-health-deep-dive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rl-job-health-deep-dive", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent rl-job-health-deep-diveInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/rl-job-health-deep-dive .github/skills/rl-job-health-deep-dive && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rl-job-health-deep-dive" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/rl-job-health-deep-dive into .github/skills/rl-job-health-deep-dive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rl-job-health-deep-dive", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent rl-job-health-deep-dive --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/rl-job-health-deep-dive .opencode/skills/rl-job-health-deep-dive && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rl-job-health-deep-dive" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/rl-job-health-deep-dive into .opencode/skills/rl-job-health-deep-dive/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rl-job-health-deep-dive", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rl-job-health-deep-diveDeep single-RL-job health probe → a KILL / NO-KILL / ERROR recommendation for the supervisor.
Rl Job Health Deep Dive is an agent skill from open-thoughts/OpenThoughts-Agent. Deep single-RL-job health probe → a KILL / NO-KILL / ERROR recommendation for the supervisor. Dispatched as a subagent on every monitor tick for RL jobs in NEW/UNTESTED settings (new config/geometry/model, "debug" or "smoke-test" flavor, first launches after a code/config change), and whenever a running job looks starved or wedged. Goes BEYOND state-poll + table metrics: captures the job's logs + tracejobs + GPU view, then runs four gates — (A) liveness, (B) resource utilization / engine subscription, (C) rollout…
Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Subagents and QA and bug reports. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlpythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Rl Job Health Deep Dive loads about 5.8k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 2,514 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 2,514 words, ~5,762 tokens.
.claude/skills/rl-job-health-deep-dive/SKILL.md (or your agent's skills folder).⚠ Do not add comments to YAMLs. Report your recommendations directly to the supervisor.
Probe one RL job when it is new or untested (config, geometry, model, image, debug/smoke launch, or first launch after a change), or looks starved or wedged. This distinguishes genuine progress from a silent death.
You are a SUBAGENT — you do NOT execute the kill (standing guardrail: never kill a RUNNING job without explicit
permission). When genuinely uncertain, prefer NO-KILL + escalate — a wrongly-killed healthy run wastes a whole
bring-up; a wrongly-kept dead one wastes one sweep. But if you couldn't get the evidence, the answer is ERROR,
not a hedged NO-KILL (see §0).
Before any gate, sync logs locally with this utility and analyze those files. Do not make repeated live
iris job logs or kubectl exec greps the primary method.
set -a; source /Users/benjaminfeuer/Documents/secrets.env; set +a
PY=/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python
$PY /Users/benjaminfeuer/Documents/MarinSkyRL/infra/sync_rl_logs.py \
/benjaminfeuer/<job> --cluster <cw-us-east-02a|cw-rno2a> [--run run-<ts>] [--dest DIR] [--trace-jobs]
# → ./<slug>-<run>/finelog.log + ./<slug>-<run>/ray_session_logs/
# (+ ./<slug>-<run>/<slug>_trace_jobs.tar.gz when --trace-jobs is passed — the Harbor rollout
# artifacts streamed into ONE archive; opt-in, since trace_jobs can be very large)It pulls both required sources:
finelog.log = the aggregated controller/job stream — this is where the terminating exception / NCCL timeout /
store->get(...) got error / Worker rank N received signal lives. The per-actor ray logs are frequently
traceback-less (Ray's C++ core logs); a hang's root cause routinely appears ONLY in the finelog. Read this FIRST.ray_session_logs/ = the per-actor worker-*.out/.err + python-*/raylet/gcs — for per-rank NCCL Init COMPLETE/nranks, weight-sync group construction, py-spy correlation, and the failing rank's own stderr.Analyze files with grep or Python. For NCCL/weight-sync structure, run
experiments/active/tasktrove-dq-sweep-opencode/artifacts/parse_collectives.py on ray_session_logs/. The utility
is idempotent and selects kubeconfig by cluster.
Hand the supervisor the local file paths and quoted lines. Live probes are reserved for a py-spy on a still-running wedge; all other analysis is file-based.
Return VERDICT: ERROR whenever you could not obtain the evidence: a required tool failed (auth / PATH /
resolver bug / timeout), a log can't be fetched or parsed, you can't separate policy-mesh from engine GPUs, or two
authoritative signals disagree and you can't reconcile them. Stop — give (a) the exact command you ran, (b) its exact
failure output, (c) what evidence is therefore missing, (d) what you did establish. Do NOT emit KILL/NO-KILL,
do NOT default to NO-KILL, NEVER substitute a plausible guess for a missing measurement.
No gate verdict without its named, quoted artifact:
| Gate | A PASS/FAIL requires you to have READ + QUOTED | else that gate is |
|---|---|---|
| A liveness | the authoritative state-poll line and the newest phase-Timer / step line + its timestamp | ERROR |
| B resources | per-rank GPU util with policy ranks separated from engine ranks, and the engine subscription line (running vs waiting vs the serving cap) | ERROR |
| C rollouts | actual reward values / trial exception files you OPENED (not a count you assumed) | ERROR |
When engines are under-subscribed or idle (Running low, Waiting=0, KV≈0, ⅓-TDP power) and the generation buffer
stalls, measure the gen→dispatch→train pipeline live:
num_coordinators (K) coordinator processes CPU/GIL-pegged on
submit_batch/gather/post-gather shaping? (py-spy / top them.) → dispatcher-bound → the lever is K, not n/npgw.max_staleness_steps because policy_train is the slow
phase (training-bound)?generate() calls reaching the engine but Waiting stays 0 (engines drain
instantly)? → the bottleneck is upstream dispatch rate, not engine capacity.(The saturation-READ tuple + "SM-util% is a trap" in marinskyrl's "Saturating vLLM engines" section is still valid; its CAUSAL "raise concurrency to beat the duty cycle" story is REFUTED — see the caveat there.) Diagnose while the job is ALIVE (py-spy dies with it). A verdict on an unmeasured starvation is ERROR-quality.
Read the relevant pointers first; this skill does not restate their contents.
.agents/ops/<cluster>/:ops/iris/ops.md — §Access (kubeconfig per cluster: East vs cw-rno2a), §Observability (the state-poll primitive iris_ops.py, JobState codes, finelog fetch, and the Poll/tooling pitfalls: rno2a job summary flakiness, the *_ms query columns, the analyze_iris_harbor_job --config/RL-output-dir limits, no server-side job logs grep), §Scheduling (gang/Kueue admission + node-headroom math), §Daytona (orgs, sandbox lifecycle, concurrency headroom), §Monitoring & debugging practices (incl. py-spy). Node shape → ops/iris/ops.md.ops/leonardo/ops.md; TACC → ops/tacc/ops.md.iris-log-resource-discipline): state-poll for liveness, bounded/filtered fetch for metrics — never dump a long log into the Mac. Exact bounded-fetch commands in ops §Observability..agents/projects/<dep>/:projects/marinskyrl/marinskyrl.md — the phase-Timer/step vocabulary, [MoE-PATH] grouped_mm-vs-for-loop, the colocated-engine/rank-0-logging deception, engine saturation (n_concurrent_trials scaling), the 80B GDN-GIL/HeartbeatMonitor death + FlashQLA, SKYRL_W13_RELOAD_BRACKET token-salad, SKYRL_R3_RESIDENT, the benign prob_diff_mean artifact, config-schema (Hydra struct) rules, runtime knobs. Read the section matching what you're chasing — do not guess a log line's meaning.projects/vllm/vllm.md — the serve engine: MoE/DCP/R3 flags, enforce_eager (CUDA-graphs) throughput cliff, serving-throughput expectations, the benign engine heartbeats (shm_broadcast … 60/600s).projects/harbor/vllm/daytona/ — the rollout/trial layout, verifier/reward path, passthrough exceptions, sandbox failure modes.scripts/iris/analyze_coreweave_rl_job_live.sh (CoreWeave artifact pull),
scripts/iris/iris_ops.py (state-poll), scripts/iris/analyze_iris_harbor_job.py (finelog science — mind the RL-job/rno2a limits in the ops note). The per-rung bring-up ladder is rl-agentic-launch-iris §8.Inputs (from the dispatch; most derivable): cluster, job id, pod-name substring, model + size (dense vs MoE active-B), the run's stage, and what's "new/untested" about it (scrutinize that hardest).
Check restart burn first: record B burned / K max, prior terminal errors, and whether failures repeat.
Repeated identical crashes are deterministically doomed; a recovered transient is benign. See ops/<cluster>/ for mechanics.
Capture the artifacts = STEP 0's sync_rl_logs.py (finelog + ray_session_logs) — already done before you reach
this gate. Do NOT hand-roll a kubectl/R2 sync or a live iris job logs grep; work from the local files STEP 0
pulled. (scripts/iris/analyze_coreweave_rl_job_live.sh is still the tool for the per-trial opencode.txt/rollout view when you
need rollout-quality detail beyond the logs.) What the phase-Timers / step counter / [MoE-PATH] markers mean →
projects/marinskyrl. The phase Timers are the progress truth — not a trace count, not a progress bar. (0
trials at +15 min on a long-episode arm is normal, not "done" and not "dead.")
⚡ ON ANY DEATH/WEDGE — READ BOTH finelog.log AND ray_session_logs/ (STEP 0 pulled both). The root cause hides in ONE of them; which one varies. Two failure classes, two homes:
store->get('...') got error: wait timeout after 1800000ms + the collective_rpc/broadcastUniqueNCCLID traceback of a het weight-sync bootstrap hang, or a Worker rank N received signal. The per-actor ray logs frequently do NOT carry this (worker-*.err may hold only a benign metrics-RPC line). Read finelog.log FIRST (grep it for got error/timeout/Traceback/received signal/Init COMPLETE).EngineCore fatal, a Ray ObjectLostError/OwnerDiedError, an actor crash, per-rank NCCL Init COMPLETE/nranks, weight-sync group construction. Grep worker-*.out/worker-*.err (NOT python-core-*.log, Ray's traceback-less C++ core logs); for NCCL/weight-sync structure run parse_collectives.py on the dir (STEP 0). A colocated-engine job that dies exit=0 with NO finelog traceback (only Killed ray::IDLE/raylet/gcs) has its killer here.exit=0 teardown with node RAM low is NOT a verdict — it's a missing-evidence ERROR until you've read BOTH local files. Worked examples: keep1-v22 (silent ~34-min death = vLLM EngineDeadError reshape_and_cache_flash … Meta tensors at first-step weight-sync, found in ray_session_logs); keep1-ncclnet (het weight-sync bootstrap hang = store->get('/skyrl/skyrl//cuda//0') 1800000ms timeout via collective_rpc, found ONLY in the finelog). Deeper background → ops/iris/ops.md §object-store / RAY ACTOR LOGS.Liveness = authoritative state poll plus log freshness, never one log-string grep. Use ops/iris/…
§Observability, then compare captured-log freshness with expected cadence and read wedge/death signatures.
⚠ Multi-mesh RL hides a wedged policy behind live engines (the colocated-engine deception). On any FSDP×EP×CP RL job the vLLM engines + RolloutCoordinator are SEPARATE actors from the policy mesh — a hung policy collective can read
running / fresh-heartbeat / high-utilwhile doing nothing. The three signals that don't lie (Ray actor-death logs, the NCCL watchdog, is the TRAINER/DRIVER log ADVANCING with policy-rank GPUs separated from engine GPUs) and the exact strings to grep for each —projects/marinskyrl(colocated-engine deception + rank-0-logging trap) andprojects/vllm(benign engine noise vs realEngineDeadError). Do not call Gate A PASS on a multi-mesh job without clearing all three.
Gate A verdict: DEAD/TERMINAL (state-poll failed/absent, 0 pods) → nothing to kill; report the root-cause
traceback + transient-vs-deterministic. WEDGED (running + a real hang signature + stale logs, no benign
explanation) → lean KILL (but if it's a starvation wedge, capture the live py-spy first — §0). A py-spy
barrier-snapshot + a lone NCCL Watchdog … ran for N ms LOG LINE is NOT a wedge by itself — a real tripped
watchdog ABORTS the process, so require pod-restarts==0 + an actual abort/terminal state + stalled FRESH logs (all
nodes) and reconcile the cited timeout against the run's timeline before calling wedge (the CW py-spy command + this
caveat: ops/iris/… §Monitoring & debugging practices). ALIVE + fresh → Gate B.
Live-poll GPUs; separate policy ranks from engine ranks and never average them. See ops/<cluster>/ for commands.
Running requests, Waiting ≈ 0, GPUs resident-but-~0%-util, the generation buffer barely filling (or a
frozen Generation Buffer Progress: N/M heartbeat — same N, growing elapsed). That is NOT a Daytona fault (§0);
the first hypothesis is rollout concurrency too low to saturate the engines, and the Daytona side has large
concurrency headroom to scale into. *The saturation math (n_concurrent_trials = 2·num_parallel_generation_workersprojects/marinskyrl (Saturating vLLM
engines). Throughput-vs-hardware expectations + the enforce_eager CUDA-graph cliff to rule out first →
projects/vllm + the H100 node shape in ops/iris/ops.md. The opposite failure —
Waiting ≫ Running with flat throughput — is over-subscription thrash. Report the concrete counts, not "looks fine."OOMKilled), policy/ref ranks actually compute-bound (high util + power) during a step — all-0%
with no log advance during "training" = wedge. OOM-detection mechanics → ops/<cluster>/; host-RAM breakdown
vocabulary + the 80B optimizer-spike danger window → projects/marinskyrl.Gate B verdict: engines under-subscribed/idle with a stalled buffer, throughput floored with Running>0 and
enforce_eager:false, or a training-stage OOM / all-ranks-0%-no-progress → lean KILL or a config fix (for
under-subscription, the fix is usually a concurrency bump, not a kill). All engines fed + generating, or training
steps advancing without OOM → healthy; Gate C.
State + GPUs can be green while the run produces garbage (e.g. a degraded weight-sync serving token-salad → all
reward-0 → no learning signal). Read the literal rollouts, qualitatively: trials INITIALIZING? COMPLETING
(count the reward markers — 0 completed at +15 min on a long arm is expected)? any rewards NON-ZERO? TURNS
completing (avg≈1 turn = dead-engine/broken-loop)? agent outputs sane (read 3–5)? verifiers scoring real attempts
vs erroring? The trace/reward/verifier layout + the known failure fingerprints (incoherent output ⇒
weight-sync/geometry fault — check SKYRL_W13_RELOAD_BRACKET; genuine infra exceptions — name them from a trial
exception file you OPENED, don't assume; ⚠ engine STARVATION is a §3 dispatch problem, NOT "every trial threw a
Daytona exception") → projects/harbor + projects/marinskyrl + ops/iris/… §Daytona.
⚠ 100%
AddTestsDirErroron a known-good dataset = a CONTAINER problem, not the dataset. When every rollout batch failsAddTestsDirError("Failed to add tests directory to environment") on a fresh/bespoke image but the dataset has run cleanly across prior experiments, the Daytona sandbox is never built (self._sandbox is None→ "Sandbox not found. Please build the environment first.") and the verifier'supload_dirinto the missing sandbox is what surfaces asAddTestsDirError. Root cause is usually container TRANSITIVE-dep drift, NOT Harbor and NOT the data:harbor[daytona]installed without--no-depslets the Daytona SDK + its transitives (e.g.websockets, litellm) re-resolve against a changed base env (a transformers/megatron bump), breaking sandbox-create vs the last-good image. Diff the failing image's Daytona-path deps against the last-good working image, pin them back, and validate sandbox-CREATE cheaply (a 1-pod throwaway that actually instantiates the sandbox) — an import smoke is NOT enough (the break is at create, not import) — before any GPU relaunch. Prove any single-variable hypothesis (e.g. "wrong Harbor") with hard evidence before rebuilding on it.
Gate C verdict: incoherent/all-reward-0 from a serving/sync/verifier fault on a new geometry, or a path that yields zero learning signal with no transient explanation → lean KILL (+ the fix). Trials completing with some non-zero rewards, or coherent multi-turn attempts on genuinely-hard tasks even at low pass-rate → NO-KILL, learning.
Run when an agentic RL job shows an engine sawtooth (inference Running peaks then troughs) and you must
decide whether each trial's throughput is capped by LLM generation vs sandbox-lifecycle churn vs
tool-exec vs error/retry — i.e. to put NUMBERS behind (or refute) a "sandbox churn" claim. Never assert
"sandbox churn" from the sawtooth alone; no numbers → ERROR per §0.
Source + discipline: the clean per-trial breakdown is each trial's result.json TimingInfo (NOT finelog).
The reusable recipe — field-to-phase map, duty-cycle fraction math, lease-race / burst≠churn checks (a harbor
trial artifact) — is in projects/harbor/ops.md §"Per-trial TimingInfo duty-cycle recipe". On cw-rno2a
the trials bucket is in-cluster-only, so aggregate in-pod and transfer aggregates only (the cluster-specific
kubectl exec + boto3 access + trials path is in ops/iris/ops.md §Observability "Per-trial
TimingInfo duty-cycle read"). Read only a bounded sample (newest ~200 for the duty cycle, ~500 for error/re-provision tails).
Compute + read (per trial, then median + p10/p90/max):
environment_setup + teardown-gap) / total = the sandbox-churn tax.n_concurrent_trials / generation-buffer depth (feed more trials to the
engines), **NOT sandbox optimization**. A **material sandbox fraction** (create/teardown >~10%, or an
environment_setup heavy >10 s tail) = real re-provision churn → carry the numbers to the verdict.verifier.finished_at > finished_at — expected 0 (harbor
runs verifier before finalize; the shielded stop/delete follows). Non-zero = release-race signature. Teardown gap
(finished_at − last-phase-finish) is sub-second on a clean run.LastModified slot. A time-clustered
DaytonaAuth/401 spike (concentrated in one slot, absent before/after) is the transient server-side 401 flake
(the _sandbox_exec hot-path missing a retry wrap — same root cause as the eval side), NOT steady-state sandbox
lifecycle and NOT a lease race — report it as transient, do not KILL for it. Only a steady per-slot error rate
is a standing fault.(Reference measurement + numbers: agent_logs/2026-07-15_per-trial-dutycycle-measurement.md.)
RL-JOB-HEALTH — /benjaminfeuer/<job> (<model>, <geometry>, <stage>) captured: <dir>
VERDICT: KILL | NO-KILL | ERROR confidence: high|medium|low
(ERROR = couldn't get the evidence — §0. Give the failed command + its output + what's missing;
do NOT emit KILL/NO-KILL and do NOT default to NO-KILL.)
Evidence I actually read: <quote the state-poll line; the policy-vs-engine util split; the engine
subscription counts; the reward values / exception files. A blank row ⇒ that gate is ERROR, not PASS.>
Restarts: <B/K burned, remaining> — <none | same failure each attempt: … | transient, recovered>
Gate A (liveness): PASS|FAIL|ERROR — <state-poll + log-freshness + any wedge/death signature>
Gate B (resources): PASS|FAIL|ERROR — <policy-vs-engine util; engine subscription (Running/Waiting vs cap);
enforce_eager; OOM? — for under-subscription, the concurrency-bump fix + the live py-spy>
Gate C (rollouts): PASS|FAIL|ERROR — <trials started/completed; rewards; turns; coherence; verifier sanity>
REASONING: <2–4 sentences — the load-bearing evidence, esp. for whatever was "new/untested">
NEXT STEPS: <if KILL: root cause + the concrete fix (config knob / weight-sync / image / infra) + relaunch-or-hold.
if NO-KILL: what to watch next tick + the specific signal that would flip it.
if ERROR: what to fix in the tooling/access so the next probe gets the evidence.>Verdict rules:
B/K + the
traceback + the fix that must land first.You never run the kill. The supervisor executes teardown + relaunch (per rl-agentic-launch-iris /
rl-*-launch-*) on the corrected setting. Log the probe + verdict to ~/Documents/agent_logs/ (dated).
monitor-cron-sweep is the breadth pass..agents/ops/<cluster>/ and code/config semantics in .agents/projects/<dep>/.© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/rl-job-health-deep-dive of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Rl Job Health Deep Dive next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Rl Job Health Deep Dive this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | |
| StandardsItamarZand88/CLI-Anything-WEB | 231 | — | ~4.2k | Automated safety check: Pass | MIT | |
| Cloud UI Change Verificationwarpdotdev/warp | 65k | — | ~1.2k | Automated safety check: Pass | AGPL-3.0 | |
| Dev CompleteFHIR/fhir-codegen | 155 | — | ~9.6k | Automated safety check: Pass | MIT | |
| Webapp QASpecterOps/skills | 706 | — | ~506 | Automated safety check: Pass | Apache-2.0 | |
| Dependabot Fixaxsaucedo/kaos | 280 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 |
ItamarZand88/CLI-Anything-WEB
Runs Phase 4 review/publish/verify for a cli-web- CLI: implementation review by 3 parallel agents, the tiered quality checklist (Tier 1 critical fail-fast, then comprehensive), pip install + smoke…
warpdotdev/warp
Verifies a Warp client change by pushing it to a branch and spawning a cloud agent with computer use that follows the test-warp-ui skill, only when you ask for it.
FHIR/fhir-codegen
Drives the entire local inner loop in one invocation, as a conductor over the skills that own each role.
SpecterOps/skills
Quick-invoke QA testing for a web app URL. An agent skill from SpecterOps/skills.
axsaucedo/kaos
Comprehensively diagnose and fix a failing Dependabot PR. An agent skill from axsaucedo/kaos.
wp-media/wp-rocket
User-facing entry point for the wp-rocket issue workflow. An agent skill from wp-media/wp-rocket.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
Categories
Deep single-RL-job health probe → a KILL / NO-KILL / ERROR recommendation for the supervisor. Rl Job Health Deep Dive is an agent skill from open-thoughts/OpenThoughts-Agent. Deep single-RL-job health probe → a KILL / NO-KILL / ERROR recommendation for the supervisor.
Rl Job Health Deep Dive fits situations like: tasks that involve Subagents; tasks that involve QA and bug reports.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a claude-code`. Or copy the skill folder (.agents/skills/rl-job-health-deep-dive in open-thoughts/OpenThoughts-Agent) into .claude/skills/rl-job-health-deep-dive in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a codex`. Or copy the skill folder (.agents/skills/rl-job-health-deep-dive in open-thoughts/OpenThoughts-Agent) into .agents/skills/rl-job-health-deep-dive in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill rl-job-health-deep-dive -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rl-job-health-deep-dive, .gemini/skills/rl-job-health-deep-dive, .github/skills/rl-job-health-deep-dive and .opencode/skills/rl-job-health-deep-dive in your project.
Going by SKILL.md and its folder, Rl Job Health Deep Dive needs the command-line tools its instructions call (kubectl and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Rl Job Health Deep Dive is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Rl Job Health Deep Dive: Standards (ItamarZand88/CLI-Anything-WEB, 231 stars), Cloud UI Change Verification (warpdotdev/warp, 65k stars), Dev Complete (FHIR/fhir-codegen, 155 stars) and Webapp QA (SpecterOps/skills, 706 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.