Schedule
TinyAGI/tinyagi
Create, list, and delete scheduled tasks (recurring or one-time) that send messages to agents.
Re-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent monitor-restore --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/monitor-restore .claude/skills/monitor-restore && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "monitor-restore" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/monitor-restore into .claude/skills/monitor-restore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitor-restore", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/monitor-restoreType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent monitor-restore --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/monitor-restore .agents/skills/monitor-restore && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "monitor-restore" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/monitor-restore into .agents/skills/monitor-restore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitor-restore", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent monitor-restore --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/monitor-restore .cursor/skills/monitor-restore && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "monitor-restore" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/monitor-restore into .cursor/skills/monitor-restore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitor-restore", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/monitor-restore--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent monitor-restore --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/monitor-restore .gemini/skills/monitor-restore && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "monitor-restore" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/monitor-restore into .gemini/skills/monitor-restore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitor-restore", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent monitor-restoreInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/monitor-restore .github/skills/monitor-restore && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "monitor-restore" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/monitor-restore into .github/skills/monitor-restore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitor-restore", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent monitor-restore --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/monitor-restore .opencode/skills/monitor-restore && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "monitor-restore" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/monitor-restore into .opencode/skills/monitor-restore/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "monitor-restore", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
monitor-restoreRe-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost.
Monitor Restore is an agent skill from open-thoughts/OpenThoughts-Agent. Re-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost. The loop is session-only and is dropped on any session restart, so re-establish it at the start of a new session or whenever the user asks to restore/restart the 3h sweep/cron/monitor. Sets a /loop 3h (or equivalent recurring cron) whose task is the canonical sweep prompt below: status-table active/pending/completed jobs…
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Productivity & Automation, covering Scheduled and recurring tasks and MLOps. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvpythongithfFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Monitor Restore loads about 4.1k tokens when it runs. Until then it costs about 186 tokens; SKILL.md has 410 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 410 words, ~4,143 tokens.
.claude/skills/monitor-restore/SKILL.md (or your agent's skills folder).Re-creates the recurring 3-hour cluster sweep — CoreWeave(iris) + TACC(Vista) + EmpireAI(Beta) (Leonardo DROPPED 2026-07-17 per operator — re-add its sections from git history when it returns; Jupiter SKIPPED, MDC maintenance until ~2026-07-12). The loop is session-only (lost on every restart), so this skill is the durable source of truth for re-installing it.
CronCreate recurring job (auto-expiry).CronList (+ any active /loop). If a recurring sweep mentioning "3-hour … leonardo, coreweave, tacc" (or legacy "jupiter, leonardo") is already live, do not duplicate (doubles the ssh/SQL/Daytona/iris load). Otherwise:/loop: interval 3h, maximum duration the harness allows, task = the canonical prompt below.CronCreate: cron: 17 */3 * * * (off the :00 mark to avoid contention), recurring: true, prompt: = the canonical block verbatim.monitor-cron-sweep (per-cluster Leonardo / CoreWeave / TACC gather+triage); monitor-job-tables + /Users/benjaminfeuer/Documents/notes/ot-agent/job_monitor_table.md (per-type table formats + metric/red-flag definitions; the unified RL table spans Leonardo + CoreWeave).rl-agentic-job-cleanup (agentic), rl-standard-job-cleanup (standard GRPO), sft-job-cleanup, datagen-job-cleanup, eval-agentic-cleanup, eval-standard-cleanup.rl-agentic-launch-iris (CoreWeave RL), rl-standard-launch-leonardo, sft-launch (Leonardo via ops/leonardo/ops.md §SFT), datagen-launch, eval-agentic-launch, eval-standard-launch. (rl-*-jupiter skills apply when Jupiter returns.).agents/ops/leonardo/ops.md, .agents/ops/iris/ops.md (+ ops.md), .agents/ops/tacc/ops.md, .agents/ops/local/ops.md. Dependency facts: .agents/projects/{marinskyrl,harbor,vllm,llama-factory,daytona}/.CLAUDE.md and any memory/skill on conflict./Users/benjaminfeuer/Documents/MarinSkyRL (the user's usual prompt says …/SkyRL). All other paths verbatim..agents/ops/jupiter/ops.md; launch/cleanup = *-jupiter variants) in both this block and the live loop.scripts/iris/iris_ops.py) are a complementary finer-grained layer for actively-debugging runs — not a substitute for the 3h sweep, or vice-versa./loop 3h or CronCreate)3-HOURLY CLUSTER SWEEP — clusters: coreweave(iris), tacc, empireai(beta) (active session only, run to max duration).
[Leonardo DROPPED 2026-07-17 (operator) — re-add its STEP-0 gotchas + gather + the flawed_summ/rl_dlp CAMPAIGN DRIVER from git history when it returns. Jupiter SKIPPED — MDC maintenance until ~2026-07-12.]
⚠ NO EXPERIMENT-SPECIFICS IN THIS PROMPT (they go stale): per-campaign values (in-flight TARGET, refill cluster/grouping/order, harvest gates + discriminators, region/launch gotchas, current bugs) live in the EXPERIMENT TRACKERS under `~/Documents/experiments/active/` (+ the `*-launch` / `*-cleanup` / `analyze-*` skills). READ the relevant tracker each sweep and drive off IT; never hardcode a number/rule here. The CAMPAIGN DRIVER sections below name WHICH tracker to read, not its contents.
STEP 0 (do this FIRST, every sweep): read EACH cluster's ops doc for its BINDING gotchas before touching it —
- empireai(beta) → `.agents/ops/empireai/ops.md`: 2FA keyboard-interactive login — RIDE the operator's ControlMaster socket (`ssh EmpireAI_Beta`; cannot 2FA headless); ⚠ the Beta master socket is REAPED in minutes by the login-node session-killer — if a non-interactive `ssh -o BatchMode=yes` fails, the socket needs an operator reconnect (they keep a warm activity loop) → SKIP + note, don't block; ALWAYS wrap remote cmds in `bash -lc` (SLURM/Pyxis invisible otherwise); user `bf996`, `--account=ny_chinmayh_datacomp`; SBATCH-DETACH all multi-minute work (survives socket death), poll via brief `ssh … sacct`/`tail` windows; storage = HOME `/mnt/home/bf996` (VAST); compute = Pyxis/Enroot containers mandatory.
- coreweave(iris) → `.agents/ops/iris/ops.md` (+ `ops.md`): NO ssh/login — drive via the iris SDK from the Mac; `export KUBECONFIG=~/.kube/coreweave-iris-gpu` is a HARD prereq in the same shell (Mac default kubeconfig points at a DIFFERENT cluster → wrong-context "0 pods/not found"); use the OTAGENT-ENV iris binary `/Users/benjaminfeuer/miniconda3/envs/otagent/bin/iris` (the marin `.venv` iris has a broken `kubernetes` import); all `iris`/`kubectl` calls SYNCHRONOUS (never background); CoreWeave nodes have egress (NO `HF_HUB_OFFLINE`).
- tacc(Vista) → `.agents/ops/tacc/ops.md`: `ssh TACCVista` (hardened single-string ssh); `salloc` BLOCKED → use sbatch; compute nodes have FULL internet (NO proxy/SOCKS/step-ca cert — contrast Leonardo); GPUs are NOT a SLURM gres (whole-node alloc) and RealMemory is misreported; uv OOMs on the shared login node → any build/install goes in a CPU `-p gg` sbatch, never the login node.
PER-CLUSTER GATHER (validate each before trusting; procedure = `monitor-cron-sweep`, particulars = each ops doc):
- EMPIREAI(beta) — `squeue -u bf996` + `sacct -u bf996 -S now-3hours -X` via `ssh -o BatchMode=yes EmpireAI_Beta "bash -lc '…'"` (filter the `module: command not found` / `Loading gcc` noise). If BatchMode ssh FAILS (socket reaped), NOTE "EmpireAI socket down — needs operator reconnect" and SKIP (can't 2FA headless) — do NOT block. For any live Axolotl SFT job: poll `~/logs/<name>_<jobid>.out` for the latest step + read the latest `checkpoint-<N>/trainer_state.json` (`log_history[]` step/loss/grad_norm; `global_step`/`max_steps`); fetch that small JSON to the relevant experiment dir per the ops-doc recipe, NEVER checkpoints. READ the active EmpireAI experiment tracker under `experiments/active/` and drive off IT.
- COREWEAVE(iris) — STATE-POLL the authoritative iris lifecycle, NOT a log-string watch (a clean kill/eviction/preempt emits no terminal log line + reaps the pods):
KUBECONFIG=~/.kube/coreweave-iris-gpu; PY=/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python
$PY scripts/iris/iris_ops.py /benjaminfeuer/<job> --once --json # per active job (auth state now)
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/iris --cluster=cw-us-east-02a job summary --json # authoritative
Treat "running-but-0-pods / record disappeared" as TERMINAL (silent-wedge signature). `iris … query` over the jobs table lists live jobs (state 1/2/3). Full log (init→crash) via `iris … job logs --since-ms <submitted_at_ms> --no-tail` (finelog keeps the WHOLE log; only `--tail` caps lines).
- TACC(Vista) — `squeue -u penfever` + `sacct -u penfever -S now-3hours -X` via `ssh TACCVista`.
IN-FLIGHT / ACTIVE jobs → report in a UNIFIED TABLE per job type, spanning all clusters (RL table = CoreWeave agentic/MoE rows; SFT/build = EmpireAI mega-container rows; Eval = TACC agentic rows). Structure = `monitor-job-tables` / notes/ot-agent/job_monitor_table.md (box-drawing, not markdown). RL rows MUST include entropy + collapse signals (grad_norm / log_ratio), not just step+reward.
CHAIN-RESTART TIMEOUTs are NORMAL (note the afterany successor), not failures. On CoreWeave, `--max-retries` re-brings-up the gang on a transient HF-weight-resolution flake — a single retry is a normal time-cost, not a fault.
EMPIREAI (SFT workstream — monitor to completion):
- READ the active EmpireAI experiment tracker under `experiments/active/` each sweep and drive off IT — never hardcode a run/campaign here (they go stale). Table each live Axolotl SFT run in the SFT bucket (step/loss/grad_norm from the latest `checkpoint-<N>/trainer_state.json`; `global_step`/`max_steps`). Fetch the small result JSONs to the experiment dir per the ops-doc recipe — never the checkpoints.
- On an SFT run COMPLETED → `sft-job-cleanup` (consolidate → HF upload the ZeRO-3-shard path → reclaim disk) → then the downstream eval per the tracker. Sequential-run campaigns fire the next run per the tracker.
- All SFT is SBATCH-DETACHED; the Beta socket is flaky (skip+note if `ssh -o BatchMode=yes` fails — needs an operator reconnect, can't 2FA headless).
COREWEAVE RL (agentic SkyRL/MoE via `rl-agentic-launch-iris`):
- Report BRING-UP for fresh launches: gang/leafgroup admission (Kueue, pods SchedulingGated until atomically admitted is normal), `apply_ep` / mesh-load, weights resolving. `shm_broadcast: …60s` + a transient ghcr EOF → ImagePullBackOff self-heal are BENIGN bring-up noise.
- EP=8 science greps (the 131k arm) — `sel_rows` / `EPDIAG` via `scripts/iris/analyze_iris_harbor_job.py` (log-content greps are SCIENCE/throughput ONLY, never liveness — liveness = the state-poll above).
- On COMPLETED → route by flavor WITHOUT asking: AGENTIC (Harbor/Daytona/terminal_bench) → `rl-agentic-job-cleanup` (FULL checklist incl. trace upload + metrics); STANDARD/non-agentic GRPO → `rl-standard-job-cleanup`.
- Per-run Monitors (bring-up/wedge watch) are a complementary finer-grained layer; this 3h cron is the baseline — don't let one substitute for the other.
TACC EVAL HARVEST (when present — newly-integrated, validated by a canary):
- TACC agentic eval runs through the front door `python -m hpc.launch --job_type eval_listener --cluster-config tacc` (bare name resolves from `HPC.eval_cluster_view`; `sbatch_script` = `eval/tacc/eval_harbor.sbatch`, `eval_jobs_dir` = `/scratch/10635/penfever/eval_jobs`; whole-node alloc, no `--gres`/`--mem`; compute nodes have egress → NO proxy/cert). Once a leg is RUNNING, harvest finished TACC evals the same way as Leonardo (`eval-agentic-cleanup` if auto-register failed). Sanity-check the canary's traces uploaded + registered before relying on it.
ON SUCCESSFUL COMPLETION (SFT / RL / datagen / eval) on ANY cluster → note it + summary stats, then route WITHOUT asking:
- RL → route by flavor: AGENTIC (Harbor/Daytona/terminal_bench) → `rl-agentic-job-cleanup`; STANDARD / non-agentic GRPO (Delphi/rlvr/dapo math cells) → `rl-standard-job-cleanup` (model + metric CSVs only; size suffix from the exported weights; DB-register only if the series is DB-registerable). SFT → `sft-job-cleanup`. (Leonardo HF upload = the sbatch-tunnel path, NOT the login node — it SIGKILLs long processes at ~100s; needs the fresh step-ca cert.)
- Datagen → verify traces uploaded to HF (penfever org); if NOT, dispatch a subagent (`datagen-job-cleanup`).
- Eval where DB registration FAILED for a technical reason → dispatch a subagent through ALL steps of `eval-agentic-cleanup`; confirm each completed, dispatching another if any were missed.
- INODES (Leonardo/GPFS): every cleanup MUST `rm` the on-disk artifact tree (`trace_jobs/`/`tasks/`) after HF upload confirmed + verify reclaim — the #1 inode leak. (CoreWeave artifacts go to HF / R2, not POSIX scratch; no on-disk tree to reap there.)
ON ANY JOB THAT FAILED since the last check → dispatch a subagent to determine cause + propose fixes. ANNOUNCE the choices, SELECT one, and apply changes + relaunch via another subagent. Keep a running DATED log of failures (job ID + remediation) in /Users/benjaminfeuer/Documents/agent_logs/.
- If an RL job EXHAUSTED all restarts WITHOUT reaching max steps AND the failure looks recoverable (transient) → queue 5 more restarts (Leonardo) / re-launch with `--max-retries ≥1` (CoreWeave). Spike-mitigation ablations are exempt from auto-cancel — observing the recovery IS the experiment (`monitor-cron-sweep`).
CODE / CONFIG EDITS → edit LOCALLY on the active branches:
/Users/benjaminfeuer/Documents/{OpenThoughts-Agent,vllm,harbor,MarinSkyRL}.
Local clones are GROUND TRUTH — clusters never diverge (no untracked/divergent changes, no hand-editing, no patch-by-rsync). Sync the Python repos by commit+push then `git pull` on the SLURM clusters (editable installs, live after pull); CoreWeave has NO clone to pull — the iris launcher uploads the local workspace to `/app` so a local commit takes effect on the next launch. EVERY SWEEP, run `git status --short` on each SLURM cluster repo (leonardo, tacc) and triage drift back to local: TRACK reusable files (commit local → push), GITIGNORE recurring transient junk (`*.bak`, `*_manifest.txt`, `&1`, ephemeral `reeval_priority_*`); reconcile with `git pull`, NEVER `git reset --hard` while live jobs depend on uncommitted state (`monitor-cron-sweep` §4). vLLM (compiled fork) → commit+push the fork, then BUILD FROM SOURCE on each cluster from that commit (never rsync / hand-patch); CoreWeave rebuilds the gpu-rl image (bump the digest) only when the compiled vLLM fork changes — first-party + MarinSkyRL fixes go live without a rebuild.
LOCAL WORKTREE + MAIN HYGIENE (every sweep — operator 2026-07-17): subagents spawn git WORKTREES (marin-fork PR flow) that pile up. (1) PRUNE stale worktrees — `git worktree list` per repo (OpenThoughts-Agent, vllm, harbor, MarinSkyRL, marin, evalchemy); `git worktree remove` (NO `--force` — git refuses a dirty one so no branch/commit is lost; the branch always survives on origin) any whose branch is MERGED/abandoned or whose work is done; HOLD only worktrees a LIVE job or an ACTIVE subagent is using. (2) KEEP PRIMARY CLONES ON CANONICAL BRANCH — marin forks (MarinSkyRL/marin/evalchemy) on `main`, OT-Agent/vllm on `penfever/working`; a clone parked on a feature branch is a live footgun (a launch from that dir uploads that branch) → reset it (`git checkout main && git pull --ff-only`) once its branch is pushed/clean. (3) KEEP `main` CLEAN — no uncommitted tracked drift on a primary clone.
ACTIVELY-DEBUGGING jobs → monitor more closely than stable ones. For any FRESH launch, set one-time checks at 15 min and 30 min after launch to catch new failures early.
LAUNCHING FRESH JOBS → follow the per-job-type launcher instructions in CLAUDE.md (+ the `*-launch-*` skills and `.agents/projects/ot-agent/ot-agent.md`). If unclear, ASK.
EXPERIMENT LOG → each launch / state change logged as a standalone dated file under /Users/benjaminfeuer/Documents/agent_logs/ (YYYY-MM-DD_<topic>.md) — no monodoc.
STANDING CONSTRAINTS (do not violate without explicit permission): enable_db_registration stays false in YAMLs (manual DB register only); Daytona RUNNING RL ≤ 6 per cluster; a3 series is CONCLUDED (no launch/refill/auto-advance); Daytona snapshot caps are HARD (clean stale, never raise); cross-user FK safety pre-check before any Supabase delete/mutate; HF uploads default PUBLIC to laion/. NEVER kill/restart a RUNNING job (or `iris cluster restart`) without express permission. Skip an unreachable cluster (note it) rather than blocking. This prompt OVERRIDES any memory/skill on conflict.If you change the cadence or scope, update BOTH the block above AND the live loop/cron (delete + recreate) so this skill stays the canonical copy.
© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/monitor-restore of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Monitor Restore next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Monitor Restore this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | |
| ScheduleTinyAGI/tinyagi | 3.6k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Send User MessageTinyAGI/tinyagi | 3.6k | — | ~829 | Automated safety check: Pass | MIT | |
| Cron Opsczl9707/build-your-own-openclaw | 1.9k | — | ~593 | Automated safety check: Pass | MIT | |
| X Bookmarkssharbelxyz/x-bookmarks | 289 | — | ~2k | Automated safety check: Notes | None | |
| Jobsphysiclaw/PhysiClaw | 386 | 1 repos | ~1.1k | Automated safety check: Pass | MIT |
TinyAGI/tinyagi
Create, list, and delete scheduled tasks (recurring or one-time) that send messages to agents.
TinyAGI/tinyagi
Send a proactive message to a paired user via their channel (Discord, Telegram, or WhatsApp).
czl9707/build-your-own-openclaw
Create, list, and delete scheduled cron jobs. An agent skill from czl9707/build-your-own-openclaw.
sharbelxyz/x-bookmarks
Fetch, summarize, and manage X/Twitter bookmarks via bird CLI or X API v2.
physiclaw/PhysiClaw
A skill your agent uses when the task involves scheduling future work — any "remind me at …", "every weekday …", "check again in 30 min", or closing a fired cron job.
Automattic/agent-skills
A skill your agent uses when working with WP-CLI (wp) for WordPress operations: safe search-replace, db export/import, plugin/theme/user/content management, cron, cache flushing, multisite, and…
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
Categories
Re-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost. Monitor Restore is an agent skill from open-thoughts/OpenThoughts-Agent. Re-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost.
Monitor Restore fits situations like: tasks that involve Scheduled and recurring tasks; tasks that involve MLOps.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a claude-code`. Or copy the skill folder (.agents/skills/monitor-restore in open-thoughts/OpenThoughts-Agent) into .claude/skills/monitor-restore in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a codex`. Or copy the skill folder (.agents/skills/monitor-restore in open-thoughts/OpenThoughts-Agent) into .agents/skills/monitor-restore in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill monitor-restore -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitor-restore, .gemini/skills/monitor-restore, .github/skills/monitor-restore and .opencode/skills/monitor-restore in your project.
Going by SKILL.md and its folder, Monitor Restore needs the command-line tools its instructions call (uv, python, git and hf). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Monitor Restore is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Monitor Restore: Schedule (TinyAGI/tinyagi, 3.6k stars), Send User Message (TinyAGI/tinyagi, 3.6k stars), Cron Ops (czl9707/build-your-own-openclaw, 1.9k stars) and X Bookmarks (sharbelxyz/x-bookmarks, 289 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.