GPU Optimizer
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch-iris --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-agentic-launch-iris .claude/skills/eval-agentic-launch-iris && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-agentic-launch-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch-iris into .claude/skills/eval-agentic-launch-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch-iris", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch-irisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch-iris --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/eval-agentic-launch-iris .agents/skills/eval-agentic-launch-iris && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-agentic-launch-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch-iris into .agents/skills/eval-agentic-launch-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch-iris", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch-iris --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/eval-agentic-launch-iris .cursor/skills/eval-agentic-launch-iris && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-agentic-launch-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch-iris into .cursor/skills/eval-agentic-launch-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch-iris", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/eval-agentic-launch-iris--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch-iris --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/eval-agentic-launch-iris .gemini/skills/eval-agentic-launch-iris && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-agentic-launch-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch-iris into .gemini/skills/eval-agentic-launch-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch-iris", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch-irisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/eval-agentic-launch-iris .github/skills/eval-agentic-launch-iris && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-agentic-launch-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch-iris into .github/skills/eval-agentic-launch-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch-iris", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch-iris --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/eval-agentic-launch-iris .opencode/skills/eval-agentic-launch-iris && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-agentic-launch-iris" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch-iris into .opencode/skills/eval-agentic-launch-iris/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch-iris", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-agentic-launch-irisLaunch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint.
Eval Agentic Launch Iris is an agent skill from open-thoughts/OpenThoughts-Agent. Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint. Use when asked to start, watch, or kill a model evaluation (evalchemy / agent-harness benchmarks) on Iris.
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering GPU and accelerator computing and Machine learning. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythoncondahfgitgsutilFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
iris.oa.devFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DAYTONA_API_KEYDAYTONA_B_KEYDAYTONA_RL_API_KEYDAYTONA_DATA_API_KEYSUPABASE_SERVICE_ROLE_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Agentic Launch Iris loads about 3.8k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 1,508 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,508 words, ~3,825 tokens.
.claude/skills/eval-agentic-launch-iris/SKILL.md (or your agent's skills folder).📍 Iris orientation — read first. Read the Iris tools catalog (
.agents/ops/iris/ops.md) and the Iris ops directory (.agents/ops/iris/— CoreWeave GPU inops.md, TPUmarininops.md) before acting.
Launch → monitor → manual cleanup of an eval job via eval/cloud/launch_eval_iris.py (Iris analog of the SkyPilot launch_eval_cloud.py). For datagen/tracegen use datagen-launch-iris instead.
🚪 Iris/cloud launchers bypass
hpc.launchentirely — thepython -m hpc.launch --job_type eval_listenerSLURM front door does NOT apply here. The Iris launcher resolves per-model serve config frommodel_config/(viamodel_config/resolver.py+hpc/model_config_apply.py): merges/forwards the model'sagent_kwargs, and applies serve intrinsics (max_model_len/limit_mm/extra_args) on the worker. It also appliesn_attemptsfrom the preset (CLI-overridable via--n_attempts) and ignores SLURM-only fields plustp_size(TPU chip count),harbor_config(CLI-required),agent_name(from harbor config). It prints[eval-iris] preset <name>: applied {…}; ignored {…}. Precedence: explicit CLI/--preset>model_config/. Edit the source atmodel_config/<org>/<slug>.yaml(not the generated registry).
model — model id for --model (HF id or GCS/served path), OR pass --datagen_config <yaml> (model inferred from its engine.model).dataset — for standard benchmarks, use --preset <name> (below), which selects the dataset. Pass an explicit dataset only for a custom benchmark or to override a preset:--dataset <harbor slug> — harbor resolves/snapshots it.--dataset_path <tasks dir | HF dataset id> (mutually exclusive with --dataset). A bare HF id has exactly one /, no leading ./,/,~; the worker's run_eval.py resolves it (snapshot_download + convert_parquet_to_tasks) — the launch host does NOT.harbor_config — REQUIRED, an eval harbor YAML from hpc/harbor_yaml/eval/:dcagent_eval_defaults.yaml — DEFAULT. Iris-adapted port of the eval team's canonical config (hpc/harbor_yaml/eval/configs/dcagent_eval_config.yaml, the SLURM listener's EVAL_CONFIG_YAML): terminus-2, timeout_multiplier: 1.0, n_attempts: 3, agent max_timeout_sec: 7200, verifier max_timeout_sec: 14400. Iris numbers match the eval team's SLURM numbers. Only deviation: force_build: true (Iris builds sandboxes at runtime).eval_ctx32k.yaml / eval_ctx131k.yaml — terminus-2 with timeout_multiplier: 8.0 (8GB/4GB sandbox). Extended-budget mode — only for deliberate 8× timeout. Don't use for normal reg eval.eval_openhands_ctx32k_* / eval_mini_swe_ctx32k.yaml / swe_agent_ctx32k_eval_.yaml — alternate harnesses (OpenHands / mini-SWE / SWE-agent). Only when reproducing a paper's harness.--preset, shared with the SLURM listener)--preset <name> pulls run defaults from eval/presets/ (one YAML per preset, same catalog the SLURM eval/unified_eval_listener.py consumes). Choices: aider, bfcl, financeagent, gaia, medagentbench, swebench, swebench_full, tb2, v1, v2. Precedence: explicit CLI flags ALWAYS override preset values.
What the Iris launcher does with each preset field:
datasets[0] → --dataset_path (bare HF id, resolved on the worker) when neither --dataset nor --dataset_path was passed (extra datasets skipped, logged); n_concurrent → --n_concurrent when not passed.eval/jupiter/eval_harbor.sbatch does): agent_parser → harbor --agent-kwarg parser=<value> (e.g. swebench → parser=xml) unless you passed a parser=; each preset agent_kwargs list entry → its own --agent-kwarg key=value (your --agent_kwarg with the same key overrides).agent_kwargs from model_config/, so thinking IS auto-applied per-model for models carrying agent_kwargs: [extra_body={…enable_thinking:true}]. For a model with no model_config/ entry, thinking falls back to the served model's chat-template default (Qwen3 = ON). For a default-OFF template model not in model_config/ (e.g. Qwen3.5/3.6), pass --agent_kwarg 'extra_body={"chat_template_kwargs":{"enable_thinking":true}}' (the live nested form vLLM applies; a bare enable_thinking=true is DEAD — terminus-2 has no such param). There is no --enable-thinking flag.slurm_time, vllm_max_retries, gpu_memory_util, sbatch_script, check_hf_exists, log_suffix, error_threshold, config_yaml, agent_envs, auto_snapshot.--preset composes with --harbor_config (required), --model, --upload_to_database, etc.
The standard/core evals are presets — launch by name (preset sets dataset, concurrency, parser; do not pass --dataset*):
| Benchmark | Command | preset sets |
|---|---|---|
| SWE-bench-verified (random 100) | --preset swebench | DCAgent2/swebench-verified-random-100-folders, n_concurrent 32, parser=xml |
| terminal-bench 2.0 | --preset tb2 | DCAgent2/terminal_bench_2, n_concurrent 32 |
Presets do not set thinking — see the note above. Both require --harbor_config hpc/harbor_yaml/eval/dcagent_eval_defaults.yaml (terminus-2 @ 32k, eval-team default budget, the Cat 1 "reg eval" harness per docs/EVAL_GUIDE.md), fit a v6e-4 for an 8B model, and should launch with --upload_to_database. For full parity with the eval team's SLURM runs, pass --n_concurrent 128 (their CLI default). (terminal-bench 2.0 also exists as slug --dataset terminal-bench@2.0; prefer the preset.)
Eval does NOT pre-build Daytona snapshots and does NOT call hpc/snapshot_manager.ensure_snapshots. Eval harbor configs set environment.force_build: true — harbor builds each task's sandbox at runtime on the worker, no launch-host prebuild, no 60-snapshot cap, no SnapshotCapExceeded. Datagen is the opposite (force_build: false → pre-builds; see datagen-launch-iris).
Always run eval out of the MAIN Daytona org (DAYTONA_API_KEY, carried via --secrets-env). Do NOT use DAYTONA_B_KEY / DAYTONA_RL_API_KEY / DAYTONA_DATA_API_KEY (other workloads).
Launch from the py3.12 otagent conda env, source "$DC_AGENT_SECRET_ENV" (see .agents/secret.md; pass --secrets-env), and git pull the marin checkout if the iris client is reported too old. Harbor env defaults to daytona (the only sandbox backend that works on iris workers).
cd /Users/benjaminfeuer/Documents/OpenThoughts-Agent
source /Users/benjaminfeuer/miniconda3/etc/profile.d/conda.sh && conda activate otagent
source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"
TS=$(date +%Y%m%d-%H%M%S)
python eval/cloud/launch_eval_iris.py \
--preset <name> \ # e.g. swebench, tb2 — seeds dataset + concurrency + parser
--harbor_config hpc/harbor_yaml/eval/dcagent_eval_defaults.yaml \ # eval-team defaults (timeout_multiplier 1.0); use eval_ctx32k.yaml only for 8x budget
--model <hf-or-gcs-model-id> \
--tpu v6e-4 --preemptible \
--job_name "eval-<model-slug>-<bench>-${TS}" \
--secrets-env "$DC_AGENT_SECRET_ENV" \
--upload_to_database \
--no-wait
# Custom benchmark (no preset): drop --preset and pass --dataset <harbor-slug>
# or --dataset_path <tasks dir | HF id>, plus --n_concurrent <N>.cw-us-east-02a)Pass --gpu H100x8 (mutually exclusive with --tpu) for one CoreWeave H100x8 node. The launcher defaults to the gpu-8x OT-Agent image, the cw-us-east-02a iris config, the datagen extra (not datagen-tpu), and skips the TPU iris-serve/patch_tpu_inference path. export KUBECONFIG=~/.kube/coreweave-iris-gpu first. Single-node only — do NOT pass --replicas > 1 (task sharding + shared multi-node vLLM not implemented for GPU eval); --gpu is limited to H100x8. Use a model known to serve on the runtime (Qwen/Qwen3-32B works).
Daytona/OpenCode against a separately served CoreWeave model: use native federated ingress, not the peer controller's public host. Submit the serving job through Marin with --target-cluster cw-us-east-02a and forward the Marin login; wait until iris --cluster=marin endpoints list <endpoint> --exact shows the mirrored peer endpoint, then mint the scoped URL at Marin and pass only https://iris.oa.dev/proxy/t/<token>/<endpoint>/v1 to the eval. A token minted with iris --cluster=cw-us-east-02a endpoints mint is peer-signed and cannot authorize the Marin parent route; iris-cw-us-east-02a.oa.dev is IP-locked and not Daytona-reachable. scripts/iris/launch_external_opencode_eval.py is the one-command Grug profile: it submits the parent-delegated serve, waits for the mirrored ready endpoint, mints at Marin, and submits the durable-S3 Harbor eval. Its defaults use the established Grug serve topology and CoreWeave S3 root; use explicit flags only to override that profile or attach an existing endpoint.
export KUBECONFIG=~/.kube/coreweave-iris-gpu
python eval/cloud/launch_eval_iris.py \
--preset swebench \
--harbor_config hpc/harbor_yaml/eval/dcagent_eval_defaults.yaml \
--model Qwen/Qwen3-32B \
--gpu H100x8 --replicas 1 \
--n_concurrent 3 --n_attempts 1 \
--harbor_extra_arg=--n-tasks=3 --harbor_extra_arg=--max-retries=0 \ # small subset for fast iteration
--job_name "eval-<slug>-cw-gpu-${TS}" \
--secrets-env "$DC_AGENT_SECRET_ENV" \
--upload_to_database --no-wait--upload_to_database (opposite of datagen's --skip_register). Registers result abstracts to Supabase and uploads traces to HF (repo auto-derived from --job_name when --upload_hf_repo omitted). Requires SUPABASE_URL + SUPABASE_SERVICE_ROLE_KEY in --secrets-env. Companion flags: --upload_username (attribution; defaults $UPLOAD_USERNAME/current user), --upload_error_mode {skip_on_error,rollback_on_error}, --upload_forced_update. No --register/--skip_register — sync is OFF by default, ON solely via --upload_to_database.--upload_hf_repo alone (no --upload_to_database) = HF-only, no Supabase.--upload_hf_repo pushes results to HF on completion (image ae085bc8+ wires harbor's --export-push); omit for local/GCS-only.--tpu defaults to v6e-4 for eval (vs v5p-8 for S1 datagen); set per model footprint.--model is optional only when --datagen_config is given (model inferred); otherwise required.docs/EVAL_GUIDE.md (benchmark/harness catalog) and scripts/iris/EVAL_GUIDE.md/README.md (eval-analysis tooling).--local-sync-dir while the job runs (local eval-analysis tooling sees files). Pass --output-mode gcs (and OMIT --gcs-output-dir) to write straight to a co-located single-region bucket (gs://marin-us-east5/ot-agent, …). An explicit --gcs-output-dir gs://marin-models-us/ot-agent opts OUT of the pin (pricier multi-region) — only for the stuck-PENDING dodge when a TPU pool has collapsed.--output-mode local — Harbor writes trace_jobs to pod-local NVMe and run_eval --upload_to_database registers to Supabase + HF in-pod before the ephemeral pod tears down (same path TPU/SLURM use). --upload_to_database IS supported on GPU. For durable raw Harbor artifacts: --output-mode s3 --s3-output-dir s3://marin-us-east-02a/tmp/ttl=7d/ot-agent/evals/<user> (CW object store). ⚠ Prefer deriving the output dir off marin_prefix() (rigging.filesystem — auto-resolves the storage root; don't hardcode the region bucket); the literal is a fallback. Never use s3://marin-na (R2) — pods can't reach it.AWS_*/LAION_*/MARIN_HMAC_* from the pod (can't clobber the R2 creds the cw-us-east-02a cluster injects via the iris-task-env envFrom Secret). Do not re-add them.Confirm placement (same as datagen):
/Users/benjaminfeuer/Documents/marin/.venv/bin/iris --cluster=marin query \
"SELECT job_id, state FROM jobs WHERE job_id='/benjaminfeuer/<job>'" -f csvSame job-agnostic analyzer as datagen:
/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python \
/Users/benjaminfeuer/Documents/OpenThoughts-Agent/scripts/iris/analyze_iris_harbor_job.py \
/benjaminfeuer/<job> --output /tmp/<job>_history.md --resyncFor eval, the signals of interest are completion + productive trial rate (non_empty_trials/total_trial_dirs) and the harness exception stats, more than gen tok/s. Scores land in the synced outputs, not the analyzer sidecar:
--local-sync-dir on the launch host;--output-mode gcs → under the pinned single-region bucket (e.g. gs://marin-us-east5/ot-agent/<job>/; resolve with python -m hpc.iris.job_output_resolver <job> --cluster …/marin.yaml).Per-task progress / resume helpers: scripts/iris/check_progress.py and check_resume_needed.py.
Kill (only with explicit user permission for a RUNNING job):
/Users/benjaminfeuer/Documents/marin/.venv/bin/iris --cluster=marin job kill /benjaminfeuer/<job>Recover partial results: outputs are already on the launch host (--local-sync-dir) or in GCS (--output-mode gcs). To re-pull a GCS job dir, resolve the recorded output prefix first (never hardcode): OUT=$(python -m hpc.iris.job_output_resolver <job> --cluster …/marin.yaml) then gsutil -m rsync -r "$OUT/<job>/" /tmp/<job>_eval/. If HF upload didn't fire and you need traces on the Hub, use the same make_and_upload_trace_dataset.py recipe as datagen-launch-iris against the local job dir.
Daytona snapshot cap: N/A for eval (no pre-build, no ensure_snapshots, eval configs use force_build: true). If you see SnapshotCapExceeded, you're on the wrong (datagen) path or wrong harbor config.
Stuck PENDING: relaunch with --output-mode gcs --gcs-output-dir gs://marin-models-us/ot-agent (unpinned — deliberate override drops the single-region pin so iris places on any free TPU in the US). Kill the stuck submission first only with user permission.
DAYTONA_API_KEY) — never the B/RL/DATA orgs. Eval builds sandboxes at runtime (force_build: true); it does not pre-build or call ensure_snapshots.--harbor_config to the model's context window and the benchmark's harness (plain vs OpenHands/mini-SWE/SWE-agent) — a mismatch fails at runtime.© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/eval-agentic-launch-iris of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Eval Agentic Launch Iris next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Agentic Launch Iris this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| GPU OptimizerMathews-Tom/armory | 328 | — | ~3.5k | Automated safety check: Notes | MIT | |
| DGX Spark Memory and Thermal Opswshobson/agents | 40k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Ray Train Distributed TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.7k | Automated safety check: Pass | MIT | |
| Optimize For GPUK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Pass | MIT |
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
Orchestra-Research/AI-Research-SKILLs
Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.
Orchestra-Research/AI-Research-SKILLs
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
majiayu000/claude-skill-registry
GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
Categories
Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint. Eval Agentic Launch Iris is an agent skill from open-thoughts/OpenThoughts-Agent. Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint.
Eval Agentic Launch Iris fits situations like: kill a model evaluation (evalchemy / agent-harness benchmarks) on Iris; tasks that involve GPU and accelerator computing; tasks that involve Machine learning.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a claude-code`. Or copy the skill folder (.agents/skills/eval-agentic-launch-iris in open-thoughts/OpenThoughts-Agent) into .claude/skills/eval-agentic-launch-iris in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a codex`. Or copy the skill folder (.agents/skills/eval-agentic-launch-iris in open-thoughts/OpenThoughts-Agent) into .agents/skills/eval-agentic-launch-iris in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch-iris -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-agentic-launch-iris, .gemini/skills/eval-agentic-launch-iris, .github/skills/eval-agentic-launch-iris and .opencode/skills/eval-agentic-launch-iris in your project.
Going by SKILL.md and its folder, Eval Agentic Launch Iris needs the command-line tools its instructions call (python, conda, hf, git and gsutil) and credentials named DAYTONA_API_KEY, DAYTONA_B_KEY, DAYTONA_RL_API_KEY and DAYTONA_DATA_API_KEY. Our summary lists: Python 3; A credential in DAYTONA_API_KEY; A credential in DAYTONA_B_KEY.
SKILL.md names 1 domain. In commands or code: iris.oa.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Agentic Launch Iris is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Agentic Launch Iris: GPU Optimizer (Mathews-Tom/armory, 328 stars), DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars), PyTorch Lightning Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Ray Train Distributed Training (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.