Agent skill

Eval Agentic Cleanup

by open-thoughts in open-thoughts/OpenThoughts-Agent

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Eval Agentic Cleanup

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-cleanup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-agentic-cleanup .claude/skills/eval-agentic-cleanup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-agentic-cleanup
GitHub stars
301
Token cost
~5.8k tokens
SKILL.md length
2,221 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

  • Works in 6 steps: Read-only completeness + health AUDIT —… → Manual upload + DB register —… → ⚠️ Verify + fix the MODEL name (with… → …
  • Verify an eval is truly complete
  • SKILL.md covers 0. Read-only completeness +…, 1. Manual upload + DB register…, 2. ⚠️ Verify + fix the MODEL… and 2b. ⚠️ ALSO verify + fix the…, plus 4 more sections
  • Calls python, gsutil and python3; needs LLM_API_KEY and DAYTONA_API_KEY

What it does

Eval Agentic Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent. Audit + recover a finished agentic eval. ALWAYS start with the read-only, idempotent completeness/health audit (§0): job finished? score present + non-zero + not obviously broken? HF traces present + linked? trial count ≈ nrep × benchmarksize? — it writes nothing and recommends an action per check. Then run only the flagged remediations: manual HF-trace upload + Supabase DB-registration via manualdbevalpush.py, the vLLM-numeric-ID → real-HF-model-name fix (with cross-user FK safety), verify, free disk. Use to…

Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and Supabase. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Verify an eval is truly complete
  • A sweep finds an eval that didnt upload/register
  • Has a broken/zero score
  • Short trial count

Example prompts

  • “/eval-agentic-cleanup”

Requirements

  • Python 3
  • A credential in LLM_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Read-only completeness + health AUDIT — run first (idempotent, ZERO writes)
  2. Manual upload + DB register — manual_db_eval_push.py
  3. ⚠️ Verify + fix the MODEL name (with cross-user FK safety)
  4. Verify it landed
  5. Clean up disk (only after upload + register verified)
  6. IRIS TPU evals — the idempotent path (GCS-backed, no SLURM)

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • gsutil
    • python3
    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LLM_API_KEY
    • DAYTONA_API_KEY
    • SUPABASE_ANON_KEY
    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Agentic Cleanup loads about 5.8k tokens when it runs. Until then it costs about 216 tokens; SKILL.md has 2,221 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~216
When it runs · the whole SKILL.md, loaded when a task matches
~5.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 2,221 words, ~5,821 tokens.

Download SKILL.mdSave it as .claude/skills/eval-agentic-cleanup/SKILL.md (or your agent's skills folder).
name
eval-agentic-cleanup
description
Audit + recover a finished agentic eval. ALWAYS start with the read-only, idempotent completeness/health audit (§0): job finished? score present + non-zero + not obviously broken? HF traces present + linked? trial count ≈ n_rep × benchmark_size? — it writes nothing and recommends an action per check. Then run only the flagged remediations: manual HF-trace upload + Supabase DB-registration via manual_db_eval_push.py, the vLLM-numeric-ID → real-HF-model-name fix (with cross-user FK safety), verify, free disk. Use to verify an eval is truly complete, when a sweep finds an eval that didn't upload/register or has a broken/zero score or short trial count, or to re-register/correct an eval's model/traces. Distinct from the model- publishing cleanups (rl-agentic-job-cleanup / sft-job-cleanup) and datagen-job-cleanup — this is the EVAL path.

eval-agentic-cleanup

Always start with the read-only audit (§0): it verifies the four conditions that actually matter, then run only the remediations it flags. "The score is in the leaderboard" does NOT mean the eval is complete — the auto-harvest path can land a score while silently skipping trace upload, or record a bogus 0 from an eval that never reached the model.

0. Read-only completeness + health AUDIT — run first (idempotent, ZERO writes)

Reads only (sacct, run-dir result.json, HF API, Supabase reads). Emits a per-eval verdict + recommended action per check; run only the flagged remediations.

Pre-gate — EXEMPT: grid / throughput / OOM-test measurement runs targeting DCAgent2/* → skip (upload only production winners).

#CheckHow (read-only)PASSIf it FAILS → recommend
1Job finishedsacct -j <id> -X -o StateCOMPLETEDRUNNING/PENDING → wait & re-audit; FAILED/TIMEOUT/CANCELLED → diagnose + relaunch (see eval-agentic-launch gotchas #1–#5: jobs_dir perm / hosted_vllm 2-slash / model_info / reservation / stale-row)
2Score present & not brokenSupabase sandbox_jobs score, or run-dir result.json → stats.evals.<k>.metrics[0].meannon-null real number; 0 OK only if check 4 shows trials rannull/missing/0 with ~0 trials/POSTs = infra-broken (eval never reached the model — classic jobs_dir/hosted_vllm/model_info failure) → re-run. High VerificationNotCompletedError rate → score may be biased (note below).
3HF traces present + linkedtrace HF repo exists + non-empty (HF API 200) AND DB/LB row has trace/url field setrepo 200 AND linkedmissing/empty OR unlinked → upload traces + register link (§1–§3)
4VALID trial count (PARSE, don't file-count)Parse each per-trial result.json for a numeric reward (snippet below). An errored trial STILL writes result.json with exception_info and no reward, so file-count overstates coverage.VALID is ~90–100% of n_rep_eval × benchmark_size, with at most n_rep_eval retained attempts per taskVALID or total above plan / more than n_rep_eval attempts for a task → deduplicate (§0a), then re-audit. Short VALID (<~90%, a missing rep, or high hard-error 15–25%) → resume errored trials. Expected: swebench-verified-random-100=100, dev_set_v2=100, terminal_bench_2=89; × n_rep_eval (default 3) → ~300/~300/~267.

Output one row per eval: ✅/⚠️ per check + the single recommended next action. Re-run the audit after any remediation to confirm ✅.

No-reward trial classes (VNC / infra errors) — bias direction
  • VerificationNotCompletedError (harbor 9203989f) / pre-fix AgentTimeoutError WITH NO reward = the verifier never produced a result. Missing because the agent burned its budget → true score ≈ 0 (SWE-bench ~pass/fail). Harbor's JobStats mean DROPS these from the denominator → biased UP; the DB/leaderboard harvest counts them as 0 → unbiased. Either way retriable: resume on a healthy sandbox — the payoff is usually recovering real passes hidden as 0s.
  • Infra errors (DaytonaAuthenticationError/DaytonaValidationError/etc.) = failure uncorrelated with solvability → dropping is ~unbiased (lowers N). Re-run for completeness only; do NOT impute 0.
  • AgentTimeoutError WITH a reward = verifier scored it → passthrough VALID 0. Don't filter/re-run.
  • Discriminate by reward presence, not the error label.
<90% COMPLETE → RESUME, do NOT REGISTER (and DE-REGISTER if it slipped through)

An eval whose VALID-trial count is < ~90% of planned n_rep × benchmark_size (swebench-100=300, tb2=267, dev_set_v2=300) is INCOMPLETE → RESUME to completion (Check-4 resume), NOT register. A partial score silently MIS-RANKS the model. If the auto-harvest already registered a partial (Pending/Started, or a prematurely-Finished partial) → DE-REGISTER that ONE row (cross-user FK-safe per §Remediations); it re-registers cleanly on resume. ≥90% / a ~99% tail-short is fine to finalize. Never run §1's register/upload on a <90% eval. Skip a candidate whose slurm job is still RUNNING.

>10% error-fraction HARD GATE — the authoritative metric

Compute via the ONE utility scripts/database/eval_guardrail.py (guardrail_counts(stats, planned).invalid_error_count), the Python port of the leaderboard guardrail (OT-Agent-Leaderboard server/storage.ts:416-449). It EXCLUDES the benign passthrough error types — never hand-restate the BENIGN set (that's how sibling copies drifted). error-fraction = invalid_error_count / total_attempted; a row is an incomplete partial (→ de-register + resume) iff > 10%. There is NO "borderline"/"close enough" discretion: 10.1% fails like 30%. Do NOT re-derive as a naive "no-reward / valid-reward<90%" count — that folds in the benign passthroughs and over-flags near-complete evals. Exception: an explicit, logged user grandfather of a specific already-pushed row.

Run this gate after deduplication. Duplicate attempts inflate the denominator and can falsely turn a broken run into a pass.

0a. Duplicate-contaminated trials — deduplicate before every later step

An eval is duplicate-contaminated when total trials exceed n_rep_eval × benchmark_size or any task has more than n_rep_eval attempts. Do not upload, score, register, resume, or delete the source run before deduplication.

Use the bundled deterministic script. Run its dry-run first. If after = benchmark_size × n_rep_eval, apply and re-run §0. If fewer unique attempts remain, apply the deduplication, then follow check 4 and resume the missing replicas in the same run:

bash
python scripts/harbor/dedupe_harbor_trials.py \
  --job-dir trace_jobs/<RUN_TAG> --n-rep <n_rep_eval>
python scripts/harbor/dedupe_harbor_trials.py \
  --job-dir trace_jobs/<RUN_TAG> --n-rep <n_rep_eval> --apply

Selection is score-blind: globally collapse copied paths with the same Harbor trial_name before grouping by task, then keep the earliest numeric-reward attempts and the earliest no-reward attempts only if needed to reach n_rep_eval. Never choose by reward value. This preserves completed original attempts, lets successful infrastructure retries replace unscored attempts, and cannot cherry-pick passes. The reusable scripts/harbor/dedupe_harbor_trials.py utility removes only surplus trial directories, writes dedupe_manifest.json, and rebuilds result.json plus harbor_result.json.

For object-store runs, stage the exact run prefix locally, deduplicate there, verify the parsed matrix and aggregate, then replace only that run's harbor_trials tree and aggregate files. If result files are staged before large trial artifacts, exclude every manifest directory prefix and file path from the original object listing when hydrating retained artifacts; otherwise a stale pre-delete listing can copy a removed nested result back into the local tree. A second invocation must report before == after and remove nothing.

Check 4 — counting VALID trials (parse, never file-count) — STANDARD report element

Trial dirs at eval_jobs/eval-<safe_model>_<safe_dataset>/<task>__<id>/result.json (depth-1; the run-dir-root result.json is the aggregate — exclude via the */ glob). A trial is VALID iff reward is finite; else ERRORED. One pass over all evals (otagent env, read-only, ~1 min):

python
import glob, json
EJ='/e/data1/datasets/playground/ot-baf/eval_jobs'
EVALS=[('eval-laion_<model>_<safe_dataset>', 300), ...]  # (run-tag, n_rep*bench_size) per eval
for tag, exp in EVALS:
    valid=err=0
    for f in glob.glob(f'{EJ}/{tag}/*/result.json'):          # */ = per-trial only, skips root aggregate
        try: d=json.load(open(f))
        except Exception: err+=1; continue
        r=((d.get('verifier_result') or {}).get('rewards') or {}).get('reward')
        if isinstance(r,(int,float)): valid+=1   # numeric reward = VALID (incl. AgentTimeout passthrough)
        else: err+=1                             # no reward = ERRORED (hard infra/exception)
    print(f'{tag}: valid={valid} err={err} total={valid+err} / exp {exp}')

Always include the valid/errored/expected matrix in the report. Flag any eval with ERRORED rate ≳10–15% as a re-run candidate.

Check 4 — resuming errored trials (resume_chunked.py)

Resume re-runs only the errored trials of the SAME run dir (keeps valid trials). Standard driver: eval/resume_chunked.py → one unified_eval_listener.py --resume-only --force-reeval per chunk. Restate the same sizing/yaml/conda-env/Pinggy from the original fire or Harbor errors on the config.json conflict.

  • --force-reeval (NOT --force-eval) — resume-mode flag; distinct from launch-time --force-eval (which bypasses dedup in should_start_job). Bypasses the DB status check AND the active_pairs resume filter (needed when an active job for the same pair uses a different scaffold).
  • --once + --batch-size — in --once mode, freshly-submitted JIDs fold into active_ids as they submit, so the sliding-window --batch-size chain takes effect.

Cat-3 swe-agent (installed-harness) resume — MUST start a Pinggy tunnel. Cat-3 = preferred-harness reproduction (dcagent_eval_config_swe_agent.yaml, swe-agent/openhands/mini-swe-agent) where the sandboxed agent calls back to the served model over a public Pinggy tunnel. The resume branch of eval/jupiter/eval_harbor.sbatch starts the Pinggy SSH tunnel (gated on EVAL_PINGGY_URL) and exports OPENAI_API_BASE/OPENAI_BASE_URL (+ openhands LLM_BASE_URL/LLM_API_KEY, mini-swe-agent MSWEA_*/HOSTED_VLLM_API_BASE). A Cat-3 resume MUST pass the four passthroughs (note hyphens — distinct from launch-mode --pinggy_persistent_url/--pinggy_token):

bash
# otagent env, in tmux. One Pinggy pair PER invocation (one chunk per free pair for multi-model batches).
python eval/resume_chunked.py \
  --csv /tmp/resume_cands.csv --preset swebench --org eval \
  --tp-size 2 --dp-size 2 --timeout-multiplier 16.0 \
  --jobs-dir <EVAL_JOBS_DIR> --conda-env <env> \
  --tag-prefix <orig_run_tag_prefix> \
  --pinggy-url <free-pair URL, e.g. dadccqeqqf.a.pinggy.link> \
  --pinggy-token <free-pair token> \
  --config-yaml dcagent_eval_config_swe_agent.yaml \
  --agent-parser '' \
  --chunk-size 4 --sleep-between 120
  • --config-yaml = Harbor config (scaffold/parser); --agent-parser '' = disable parser for swe-agent.
  • Pinggy pair URL+token privileged — read a FREE pair from .agents/secret.md/pinggy_bank.md.
  • terminus-2 resumes leave EVAL_PINGGY_URL empty → tunnel block is a no-op (they use the normal sbatch path); only Cat-3 installed-harness resumes set the Pinggy flags.

swe-agent retry-policy (dcagent_eval_config_swe_agent.yaml, aligned to M1 SERA-32B 48.67%): ContextLengthExceededError/BadRequestError/AgentEnvironmentTimeoutError/SummarizationTimeout are RETRIED; VerifierOutputParseError/RewardFileEmptyError/RewardFileNotFoundError are terminal. A resume scored under this policy is NOT comparable to one scored under the old drifted config.

Offline-first pre-download: sbatch tries snapshot_download(local_files_only=True) first (HF cache may be read-only when HEAD advanced past the cached snapshot), falls back to online. Pinggy SSH-out is prefixed with proxychains to egress from no-internet nodes. (etag_timeout=120, 600s hard cap on the pre-download wedge.)


Remediations — the §0 audit scopes which to run (re-run §0 after to confirm ✅)

HF trace upload + Supabase registration are normal, sanctioned operations — run the ones the audit flags.

Cross-user FK safety — REQUIRED before EVERY write (update / delete)

The same (model × benchmark) pair can have MULTIPLE sandbox_jobs rows (reruns, resumes, sibling model_ids, AND other users' rows). Before ANY mutation (sandbox_jobs/sandbox_trial_model_usage/models/benchmarks): query ALL rows for the pair, disambiguate by created_at + username, assert the target row is yours (username == me; your usernames: feuer1/bfeuer00/penfever/benjaminfeuer). Scope every update/delete to id+username. NEVER mutate/delete another user's row. When you can't tell which row is yours → STOP and surface, don't guess. (An INSERT of a new own row never touches theirs — see §1.)

Show full SKILL.md (896 more words)Show less

1. Manual upload + DB register — manual_db_eval_push.py

Pass the trace_jobs/<RUN_TAG> path (where Harbor writes <task>__<id> trial dirs), NOT eval_jobs/<RUN_TAG> (that only has meta.env). Auto-resolves nested trial subdirs and auto-detects agent/model/benchmark.

bash
set -a; source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"; set +a   # SUPABASE_URL, SUPABASE_ANON_KEY, HF_TOKEN
cd <OpenThoughts-Agent>
python scripts/database/manual_db_eval_push.py --job-dir trace_jobs/<RUN_TAG> --verbose
#   --hf-repo DCAgent2/<RUN_TAG>-traces   # explicit HF repo
#   --skip-hf                              # DB only (traces already uploaded)
#   --forced-update                        # overwrite existing records
#   --benchmark-name terminal_bench_2      # EXPLICIT benchmark — PREVENTS the §2b hash bug. Pass when
#                                          # benchmark AND model both contain underscores (e.g.
#                                          # terminal_bench_2_Qwen3_5_35B_A3B_…) — derive_benchmark_from_job_dir
#                                          # can't find the split → mints a HASH benchmark.
  • A cross-user-owned (model × benchmark) pair does NOT block registration. When a gate-passing re-fire's pair is already owned by another user, REGISTER A NEW penfever row alongside it with the corrected score — don't stop, don't ask. The guardrail forbids DELETING/MUTATING a cross-user row; an INSERT never touches theirs (multiple rows per model×benchmark are allowed).
HF trace dataset — use the memory-efficient uploader

Use the streamed, last-episode-per-trial uploader (canonical invocation in rl-agentic-job-cleanup §8) — naive per-conversation extraction loads every episode into RAM (I/O-heavy on GPFS at 300 trials × hundreds of episodes):

bash
# otagent env; ALWAYS in tmux; `hf upload`, NEVER `hf upload-large-folder` (LFS-429 deadlocks)
python -m scripts.harbor.make_and_upload_trace_dataset \
  --job-dir trace_jobs/<RUN_TAG> --repo_id <org>/<RUN_TAG>-traces --episodes last

--episodes last keeps only the scored episode per trial. Eval-vs-RL: RL passes --skip_register; for EVALS register/link the trace repo onto the eval's existing sandbox_jobs row (§2/§3 — set trace/url only, do NOT create a second row or touch the score). On Leonardo, login-node hf upload is SIGKILLed at ~100s → use the sbatch+tunnel pattern (ops/leonardo/ops.md).

2. ⚠️ Verify + fix the MODEL name (with cross-user FK safety)

Auto-detect reads agent_info.model_info.name. For vLLM-served models that field is the vLLM served-model name (numeric, e.g. 1774950145766573), NOT the HF repo → can register a bogus numeric models row. Get the real name:

bash
python3 -c "import json; d=json.load(open('experiments/<RUN_TAG>/configs/<RUN_TAG>_eval_config.json')); print(d['model_hf_name'])"

Then check + repoint (FK-safe; the assert enforces it; models DELETE only if no other-user row FKs it):

python
c.table("sandbox_jobs").select("model_id,username").eq("id", "<JOB_ID>").execute()
c.table("models").select("name").eq("id", "<MODEL_ID>").execute()
import os; me = os.environ.get("USER")
job = c.table("sandbox_jobs").select("id,username,model_id").eq("id", "<JOB_ID>").execute().data[0]
assert job["username"] == me, f"FK-SAFETY STOP: job owned by {job['username']}, not {me} — do not mutate"
correct = c.table("models").select("id").eq("name", "laion/<real-model-name>").execute()
c.table("sandbox_jobs").update({"model_id": correct.data[0]["id"]}).eq("id", "<JOB_ID>").eq("username", me).execute()
c.table("sandbox_trial_model_usage").update({"model_id": correct.data[0]["id"]}).eq("model_id", "<BOGUS_ID>").execute()  # only if all FK'd rows are yours
c.table("models").delete().eq("id", "<BOGUS_ID>").execute()  # only if no other-user row FKs it

2b. ⚠️ ALSO verify + fix the BENCHMARK name (the version-hash mis-registration)

Same failure mode as the model FK. The harness can register the row under a raw 40-hex dataset version-hash (0c553bd6d05d…/377118ff…/693231ec…) instead of the benchmark name — bites when the launch passed a local hash-named snapshot dir as the dataset. A hash-named benchmark silently MIS-FILES the score and spawns orphan rows. (Fixed for new launches; old legs, the auto-harvest path, and manual registers off such a run dir can still land under a hash.)

Detect (in §0 check-2/3 AND after any register): if the row's benchmark NAME matches ^[0-9a-f]{32,64}$ → mis-keyed. Resolve from the run-dir name (<benchmark>_<model>_<ts> — the leading segment IS the benchmark) or trace repo (DCAgent2/<benchmark>_…), then map: swebench-verified-random-100-folders=cc1aca76… · dev_set_v2=b94dfab2… · terminal_bench_2=34ab93c4…. Repoint FK-safe:

python
me = ...  # your username ∈ feuer1/bfeuer00/penfever/benjaminfeuer
job = c.table("sandbox_jobs").select("id,username,benchmark_id").eq("id","<JOB_ID>").execute().data[0]
assert job["username"] == me, "FK-SAFETY STOP — not your row; surface, do not mutate"
proper = c.table("benchmarks").select("id").eq("name","terminal_bench_2").execute().data[0]["id"]
c.table("sandbox_jobs").update({"benchmark_id": proper}).eq("id","<JOB_ID>").eq("username",me).execute()

Clean the orphan hash benchmark — ONLY if 0 sandbox_jobs reference it across ALL users (cross-user pre-check): delete sandbox_benchmark_tasks CHILDREN first (the FK that blocks the delete; that table has no id column — delete by .eq("benchmark_id", <id>)), then delete the benchmark. LEAVE it if ANY row still references it.

3. Verify it landed

Confirm sandbox_jobs.<JOB_ID>.model_id → real laion/<name> row and trial scores attached. If --skip-hf wasn't used, confirm the trace dataset (DCAgent2/<RUN_TAG>-traces) is non-empty on HF.

4. Clean up disk (only after upload + register verified)

Remove the local eval run dir once traces are on HF + the DB row is correct (the trace_jobs tree is the bulk). Detach a large GPFS rm (nohup/tmux); never du/find to size it first. Do NOT delete shared canonical task dirs.


5. IRIS TPU evals — the idempotent path (GCS-backed, no SLURM)

Iris (marin v6e TPU) evals run through eval/cloud/launch_eval_iris.py, NOT SLURM — so the SLURM mechanics (sacct, cluster run-dirs, sbatch resume, GPFS rm) DON'T apply. The audit logic is identical (same 4 checks + >10% HARD gate + <90%→resume-not-register + cross-user FK safety); only the data source + resume/register MECHANICS differ.

  • Enumerate terminal evals (no sacct): the marin jobs table. iris=/Users/benjaminfeuer/Documents/marin/.venv/bin/iris $iris --cluster=marin query "SELECT job_id,state FROM jobs WHERE job_id LIKE '%eval-%' ORDER BY job_id DESC" -f csv State codes: 4=COMPLETED, 5=FAILED, 6=KILLED, 1/2/3=pending/running (SKIP — live). EXEMPT DCAgent2/* measurement runs (§0 pre-gate).
  • Results live in GCS, not a filesystem. The aggregate is <job_output_dir>/<job>/result.json. Resolve the prefix — never guess/scan buckets: OUT=$(python -m hpc.iris.job_output_resolver <job> --cluster …/marin.yaml) (registry-first, iris-fallback). It returns whatever bucket the job wrote to (new single-region gs://marin-us-east5/…, legacy multi-region gs://marin-models-us/…, or the gs://marin-eu-west4 static default). A single gsutil ls "$OUT/<job>/" finds the result.json regardless of region.
  • Audit the GCS result.json — schema is stats.evals.<key> (NOT top-level metrics). Per eval key: score = metrics[0].mean; n_trials/n_total_trials (swe/v2=300, tb2=267); exception_stats[<name>] is a LIST of trial ids → use len() (Σlen over NON-benign names / n_trials = error-fraction; same benign set as §0 check-4). n_cache_tokens=0 = prefix-cache-off fingerprint.
  • Resume (<90% OR >10% err) — no sbatch; relaunch launch_eval_iris.py with the SAME --job_name (Harbor resumes incomplete/errored trials; helpers scripts/iris/check_resume_needed.py + check_progress.py). REGION-CORRECT the relaunch (else it re-lands in eu-west4): run with BOTH export PATH=/Users/benjaminfeuer/Documents/marin/.venv/bin:$PATH (marin iris → region discovery → us-east5) AND the otagent python by FULL PATH /Users/benjaminfeuer/miniconda3/envs/otagent/bin/python (launcher needs omegaconf, which the marin venv lacks). Add --max-retries 2 (a fraction of fresh preemptible v6e-4 slices wedge at model-load model_loader.py:476 ~50min then die). Confirm Region pin: … → us-east5. Iris eval = MAIN Daytona org (DAYTONA_API_KEY), force_build: true (no snapshot pre-build / cap).
  • Register (≥90%, not auto-registered) — Iris evals auto-register via --upload_to_database on completion, so most complete legs ARE registered; for one that isn't, gsutil -m rsync -r gs://<bucket>/ot-agent/<job>/<job>/ <local_tmp>/ then manual_db_eval_push.py --job-dir <local_tmp> (+ §2/§2b FK fixes, own-rows-only). Then remove the local tmp (GCS is the durable store — no cluster-filesystem cleanup).

Sibling cleanups: rl-agentic-job-cleanup (RL model), sft-job-cleanup (SFT model), datagen-job-cleanup (trace dataset). Launching evals → eval-agentic-launch (+ the *-iris variant for TPU). Per-cluster particulars → .agents/ops/<cluster>/.


Operating notes

  • Stale >12h non-Finished (Started/Running/Pending) rows owned by us are cosmetic cruft → DELETE them (FK-cascade, own-rows-only); registration is create-if-missing so this never blocks a real score landing later.
  • notes/ot-agent/task_repos/rl_to_check.md is a QUEUE file (flat URL list, one per line, consumed line-by-line by the smoke-test runner; processed entries → rl_checked.md). "Update rl_to_check.md with the fixes" = append new repo URLs, one per line. Do NOT add markdown tables/sections (breaks the parser). Same caution for all task_repos/*.md.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/eval-agentic-cleanup of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Eval Agentic Cleanup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Agentic Cleanup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Agentic Cleanup this skillopen-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k3 repos~3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Diffusion Perf Optvllm-project/vllm-omni7.1k—~7.5kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Ascend Release Manager for vLLMvllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 3 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Ascend Release Manager for vLLM

    vllm-project/vllm-ascend

    Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

    2.9k GitHub stars~7.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Quantization

    vllm-project/vllm-omni

    Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 9 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed

Works with

Questions about Eval Agentic Cleanup

What does Eval Agentic Cleanup do?

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent. Eval Agentic Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent. Audit + recover a finished agentic eval.

When should I use Eval Agentic Cleanup?

Eval Agentic Cleanup fits situations like: verify an eval is truly complete; A sweep finds an eval that didnt upload/register; has a broken/zero score; short trial count.

How do I install Eval Agentic Cleanup in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a claude-code`. Or copy the skill folder (.agents/skills/eval-agentic-cleanup in open-thoughts/OpenThoughts-Agent) into .claude/skills/eval-agentic-cleanup in your project. Claude Code loads it when a task matches its description.

How do I install Eval Agentic Cleanup in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a codex`. Or copy the skill folder (.agents/skills/eval-agentic-cleanup in open-thoughts/OpenThoughts-Agent) into .agents/skills/eval-agentic-cleanup in your project. Codex loads it when a task matches its description.

Can I use Eval Agentic Cleanup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-agentic-cleanup, .gemini/skills/eval-agentic-cleanup, .github/skills/eval-agentic-cleanup and .opencode/skills/eval-agentic-cleanup in your project.

What does Eval Agentic Cleanup need to run?

Going by SKILL.md and its folder, Eval Agentic Cleanup needs the command-line tools its instructions call (python, gsutil, python3 and hf) and credentials named LLM_API_KEY, DAYTONA_API_KEY, SUPABASE_ANON_KEY and HF_TOKEN. Our summary lists: Python 3; A credential in LLM_API_KEY.

Does Eval Agentic Cleanup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Agentic Cleanup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Agentic Cleanup use?

Eval Agentic Cleanup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Agentic Cleanup use?

About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Agentic Cleanup?

Skills that share tags, products or a category with Eval Agentic Cleanup: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Diffusion Perf Opt (vllm-project/vllm-omni, 7.1k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Agentic Cleanup?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.