Aider Delegate
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-cleanup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-agentic-cleanup .claude/skills/eval-agentic-cleanup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-agentic-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-cleanup into .claude/skills/eval-agentic-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-cleanup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-cleanupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-cleanup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/eval-agentic-cleanup .agents/skills/eval-agentic-cleanup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-agentic-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-cleanup into .agents/skills/eval-agentic-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-cleanup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-cleanup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/eval-agentic-cleanup .cursor/skills/eval-agentic-cleanup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-agentic-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-cleanup into .cursor/skills/eval-agentic-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-cleanup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/eval-agentic-cleanup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-cleanup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/eval-agentic-cleanup .gemini/skills/eval-agentic-cleanup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-agentic-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-cleanup into .gemini/skills/eval-agentic-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-cleanup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-cleanupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/eval-agentic-cleanup .github/skills/eval-agentic-cleanup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-agentic-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-cleanup into .github/skills/eval-agentic-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-cleanup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-cleanup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/eval-agentic-cleanup .opencode/skills/eval-agentic-cleanup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-agentic-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-cleanup into .opencode/skills/eval-agentic-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-cleanup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-agentic-cleanupAudit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.
Eval Agentic Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent. Audit + recover a finished agentic eval. ALWAYS start with the read-only, idempotent completeness/health audit (§0): job finished? score present + non-zero + not obviously broken? HF traces present + linked? trial count ≈ nrep × benchmarksize? — it writes nothing and recommends an action per check. Then run only the flagged remediations: manual HF-trace upload + Supabase DB-registration via manualdbevalpush.py, the vLLM-numeric-ID → real-HF-model-name fix (with cross-user FK safety), verify, free disk. Use to…
Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and Supabase. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythongsutilpython3hfFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
LLM_API_KEYDAYTONA_API_KEYSUPABASE_ANON_KEYHF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Agentic Cleanup loads about 5.8k tokens when it runs. Until then it costs about 216 tokens; SKILL.md has 2,221 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 2,221 words, ~5,821 tokens.
.claude/skills/eval-agentic-cleanup/SKILL.md (or your agent's skills folder).Always start with the read-only audit (§0): it verifies the four conditions that actually matter, then run only the remediations it flags. "The score is in the leaderboard" does NOT mean the eval is complete — the auto-harvest path can land a score while silently skipping trace upload, or record a bogus 0 from an eval that never reached the model.
Reads only (sacct, run-dir result.json, HF API, Supabase reads). Emits a per-eval verdict + recommended action per check; run only the flagged remediations.
Pre-gate — EXEMPT: grid / throughput / OOM-test measurement runs targeting DCAgent2/* → skip (upload only production winners).
| # | Check | How (read-only) | PASS | If it FAILS → recommend |
|---|---|---|---|---|
| 1 | Job finished | sacct -j <id> -X -o State | COMPLETED | RUNNING/PENDING → wait & re-audit; FAILED/TIMEOUT/CANCELLED → diagnose + relaunch (see eval-agentic-launch gotchas #1–#5: jobs_dir perm / hosted_vllm 2-slash / model_info / reservation / stale-row) |
| 2 | Score present & not broken | Supabase sandbox_jobs score, or run-dir result.json → stats.evals.<k>.metrics[0].mean | non-null real number; 0 OK only if check 4 shows trials ran | null/missing/0 with ~0 trials/POSTs = infra-broken (eval never reached the model — classic jobs_dir/hosted_vllm/model_info failure) → re-run. High VerificationNotCompletedError rate → score may be biased (note below). |
| 3 | HF traces present + linked | trace HF repo exists + non-empty (HF API 200) AND DB/LB row has trace/url field set | repo 200 AND linked | missing/empty OR unlinked → upload traces + register link (§1–§3) |
| 4 | VALID trial count (PARSE, don't file-count) | Parse each per-trial result.json for a numeric reward (snippet below). An errored trial STILL writes result.json with exception_info and no reward, so file-count overstates coverage. | VALID is ~90–100% of n_rep_eval × benchmark_size, with at most n_rep_eval retained attempts per task | VALID or total above plan / more than n_rep_eval attempts for a task → deduplicate (§0a), then re-audit. Short VALID (<~90%, a missing rep, or high hard-error 15–25%) → resume errored trials. Expected: swebench-verified-random-100=100, dev_set_v2=100, terminal_bench_2=89; × n_rep_eval (default 3) → ~300/~300/~267. |
Output one row per eval: ✅/⚠️ per check + the single recommended next action. Re-run the audit after any remediation to confirm ✅.
VerificationNotCompletedError (harbor 9203989f) / pre-fix AgentTimeoutError WITH NO reward = the verifier never produced a result. Missing because the agent burned its budget → true score ≈ 0 (SWE-bench ~pass/fail). Harbor's JobStats mean DROPS these from the denominator → biased UP; the DB/leaderboard harvest counts them as 0 → unbiased. Either way retriable: resume on a healthy sandbox — the payoff is usually recovering real passes hidden as 0s.DaytonaAuthenticationError/DaytonaValidationError/etc.) = failure uncorrelated with solvability → dropping is ~unbiased (lowers N). Re-run for completeness only; do NOT impute 0.AgentTimeoutError WITH a reward = verifier scored it → passthrough VALID 0. Don't filter/re-run.An eval whose VALID-trial count is < ~90% of planned n_rep × benchmark_size (swebench-100=300, tb2=267, dev_set_v2=300) is INCOMPLETE → RESUME to completion (Check-4 resume), NOT register. A partial score silently MIS-RANKS the model. If the auto-harvest already registered a partial (Pending/Started, or a prematurely-Finished partial) → DE-REGISTER that ONE row (cross-user FK-safe per §Remediations); it re-registers cleanly on resume. ≥90% / a ~99% tail-short is fine to finalize. Never run §1's register/upload on a <90% eval. Skip a candidate whose slurm job is still RUNNING.
Compute via the ONE utility scripts/database/eval_guardrail.py (guardrail_counts(stats, planned).invalid_error_count), the Python port of the leaderboard guardrail (OT-Agent-Leaderboard server/storage.ts:416-449). It EXCLUDES the benign passthrough error types — never hand-restate the BENIGN set (that's how sibling copies drifted). error-fraction = invalid_error_count / total_attempted; a row is an incomplete partial (→ de-register + resume) iff > 10%. There is NO "borderline"/"close enough" discretion: 10.1% fails like 30%. Do NOT re-derive as a naive "no-reward / valid-reward<90%" count — that folds in the benign passthroughs and over-flags near-complete evals. Exception: an explicit, logged user grandfather of a specific already-pushed row.
Run this gate after deduplication. Duplicate attempts inflate the denominator and can falsely turn a broken run into a pass.
An eval is duplicate-contaminated when total trials exceed n_rep_eval × benchmark_size or any task has more than n_rep_eval attempts. Do not upload, score, register, resume, or delete the source run before deduplication.
Use the bundled deterministic script. Run its dry-run first. If after = benchmark_size × n_rep_eval, apply and re-run §0. If fewer unique attempts remain, apply the deduplication, then follow check 4 and resume the missing replicas in the same run:
python scripts/harbor/dedupe_harbor_trials.py \
--job-dir trace_jobs/<RUN_TAG> --n-rep <n_rep_eval>
python scripts/harbor/dedupe_harbor_trials.py \
--job-dir trace_jobs/<RUN_TAG> --n-rep <n_rep_eval> --applySelection is score-blind: globally collapse copied paths with the same Harbor trial_name before grouping by task, then keep the earliest numeric-reward attempts and the earliest no-reward attempts only if needed to reach n_rep_eval. Never choose by reward value. This preserves completed original attempts, lets successful infrastructure retries replace unscored attempts, and cannot cherry-pick passes. The reusable scripts/harbor/dedupe_harbor_trials.py utility removes only surplus trial directories, writes dedupe_manifest.json, and rebuilds result.json plus harbor_result.json.
For object-store runs, stage the exact run prefix locally, deduplicate there, verify the parsed matrix and aggregate, then replace only that run's harbor_trials tree and aggregate files. If result files are staged before large trial artifacts, exclude every manifest directory prefix and file path from the original object listing when hydrating retained artifacts; otherwise a stale pre-delete listing can copy a removed nested result back into the local tree. A second invocation must report before == after and remove nothing.
Trial dirs at eval_jobs/eval-<safe_model>_<safe_dataset>/<task>__<id>/result.json (depth-1; the run-dir-root result.json is the aggregate — exclude via the */ glob). A trial is VALID iff reward is finite; else ERRORED. One pass over all evals (otagent env, read-only, ~1 min):
import glob, json
EJ='/e/data1/datasets/playground/ot-baf/eval_jobs'
EVALS=[('eval-laion_<model>_<safe_dataset>', 300), ...] # (run-tag, n_rep*bench_size) per eval
for tag, exp in EVALS:
valid=err=0
for f in glob.glob(f'{EJ}/{tag}/*/result.json'): # */ = per-trial only, skips root aggregate
try: d=json.load(open(f))
except Exception: err+=1; continue
r=((d.get('verifier_result') or {}).get('rewards') or {}).get('reward')
if isinstance(r,(int,float)): valid+=1 # numeric reward = VALID (incl. AgentTimeout passthrough)
else: err+=1 # no reward = ERRORED (hard infra/exception)
print(f'{tag}: valid={valid} err={err} total={valid+err} / exp {exp}')Always include the valid/errored/expected matrix in the report. Flag any eval with ERRORED rate ≳10–15% as a re-run candidate.
resume_chunked.py)Resume re-runs only the errored trials of the SAME run dir (keeps valid trials). Standard driver: eval/resume_chunked.py → one unified_eval_listener.py --resume-only --force-reeval per chunk. Restate the same sizing/yaml/conda-env/Pinggy from the original fire or Harbor errors on the config.json conflict.
--force-reeval (NOT --force-eval) — resume-mode flag; distinct from launch-time --force-eval (which bypasses dedup in should_start_job). Bypasses the DB status check AND the active_pairs resume filter (needed when an active job for the same pair uses a different scaffold).--once + --batch-size — in --once mode, freshly-submitted JIDs fold into active_ids as they submit, so the sliding-window --batch-size chain takes effect.Cat-3 swe-agent (installed-harness) resume — MUST start a Pinggy tunnel. Cat-3 = preferred-harness reproduction (dcagent_eval_config_swe_agent.yaml, swe-agent/openhands/mini-swe-agent) where the sandboxed agent calls back to the served model over a public Pinggy tunnel. The resume branch of eval/jupiter/eval_harbor.sbatch starts the Pinggy SSH tunnel (gated on EVAL_PINGGY_URL) and exports OPENAI_API_BASE/OPENAI_BASE_URL (+ openhands LLM_BASE_URL/LLM_API_KEY, mini-swe-agent MSWEA_*/HOSTED_VLLM_API_BASE). A Cat-3 resume MUST pass the four passthroughs (note hyphens — distinct from launch-mode --pinggy_persistent_url/--pinggy_token):
# otagent env, in tmux. One Pinggy pair PER invocation (one chunk per free pair for multi-model batches).
python eval/resume_chunked.py \
--csv /tmp/resume_cands.csv --preset swebench --org eval \
--tp-size 2 --dp-size 2 --timeout-multiplier 16.0 \
--jobs-dir <EVAL_JOBS_DIR> --conda-env <env> \
--tag-prefix <orig_run_tag_prefix> \
--pinggy-url <free-pair URL, e.g. dadccqeqqf.a.pinggy.link> \
--pinggy-token <free-pair token> \
--config-yaml dcagent_eval_config_swe_agent.yaml \
--agent-parser '' \
--chunk-size 4 --sleep-between 120--config-yaml = Harbor config (scaffold/parser); --agent-parser '' = disable parser for swe-agent..agents/secret.md/pinggy_bank.md.EVAL_PINGGY_URL empty → tunnel block is a no-op (they use the normal sbatch path); only Cat-3 installed-harness resumes set the Pinggy flags.swe-agent retry-policy (
dcagent_eval_config_swe_agent.yaml, aligned to M1 SERA-32B 48.67%):ContextLengthExceededError/BadRequestError/AgentEnvironmentTimeoutError/SummarizationTimeoutare RETRIED;VerifierOutputParseError/RewardFileEmptyError/RewardFileNotFoundErrorare terminal. A resume scored under this policy is NOT comparable to one scored under the old drifted config.
Offline-first pre-download: sbatch tries
snapshot_download(local_files_only=True)first (HF cache may be read-only when HEAD advanced past the cached snapshot), falls back to online. Pinggy SSH-out is prefixed with proxychains to egress from no-internet nodes. (etag_timeout=120, 600s hard cap on the pre-download wedge.)
HF trace upload + Supabase registration are normal, sanctioned operations — run the ones the audit flags.
The same (model × benchmark) pair can have MULTIPLE sandbox_jobs rows (reruns, resumes, sibling model_ids, AND other users' rows). Before ANY mutation (sandbox_jobs/sandbox_trial_model_usage/models/benchmarks): query ALL rows for the pair, disambiguate by created_at + username, assert the target row is yours (username == me; your usernames: feuer1/bfeuer00/penfever/benjaminfeuer). Scope every update/delete to id+username. NEVER mutate/delete another user's row. When you can't tell which row is yours → STOP and surface, don't guess. (An INSERT of a new own row never touches theirs — see §1.)
manual_db_eval_push.pyPass the trace_jobs/<RUN_TAG> path (where Harbor writes <task>__<id> trial dirs), NOT eval_jobs/<RUN_TAG> (that only has meta.env). Auto-resolves nested trial subdirs and auto-detects agent/model/benchmark.
set -a; source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV to the secrets file first}"; set +a # SUPABASE_URL, SUPABASE_ANON_KEY, HF_TOKEN
cd <OpenThoughts-Agent>
python scripts/database/manual_db_eval_push.py --job-dir trace_jobs/<RUN_TAG> --verbose
# --hf-repo DCAgent2/<RUN_TAG>-traces # explicit HF repo
# --skip-hf # DB only (traces already uploaded)
# --forced-update # overwrite existing records
# --benchmark-name terminal_bench_2 # EXPLICIT benchmark — PREVENTS the §2b hash bug. Pass when
# # benchmark AND model both contain underscores (e.g.
# # terminal_bench_2_Qwen3_5_35B_A3B_…) — derive_benchmark_from_job_dir
# # can't find the split → mints a HASH benchmark.penfever row alongside it with the corrected score — don't stop, don't ask. The guardrail forbids DELETING/MUTATING a cross-user row; an INSERT never touches theirs (multiple rows per model×benchmark are allowed).Use the streamed, last-episode-per-trial uploader (canonical invocation in rl-agentic-job-cleanup §8) — naive per-conversation extraction loads every episode into RAM (I/O-heavy on GPFS at 300 trials × hundreds of episodes):
# otagent env; ALWAYS in tmux; `hf upload`, NEVER `hf upload-large-folder` (LFS-429 deadlocks)
python -m scripts.harbor.make_and_upload_trace_dataset \
--job-dir trace_jobs/<RUN_TAG> --repo_id <org>/<RUN_TAG>-traces --episodes last--episodes last keeps only the scored episode per trial. Eval-vs-RL: RL passes --skip_register; for EVALS register/link the trace repo onto the eval's existing sandbox_jobs row (§2/§3 — set trace/url only, do NOT create a second row or touch the score). On Leonardo, login-node hf upload is SIGKILLed at ~100s → use the sbatch+tunnel pattern (ops/leonardo/ops.md).
Auto-detect reads agent_info.model_info.name. For vLLM-served models that field is the vLLM served-model name (numeric, e.g. 1774950145766573), NOT the HF repo → can register a bogus numeric models row. Get the real name:
python3 -c "import json; d=json.load(open('experiments/<RUN_TAG>/configs/<RUN_TAG>_eval_config.json')); print(d['model_hf_name'])"Then check + repoint (FK-safe; the assert enforces it; models DELETE only if no other-user row FKs it):
c.table("sandbox_jobs").select("model_id,username").eq("id", "<JOB_ID>").execute()
c.table("models").select("name").eq("id", "<MODEL_ID>").execute()
import os; me = os.environ.get("USER")
job = c.table("sandbox_jobs").select("id,username,model_id").eq("id", "<JOB_ID>").execute().data[0]
assert job["username"] == me, f"FK-SAFETY STOP: job owned by {job['username']}, not {me} — do not mutate"
correct = c.table("models").select("id").eq("name", "laion/<real-model-name>").execute()
c.table("sandbox_jobs").update({"model_id": correct.data[0]["id"]}).eq("id", "<JOB_ID>").eq("username", me).execute()
c.table("sandbox_trial_model_usage").update({"model_id": correct.data[0]["id"]}).eq("model_id", "<BOGUS_ID>").execute() # only if all FK'd rows are yours
c.table("models").delete().eq("id", "<BOGUS_ID>").execute() # only if no other-user row FKs itSame failure mode as the model FK. The harness can register the row under a raw 40-hex dataset version-hash (0c553bd6d05d…/377118ff…/693231ec…) instead of the benchmark name — bites when the launch passed a local hash-named snapshot dir as the dataset. A hash-named benchmark silently MIS-FILES the score and spawns orphan rows. (Fixed for new launches; old legs, the auto-harvest path, and manual registers off such a run dir can still land under a hash.)
Detect (in §0 check-2/3 AND after any register): if the row's benchmark NAME matches ^[0-9a-f]{32,64}$ → mis-keyed.
Resolve from the run-dir name (<benchmark>_<model>_<ts> — the leading segment IS the benchmark) or trace repo (DCAgent2/<benchmark>_…), then map: swebench-verified-random-100-folders=cc1aca76… · dev_set_v2=b94dfab2… · terminal_bench_2=34ab93c4….
Repoint FK-safe:
me = ... # your username ∈ feuer1/bfeuer00/penfever/benjaminfeuer
job = c.table("sandbox_jobs").select("id,username,benchmark_id").eq("id","<JOB_ID>").execute().data[0]
assert job["username"] == me, "FK-SAFETY STOP — not your row; surface, do not mutate"
proper = c.table("benchmarks").select("id").eq("name","terminal_bench_2").execute().data[0]["id"]
c.table("sandbox_jobs").update({"benchmark_id": proper}).eq("id","<JOB_ID>").eq("username",me).execute()Clean the orphan hash benchmark — ONLY if 0 sandbox_jobs reference it across ALL users (cross-user pre-check): delete sandbox_benchmark_tasks CHILDREN first (the FK that blocks the delete; that table has no id column — delete by .eq("benchmark_id", <id>)), then delete the benchmark. LEAVE it if ANY row still references it.
Confirm sandbox_jobs.<JOB_ID>.model_id → real laion/<name> row and trial scores attached. If --skip-hf wasn't used, confirm the trace dataset (DCAgent2/<RUN_TAG>-traces) is non-empty on HF.
Remove the local eval run dir once traces are on HF + the DB row is correct (the trace_jobs tree is the bulk). Detach a large GPFS rm (nohup/tmux); never du/find to size it first. Do NOT delete shared canonical task dirs.
Iris (marin v6e TPU) evals run through eval/cloud/launch_eval_iris.py, NOT SLURM — so the SLURM mechanics (sacct, cluster run-dirs, sbatch resume, GPFS rm) DON'T apply. The audit logic is identical (same 4 checks + >10% HARD gate + <90%→resume-not-register + cross-user FK safety); only the data source + resume/register MECHANICS differ.
iris=/Users/benjaminfeuer/Documents/marin/.venv/bin/iris
$iris --cluster=marin query "SELECT job_id,state FROM jobs WHERE job_id LIKE '%eval-%' ORDER BY job_id DESC" -f csv
State codes: 4=COMPLETED, 5=FAILED, 6=KILLED, 1/2/3=pending/running (SKIP — live). EXEMPT DCAgent2/* measurement runs (§0 pre-gate).<job_output_dir>/<job>/result.json. Resolve the prefix — never guess/scan buckets: OUT=$(python -m hpc.iris.job_output_resolver <job> --cluster …/marin.yaml) (registry-first, iris-fallback). It returns whatever bucket the job wrote to (new single-region gs://marin-us-east5/…, legacy multi-region gs://marin-models-us/…, or the gs://marin-eu-west4 static default). A single gsutil ls "$OUT/<job>/" finds the result.json regardless of region.result.json — schema is stats.evals.<key> (NOT top-level metrics). Per eval key: score = metrics[0].mean; n_trials/n_total_trials (swe/v2=300, tb2=267); exception_stats[<name>] is a LIST of trial ids → use len() (Σlen over NON-benign names / n_trials = error-fraction; same benign set as §0 check-4). n_cache_tokens=0 = prefix-cache-off fingerprint.launch_eval_iris.py with the SAME --job_name (Harbor resumes incomplete/errored trials; helpers scripts/iris/check_resume_needed.py + check_progress.py). REGION-CORRECT the relaunch (else it re-lands in eu-west4): run with BOTH export PATH=/Users/benjaminfeuer/Documents/marin/.venv/bin:$PATH (marin iris → region discovery → us-east5) AND the otagent python by FULL PATH /Users/benjaminfeuer/miniconda3/envs/otagent/bin/python (launcher needs omegaconf, which the marin venv lacks). Add --max-retries 2 (a fraction of fresh preemptible v6e-4 slices wedge at model-load model_loader.py:476 ~50min then die). Confirm Region pin: … → us-east5. Iris eval = MAIN Daytona org (DAYTONA_API_KEY), force_build: true (no snapshot pre-build / cap).--upload_to_database on completion, so most complete legs ARE registered; for one that isn't, gsutil -m rsync -r gs://<bucket>/ot-agent/<job>/<job>/ <local_tmp>/ then manual_db_eval_push.py --job-dir <local_tmp> (+ §2/§2b FK fixes, own-rows-only). Then remove the local tmp (GCS is the durable store — no cluster-filesystem cleanup).Sibling cleanups: rl-agentic-job-cleanup (RL model), sft-job-cleanup (SFT model), datagen-job-cleanup (trace dataset). Launching evals → eval-agentic-launch (+ the *-iris variant for TPU). Per-cluster particulars → .agents/ops/<cluster>/.
notes/ot-agent/task_repos/rl_to_check.md is a QUEUE file (flat URL list, one per line, consumed line-by-line by the smoke-test runner; processed entries → rl_checked.md). "Update rl_to_check.md with the fixes" = append new repo URLs, one per line. Do NOT add markdown tables/sections (breaks the parser). Same caution for all task_repos/*.md.© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/eval-agentic-cleanup of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Eval Agentic Cleanup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Agentic Cleanup this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | |
| Aider DelegateamElnagdy/delegate-skills | 2.3k | 3 repos | ~3k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Diffusion Perf Optvllm-project/vllm-omni | 7.1k | — | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Ascend Release Manager for vLLMvllm-project/vllm-ascend | 2.9k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 |
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
vllm-project/vllm-ascend
Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.
vllm-project/vllm-omni
Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
Categories
Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent. Eval Agentic Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent. Audit + recover a finished agentic eval.
Eval Agentic Cleanup fits situations like: verify an eval is truly complete; A sweep finds an eval that didnt upload/register; has a broken/zero score; short trial count.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a claude-code`. Or copy the skill folder (.agents/skills/eval-agentic-cleanup in open-thoughts/OpenThoughts-Agent) into .claude/skills/eval-agentic-cleanup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a codex`. Or copy the skill folder (.agents/skills/eval-agentic-cleanup in open-thoughts/OpenThoughts-Agent) into .agents/skills/eval-agentic-cleanup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-cleanup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-agentic-cleanup, .gemini/skills/eval-agentic-cleanup, .github/skills/eval-agentic-cleanup and .opencode/skills/eval-agentic-cleanup in your project.
Going by SKILL.md and its folder, Eval Agentic Cleanup needs the command-line tools its instructions call (python, gsutil, python3 and hf) and credentials named LLM_API_KEY, DAYTONA_API_KEY, SUPABASE_ANON_KEY and HF_TOKEN. Our summary lists: Python 3; A credential in LLM_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Agentic Cleanup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Agentic Cleanup: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Diffusion Perf Opt (vllm-project/vllm-omni, 7.1k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.