Agent skill

Eval Standard Cleanup

by open-thoughts in open-thoughts/OpenThoughts-Agent

Consolidate FINISHED standard / lmeval (evalchemy) math-suite eval jobs — the Delphi 6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker.

Apache-2.0Auto-check passed

Install Eval Standard Cleanup

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-cleanup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent eval-standard-cleanup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-standard-cleanup .claude/skills/eval-standard-cleanup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-standard-cleanup
GitHub stars
301
Token cost
~1.7k tokens
SKILL.md length
684 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Consolidate FINISHED standard / lmeval (evalchemy) math-suite eval jobs — the Delphi 6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker.

  • Works in 7 steps: Reference files (source of truth —… → Per finished job — verify a real,… → Consolidate — rsync results into the… → …
  • Asked to consolidate / harvest finished standard math-eval jobs
  • SKILL.md covers 0. Reference files (source of…, 1. Per finished job — verify a…, 2. Consolidate — rsync results… and 3. Parse the scalars (the…, plus 3 more sections
  • Calls rsync

What it does

Eval Standard Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent. Consolidate FINISHED standard / lmeval (evalchemy) math-suite eval jobs — the Delphi 6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker. Per job: confirm a NON-EMPTY seed42 result (a crash leaves it empty — don't record it as a score), rsync the results.json into the local per-model dir, write scalar partials, parse the scalars (MATH-500 acc×100 / AIME24 10-seed mean±sd / gsm8k strict+flex / Raw), and flip the SCORES.md row 🚀 eval submitted → ✅ done. The artifact…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with Supabase. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asked to consolidate / harvest finished standard math-eval jobs
  • Fill the scaling-laws score grid

Example prompts

  • “/eval-standard-cleanup”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Reference files (source of truth — derive the flow from these)
  2. Per finished job — verify a real, non-empty result (do NOT record a crash as a score)
  3. Consolidate — rsync results into the local per-model dir + write scalar partials
  4. Parse the scalars (the §5.2-E method — round to 1 decimal to match the table)
  5. Update the SCORES.md row — 🚀 eval submitted → ✅ done
  6. HF-upload only — NEVER DB
  7. Disk / hygiene

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rsync

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use rsync, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Standard Cleanup loads about 1.7k tokens when it runs. Until then it costs about 196 tokens; SKILL.md has 684 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~196
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 684 words, ~1,677 tokens.

Download SKILL.mdSave it as .claude/skills/eval-standard-cleanup/SKILL.md (or your agent's skills folder).
name
eval-standard-cleanup
description
Consolidate FINISHED standard / lm_eval (evalchemy) math-suite eval jobs — the Delphi #6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker. Per job: confirm a NON-EMPTY seed42 result (a crash leaves it empty — don't record it as a score), rsync the results_*.json into the local per-model dir, write scalar partials, parse the scalars (MATH-500 acc×100 / AIME24 10-seed mean±sd / gsm8k strict+flex / Raw), and flip the SCORES.md row 🚀 eval submitted → ✅ done. The artifact is SCALAR SCORES IN A TRACKER — HF-upload-only, NEVER DB. Use when asked to consolidate / harvest finished standard math-eval jobs or fill the scaling-laws score grid. DISTINCT from eval-agentic-cleanup (the Harbor-trace + Supabase DB path).

eval-standard-cleanup

Post-run consolidation for STANDARD evals (lm_eval / evalchemy): the Delphi #6279 RL-scaling-laws math suite — MATH-500 (1 seed) + AIME24 (10-seed mean±sd) + gsm8k (strict+flex) — launched by the eval-standard-launch skill on Leonardo. The deliverable is scalar scores in a tracker: no HF trace upload, no model upload, NEVER a DB registration (see §5). For the agentic Harbor-trace + Supabase path, use eval-agentic-cleanup instead.

0. Reference files (source of truth — derive the flow from these)

Local notes dir: /Users/benjaminfeuer/Documents/experiments/active/delphi/rl-scaling-laws-6279/

  • eval-standard-launch skill — the launch + delphi_eval.sbatch + the per-model output layout it produces (…/experiments/delphi-eval/<RUN_NAME>/seed{42..51}/). Read it for what's already on the cluster.
  • main_sft_evals/SCORES.md — the master tracker for the 54-run main grid; its header describes the consolidation flow (per-model dirs main_sft_evals/<basename>/, partials main_sft_evals/.partial/).
  • EVAL_CONVENTION.md §3.4 / §5.2-E — the per-cluster naming, the idempotent-skip seed42 dir, the exact scalar-parse method, and the "COMPLETED is NOT proof — verify numeric scores" rule.
  • The earlier cold-start grid is the same flow but a different tracker (eval/SCORES.md, artifacts eval/<RUN>/) — don't cross-file; the 54-run main grid lands in main_sft_evals/.

1. Per finished job — verify a real, non-empty result (do NOT record a crash as a score)

A SLURM COMPLETED state is necessary but not sufficient: evalchemy catches engine errors, exits 0, and writes an empty results: {} JSON (silent drop). For each job:

  1. Confirm terminal state via sacct -j <jobid> --format=State -n -P | head -1 (the source of truth, not log-tailing).
  2. Confirm the job produced a NON-EMPTY seed42 result — the §3.4 idempotent-skip dir (…/delphi-eval/<RUN_NAME>/seed42/) must exist and its results_*.json must carry numeric scores (a crashed/timed-out run leaves an empty or missing seed42). An empty/missing seed42 → leave the row pending and note it; do not fabricate a score. (If gsm8k failed after MATH-500 wrote seed42, seed42 exists but is partial — note it; deleting seed42 + re-running is a launcher action, not a cleanup one.)

2. Consolidate — rsync results into the local per-model dir + write scalar partials

Pull only the per-task results_*.json (the examples payloads live inside the JSON, ~1-2 MB each; skip the multi-MB trace subdirs). For the main grid the destination is main_sft_evals/<basename>/:

bash
RUN=delphi-9e19-p33m67-k0p20-lr83-a002-magpie_lr1e5-sft   # = the SFT model basename, maps 1:1 to the row
DEST=/Users/benjaminfeuer/Documents/experiments/active/delphi/rl-scaling-laws-6279/main_sft_evals/$RUN
mkdir -p "$DEST"
rsync -avz --prune-empty-dirs --include='*/' --include='results_*.json' --exclude='*' \
  Leonardo:/leonardo_work/AIFAC_5C0_290/bfeuer00/experiments/delphi-eval/$RUN/ "$DEST/"

This pulls the MATH-500 seed42 JSON, the gsm8k seed42 JSON (written by the plain lm_eval step), and the AIME24 seed42..seed51 JSONs. Write the extracted scalars to main_sft_evals/.partial/<basename>.json. Keep the rsynced JSONs in <basename>/ as the archive. Parse by loading the JSON and reading only numeric keys — never print the huge examples list.

Show full SKILL.md (296 more words)Show less

3. Parse the scalars (the §5.2-E method — round to 1 decimal to match the table)

  • MATH-500 = results["MATH500"]["accuracy"] × 100 (it's a fraction; 0.018 → 1.8). From seed42.
  • AIME24 = results["AIME24"]["accuracy_avg"] × 100, taken over the 10 seed dirs (seed42..seed51); report mean ± sample-stdev of the per-seed values (×100). (e.g. 0.1±0.1.)
  • gsm8k (from the lm_eval output under <basename>/seed42): exact_match,strict-match → gsm8k-S and exact_match,flexible-extract → gsm8k-F, each ×100.
  • Raw = mean(MATH-500, AIME24, gsm8k-strict).
  • Note any format-fail cell (empty \boxed{} / model_answer:"") — that's a real signal, not an eval bug.

The marin compiler marin:experiments/evals/evalchemy_results_compiler.py documents the exact JSON shape (results[task]…) — reuse its extraction shape rather than rolling a new parse. (verify the precise key names against the rsynced JSON on first use; evalchemy versions have drifted.)

4. Update the SCORES.md row — 🚀 eval submitted → ✅ done

Fill that row's MATH-500 | AIME24 (mean±se) | gsm8k-S | gsm8k-F | Raw cells from the partial and flip status from 🚀 eval submitted to ✅ done (keep the eval job id in its column). Edit the single row (rows are independent lines) and preserve the table format exactly. A crashed/empty-seed42 job stays pending with a note (step 1) — never invent numbers.

5. HF-upload only — NEVER DB

These are scalar scores in a tracker, not a model/trace artifact. Do NOT call manual_db_eval_push.py / touch Supabase — that is the eval-agentic-cleanup (agentic / Harbor-trace) path. Results live in this experiment's docs + (optionally) as a training_logs/-style JSON on the model's HF repo; there is no leaderboard row. (Per project_delphi_sft_hf_only_no_db.)

6. Disk / hygiene

The local artifacts are small (~1-2 MB JSONs); no large delete is needed. On the cluster, leave the delphi-eval/<RUN_NAME>/ dirs in place — they back the §3.4 idempotent skip, so a re-harvest or re-submit stays safe. No find / du / rglob on GPFS; use canonical paths.


Launcher: eval-standard-launch. Model-publishing cleanups: rl-agentic-job-cleanup / sft-job-cleanup. Per-cluster particulars (ssh, the step-ca cert refresh for rsync, the /leonardo_work/AIFAC_5C0_290/bfeuer00/… paths) → .agents/ops/<cluster>/.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/eval-standard-cleanup of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Eval Standard Cleanup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Standard Cleanup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Standard Cleanup this skillopen-thoughts/OpenThoughts-Agent301—~1.7kAutomated safety check: PassApache-2.0
Supabase Postgres Best Practicessupabase/agent-skills2.7k24 repos~808Automated safety check: PassMIT
Supabase Development and Debuggingsupabase/agent-skills2.7k3 repos~3.6kAutomated safety check: PassMIT
Clickhouse Logs Queriessupabase/supabase111k—~2.4kAutomated safety check: PassApache-2.0
Security Reviewjewbetcha/opentrace11618 repos~3.1kAutomated safety check: NotesMIT
Review The Docssupabase/supabase111k—~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Gives the agent Postgres rules to consult before writing or changing tables, queries, indexes, RLS policies or migrations, and when diagnosing slow queries.

    2.7k GitHub starsUsed in 24 repos~808 tokens
    DatabasesAuto-check passed
  • Official

    General Supabase skill for database, auth, Edge Functions, Realtime and storage work, plus client libraries, migrations, security audits, debugging and reading logs.

    2.7k GitHub starsUsed in 3 repos~3.6k tokens
    Backend & APIsAuto-check passed
  • Clickhouse Logs Queries

    supabase/supabase

    Official

    Write, review, and migrate Supabase logs queries against the ClickHouse-backed logs table (the logs.all.otel analytics endpoint).

    111k GitHub stars~2.4k tokensUpdated today
    DatabasesAuto-check passed
  • Security Review

    jewbetcha/opentrace

    A skill your agent uses when adding authentication, handling user input, working with secrets, creating API endpoints, or implementing payment/sensitive features.

    116 GitHub starsUsed in 18 repos~3.1k tokens
    SecurityAuto-check: notes
  • Review The Docs

    supabase/supabase

    Official

    Review Supabase docs changes locally in your supabase/supabase checkout — either an open PR (triage, classify, verify) or your own branch before opening a PR (local self-review).

    111k GitHub stars~4.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Safe SQL Execution

    supabase/supabase

    Official

    A skill your agent uses whenever code will build, return, fetch, or execute SQL that runs against a user's real Postgres database — even when the request reads like an ordinary feature or bug fix…

    111k GitHub stars~4.2k tokensUpdated today
    DatabasesAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 10 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about Eval Standard Cleanup

What does Eval Standard Cleanup do?

Consolidate FINISHED standard / lmeval (evalchemy) math-suite eval jobs — the Delphi 6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker. Eval Standard Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent.md tracker.

When should I use Eval Standard Cleanup?

Eval Standard Cleanup fits situations like: asked to consolidate / harvest finished standard math-eval jobs; fill the scaling-laws score grid.

How do I install Eval Standard Cleanup in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-cleanup -a claude-code`. Or copy the skill folder (.agents/skills/eval-standard-cleanup in open-thoughts/OpenThoughts-Agent) into .claude/skills/eval-standard-cleanup in your project. Claude Code loads it when a task matches its description.

How do I install Eval Standard Cleanup in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-cleanup -a codex`. Or copy the skill folder (.agents/skills/eval-standard-cleanup in open-thoughts/OpenThoughts-Agent) into .agents/skills/eval-standard-cleanup in your project. Codex loads it when a task matches its description.

Can I use Eval Standard Cleanup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-cleanup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-standard-cleanup, .gemini/skills/eval-standard-cleanup, .github/skills/eval-standard-cleanup and .opencode/skills/eval-standard-cleanup in your project.

What does Eval Standard Cleanup need to run?

Going by SKILL.md and its folder, Eval Standard Cleanup needs the command-line tools its instructions call (rsync).

Does Eval Standard Cleanup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Standard Cleanup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Standard Cleanup use?

Eval Standard Cleanup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Standard Cleanup use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Standard Cleanup?

Skills that share tags, products or a category with Eval Standard Cleanup: Supabase Postgres Best Practices (supabase/agent-skills, 2.7k stars), Supabase Development and Debugging (supabase/agent-skills, 2.7k stars), Clickhouse Logs Queries (supabase/supabase, 111k stars) and Security Review (jewbetcha/opentrace, 116 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Standard Cleanup?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.