Agent skill

Eval Standard Launch

by open-thoughts in open-thoughts/OpenThoughts-Agent

Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Eval Standard Launch

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-launch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent eval-standard-launch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-standard-launch .claude/skills/eval-standard-launch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-standard-launch
GitHub stars
301
Token cost
~2.3k tokens
SKILL.md length
997 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints.

  • Works in 7 steps: Reference files (read these first — they… → What "newly completed" means → The launch (per cell) → …
  • Asked to eval Delphi 6279 SFT cells / update the scaling-laws score grid
  • SKILL.md covers 0. Reference files (read these…, 1. What "newly completed" means, 2. The launch (per cell) and 3. Pre-download is REQUIRED…, plus 4 more sections
  • Calls bash and ssh; needs HF_TOKEN

What it does

Eval Standard Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints. Covers finding which cells are newly-completed-but-uneval'd, the offline pre-download, the delphieval.sbatch invocation + RUNNAME/STAGE convention, the load-bearing gotchas (chat-template override, 4k context, TP per head-count, MATH500/gsm8k split), and the SCORES.md tracker update. HF-upload-only — NEVER DB-register. Use when…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asked to eval Delphi 6279 SFT cells / update the scaling-laws score grid

Example prompts

  • “/eval-standard-launch”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Reference files (read these first — they are the source of truth)
  2. What "newly completed" means
  3. The launch (per cell)
  4. Pre-download is REQUIRED (compute is offline)
  5. Load-bearing gotchas (all handled inside the sbatch — know them when debugging)
  6. After submit → tracking + consolidation
  7. Leonardo SSH quoting traps (these bite repeatedly)

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • ssh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use ssh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Standard Launch loads about 2.3k tokens when it runs. Until then it costs about 168 tokens; SKILL.md has 997 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~168
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 997 words, ~2,301 tokens.

Download SKILL.mdSave it as .claude/skills/eval-standard-launch/SKILL.md (or your agent's skills folder).
name
eval-standard-launch
description
Launch the fixed Delphi #6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lm_eval) on CINECA Leonardo, for completed SFT / RL / base checkpoints. Covers finding which cells are newly-completed-but-uneval'd, the offline pre-download, the delphi_eval.sbatch invocation + RUN_NAME/STAGE convention, the load-bearing gotchas (chat-template override, 4k context, TP per head-count, MATH500/gsm8k split), and the SCORES.md tracker update. HF-upload-only — NEVER DB-register. Use when asked to eval Delphi #6279 SFT cells / update the scaling-laws score grid. Refs: experiments/active/delphi/rl-scaling-laws-6279/.

eval-standard-launch

Downstream math-eval harness for Delphi #6279 RL-scaling-laws (task #215) — NOT the agentic tb2 eval listener. Scores Delphi checkpoints on MATH-500 (1 seed) + AIME24 (10-seed mean±se) + gsm8k (strict+flex) via evalchemy/lm_eval on Leonardo. HF-upload only — NEVER DB-registered (project_delphi_sft_hf_only_no_db).

0. Reference files (read these first — they are the source of truth)

Local notes dir: /Users/benjaminfeuer/Documents/experiments/active/delphi/rl-scaling-laws-6279/

  • EVAL_CONVENTION.md — the fixed eval protocol (suites, seeds, naming §3.4, chat-template §2.5).
  • delphi_eval.sbatch — the ONE eval job (all gotchas baked in; canonical). Cluster copy under /leonardo_work/AIFAC_5C0_290/bfeuer00/….
  • main_sft_evals/SCORES.md — master tracker for the 54-run main grid (27 midtrained ckpts × 2 cold-starts: magpie_lr1e5 math-strong / wc386k_lr1e5 math-weak). Each row: HF model laion/<basename>, eval-job column, status (✅ done / ⏳ pending). The separate earlier cold-start grid lives in coldstart_grid_evals/ (own ../eval/SCORES.md).
  • SFT_LEONARDO_INSTRUCTIONS.md — the SFT side (how cells get trained + uploaded).

1. What "newly completed" means

A cell is ready to eval when its SFT model is uploaded to HF laion/<basename> (Delphi SFT is HF-only; an uploaded repo = a finished cell). The work = the SCORES.md rows still ⏳ pending whose laion/<basename> repo now exists + is non-empty (check via huggingface_hub/hf; HF_TOKEN from your secrets env — .agents/secret.md). Rows whose model isn't uploaded yet (e.g. most large 1e21/1e22 cells mid-training) are SKIPPED until done.

2. The launch (per cell)

bash
# RUN_NAME = the exact SFT model basename; STAGE = sft (chat-template ON) | rl | base (template OFF)
RUN=delphi-9e19-p33m67-k0p20-lr83-a002-magpie_lr1e5-sft
sbatch --job-name="delphi-eval-$RUN" <leonardo>/delphi_eval.sbatch laion/$RUN $RUN sft
  • Model-specific --job-name is MANDATORY (not the generic default) so squeue + %x-%j.log + meta.env map jobid→model 1:1. The script self-renames + emits a greppable EVAL_JOBMAP line + writes <OUT>/meta.env.
  • Job shape: 1 node / 4 GPU (A100 64GB) / 8h / boost_usr_prod, conda env evalchemy-marin, runs from /leonardo_work/AIFAC_5C0_290/bfeuer00/code/evalchemy-marin. Output → …/experiments/delphi-eval/<RUN_NAME>/.

3. Pre-download is REQUIRED (compute is offline)

delphi_eval.sbatch runs HF_HUB_OFFLINE=1 (Leonardo compute has no internet). Pre-cache each model on the LOGIN node first into HF_HOME=HF_HUB_CACHE=/leonardo_work/AIFAC_5C0_290/bfeuer00/data/hub (login nodes have direct internet). If a login-node snapshot_download risks the ~100s login-killer, use the notes' documented pre-download path (tmux / a small sbatch). The eval sbatch itself needs NO SSH tunnel (offline + pre-cached); only the pre-download touches the network.

4. Load-bearing gotchas (all handled inside the sbatch — know them when debugging)

  • HOME is read-only on Leonardo (login AND compute). The sbatch redirects HOME + flashinfer/triton/inductor/vLLM/XDG caches to a writable …/delphi-eval/.cache/*; without it every vLLM worker dies PermissionError … /.cache/flashinfer.
  • delphi_v0 chat-template override (sft/rl only): the SFT/RL repos ship a plain 656-char Llama-3 template (the delphi ReasoningTemplate didn't persist into the repo) → evaluating as-is is a train/eval mismatch (empty think channel). The sbatch overrides the cached tokenizer's chat_template to OpenThoughts-Agent/chat_templates/delphi_v0.jinja2 (idempotent, leaves .plainbak) before eval. base stage skips this (no template, raw completion).
  • MAX_MODEL_LEN=4096 (NOT the marin 32768 default): Delphi ckpts are 4k-cutoff with a malformed llama3 rope_scaling block; vLLM derives 4096 and HARD-rejects 32768. Generation pinned MAX_GEN_TOKS=3584. Comparability holds within the 4k cohort; flag any model exposing >4k.
  • --max_tokens MUST be pinned to MAX_GEN_TOKS — MATH500/AIME24 are evalchemy chat_benchmarks whose max_new_tokens DEFAULTS to 32768 (not reached by --gen_kwargs max_gen_toks); unset → lm-eval computes 4096-32768 = negative → truncates prompt to empty → decoder prompt cannot be empty.
  • TP per head-divisibility: num_attention_heads % TP == 0. The small Delphi Qwen3 (14 heads) → TP=2 (TP=4 hard-fails). Pin TP per model to the largest node-supported divisor of its head count.
  • MATH500 and gsm8k run as SEPARATE sequential processes (not one --tasks call): MATH500 is an evalchemy chat_benchmark, gsm8k is lm-eval-native; in one process the second vLLM engine inits while the first is GPU-resident → OOM/WorkerProc fail and gsm8k is silently dropped. gsm8k runs via plain lm_eval (evalchemy double-builds the engine for native tasks → OOM on 64GB). AIME24 is its own pass.
  • --verbosity INFO is required (evalchemy getattr(logging, args.verbosity); default None → AttributeError after full vLLM init).
  • Idempotent skip: <OUT>/seed42 existing → the cell is skipped. If gsm8k failed after MATH500 wrote seed42, delete <OUT>/seed42 (or re-run gsm8k by hand) — the skip keys on that dir.
Show full SKILL.md (409 more words)Show less

5. After submit → tracking + consolidation

  1. Confirm queued: squeue -u bfeuer00 | grep delphi-eval; collect job ids.
  2. Update main_sft_evals/SCORES.md: set submitted rows' status → 🚀 eval submitted + put the Leonardo job id in the eval-job column. Do NOT fabricate score cells (leave —); preserve the table format exactly.
  3. On completion, per-model results_*.json rsync into main_sft_evals/<basename>/, scalar partials into main_sft_evals/.partial/<basename>.json, and SCORES.md is consolidated from them (MATH-500, AIME24 mean±se, gsm8k strict/flex, Raw). The #6279 deliverable = how MATH-500/AIME24/gsm8k move with (scale × mix) at each of the two starting points (does a strong-vs-weak math start change the midtraining ranking).

5b. Evaluate a (post-RL) checkpoint on the Delphi eval suite (reusable)

The same harness scores ANY standard Delphi checkpoint — base / post-SFT / post-RL — on the fixed suite (EVAL_CONVENTION.md §1.2: MATH500 1-seed + AIME24 10-seed mean±se + gsm8k strict/flex, pass@1, temp 0.7). rl-standard-job-cleanup defers to THIS section as its final step after the post-RL ckpt is HF-uploaded. The only deltas from §2 are the STAGE token and the tracker the result lands in:

  1. Point it at the HF-uploaded ckpt. The ckpt is laion/<run_name>-<BEST>-<size>B (the repo rl-standard-job-cleanup §6 just published — weights at root). Pre-cache it on the login node first (§3), same as any cell.
  2. STAGE = rl (chat-template ON — same delphi_v0 override as sft; only base skips it):
    bash
    RUN=<run_name>-<BEST>-<size>B
    sbatch --job-name="delphi-eval-$RUN" \
      /leonardo_work/AIFAC_5C0_290/bfeuer00/experiments/delphi-eval/delphi_eval.sbatch laion/$RUN $RUN rl
    Auto-TP=2 for the 30-head 9.7B Delphi Qwen3; max_model_len=4096/max_gen_toks=3584; all §4 gotchas (template override, 4k context, MATH500/gsm8k split, AIME24 10-seed pass) apply unchanged — baked into the canonical delphi_eval.sbatch. One node / 4 GPU / 8h / evalchemy-marin env.
  3. Results land in the RL tracker, not the SFT one. Per-model output is …/experiments/delphi-eval/<RUN>/seed{42..51} + meta.env; consolidated scores go to main_rl_evals/SCORES.md (the post-RL tracker — keyed by (scale, mix, start-point) so SFT-vs-RL deltas line up against main_sft_evals/), via the same harvest path as §5 / eval-standard-cleanup. HF-upload-only — NEVER DB (the post-RL ckpt has no models DB row and this eval doesn't create one). After submit, add the row to main_rl_evals/SCORES.md set to 🚀 eval submitted with the Leonardo job id; harvest per eval-standard-cleanup.

6. Leonardo SSH quoting traps (these bite repeatedly)

  • Do NOT use parentheses inside a bash -lc "..." double-quoted string.
  • Do NOT use single quotes inside the outer ssh '...' arg (a single quote closes it). Use escaped double-quotes / heredocs / plain words.
  • Refresh the step-ca cert if any tunnel op needs it (per CLAUDE.md) — but offline eval sbatch + a login-node pre-download (direct internet) do not need the tunnel.
  • Don't disturb the still-PENDING Delphi SFT dependency chain; evals are independent 1-node jobs.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/eval-standard-launch of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Eval Standard Launch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Standard Launch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Standard Launch this skillopen-thoughts/OpenThoughts-Agent301—~2.3kAutomated safety check: PassApache-2.0
Agent BuildershareAI-lab/learn-claude-code78k6 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.8k15 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 6 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.8k GitHub starsUsed in 15 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 10 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed

Questions about Eval Standard Launch

What does Eval Standard Launch do?

Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints. Eval Standard Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints.

When should I use Eval Standard Launch?

Eval Standard Launch fits situations like: asked to eval Delphi 6279 SFT cells / update the scaling-laws score grid.

How do I install Eval Standard Launch in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-launch -a claude-code`. Or copy the skill folder (.agents/skills/eval-standard-launch in open-thoughts/OpenThoughts-Agent) into .claude/skills/eval-standard-launch in your project. Claude Code loads it when a task matches its description.

How do I install Eval Standard Launch in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-launch -a codex`. Or copy the skill folder (.agents/skills/eval-standard-launch in open-thoughts/OpenThoughts-Agent) into .agents/skills/eval-standard-launch in your project. Codex loads it when a task matches its description.

Can I use Eval Standard Launch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-standard-launch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-standard-launch, .gemini/skills/eval-standard-launch, .github/skills/eval-standard-launch and .opencode/skills/eval-standard-launch in your project.

What does Eval Standard Launch need to run?

Going by SKILL.md and its folder, Eval Standard Launch needs the command-line tools its instructions call (bash and ssh) and credentials named HF_TOKEN.

Does Eval Standard Launch access the network?

SKILL.md contains no URLs. Its commands use ssh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval Standard Launch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Standard Launch use?

Eval Standard Launch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Standard Launch use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Standard Launch?

Skills that share tags, products or a category with Eval Standard Launch: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Standard Launch?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.