Agent skill

Eval Agentic Launch

by open-thoughts in open-thoughts/OpenThoughts-Agent

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Eval Agentic Launch

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-agentic-launch .claude/skills/eval-agentic-launch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-agentic-launch
GitHub stars
301
Token cost
~3.7k tokens
SKILL.md length
1,485 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

  • Works in 6 steps: Select the models → Harbor config + timeout multiplier… → Pinggy tunnel — installed-harness ONLY… → …
  • Asked to launch/relaunch agentic evals
  • SKILL.md covers 1. Select the models, 2. Harbor config + timeout…, 3. Pinggy tunnel —… and 4. Launch (in tmux — listener…, plus 3 more sections
  • Calls python, conda and ssh; needs DAYTONA_API_KEY and DAYTONA_DATA_API_KEY

What it does

Eval Agentic Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy served-model tunnel, submit with the right preset + flags in tmux, then VERIFY the launch actually works via the 15-min infra sanity check (pinggy auth, Daytona→cluster apibase, vLLM POSTs, trial progression — catches "RUNNING but silently dead" jobs). Cluster-AGNOSTIC: per-cluster particulars (sbatch script, gpu-mem ceiling…

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and LLM inference and serving. It works with tmux and vLLM. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Asked to launch/relaunch agentic evals
  • Eval a model on a benchmark (terminalbench2 / devsetv2 / swebench / bfcl / aider)

Example prompts

  • “RUNNING but silently dead”
  • “/eval-agentic-launch”

Requirements

  • Python 3
  • A credential in DAYTONA_API_KEY
  • A credential in DAYTONA_DATA_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Select the models
  2. Harbor config + timeout multiplier (config-by-size — usually nothing to do)
  3. Pinggy tunnel — installed-harness ONLY (not terminus-2)
  4. Launch (in tmux — listener is long-running)
  5. VERIFY the launch — 15-min infra sanity check (do NOT trust "RUNNING")
  6. Trial directory layout

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • conda
    • ssh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use ssh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DAYTONA_API_KEY
    • DAYTONA_DATA_API_KEY
    • SUPABASE_SERVICE_ROLE_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Agentic Launch loads about 3.7k tokens when it runs. Until then it costs about 197 tokens; SKILL.md has 1,485 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~197
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,485 words, ~3,743 tokens.

Download SKILL.mdSave it as .claude/skills/eval-agentic-launch/SKILL.md (or your agent's skills folder).
name
eval-agentic-launch
description
Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unified_eval_listener.py) on any cluster: select models (query_unevaled_models.py / priority lists), wire the pinggy served-model tunnel, submit with the right preset + flags in tmux, then VERIFY the launch actually works via the 15-min infra sanity check (pinggy auth, Daytona→cluster api_base, vLLM POSTs, trial progression — catches "RUNNING but silently dead" jobs). Cluster-AGNOSTIC: per-cluster particulars (sbatch script, gpu-mem ceiling, concurrency, cert/tunnel, conda env, paths, Daytona key, pre-download) live in `.agents/ops/<cluster>/`. Use when asked to launch/relaunch agentic evals, or eval a model on a benchmark (terminal_bench_2 / dev_set_v2 / swebench / bfcl / aider).

eval-agentic-launch

Launch agentic Harbor evals via the unified eval listener (eval/unified_eval_listener.py). Cluster-agnostic; read .agents/ops/<cluster>/ops.md first for the cluster's sbatch script, gpu-mem ceiling, concurrency, cert/tunnel, conda env, paths, Daytona eval-org key, and whether --pre-download is needed.

Front door: python -m hpc.launch --job_type eval_listener …. Runs the listener's main() in-process after the launcher preamble (detect_hpc + set_environment → DCFT/EXPERIMENTS_DIR/PYTHONPATH + hosted-vllm/Supabase keys + chdir to repo root), so no manual source hpc/dotenv/<cluster>.env / export PYTHONPATH / cd is needed. Forwards the listener's ~50 flags verbatim (strips only --job_type eval_listener). The raw python eval/unified_eval_listener.py … fallback still works (same public API) but you own the preamble — if you ever see FATAL: WORKDIR=... is not the OpenThoughts-Agent repo root, you used the raw script from the wrong place; switch to the front door.

⚠ Secrets from $DC_AGENT_SECRET_ENV, never hardcoded in a script/config/commit. The eval sbatch sources it (~/secrets.env; TACC $SCRATCH/keys.env) and reads the two Daytona eval-org keys: DAYTONA_API_KEY (org1) + DAYTONA_DATA_API_KEY (org2), 3:1-weighted (3/4 org2). Fails loudly (:?) if either is unset. A literal dtn_… key committed anywhere is a leak — rotate/revoke it, don't just fix-forward.

1. Select the models

  • Priority list (default): a file in eval/lists/ (models_8b_*.txt, models_32b.txt, models_131k.txt). Launch with --require-priority-list --priority-file eval/lists/<file>.
  • Find unevaled models — scripts/database/query_unevaled_models.py (resolves benchmark families via the Supabase duplicate_of field, e.g. dev_set_v2 ⊇ DCAgent_dev_set_v2/dev_set_v2_2.0x/openthoughts-tblite):
    bash
    python scripts/database/query_unevaled_models.py --benchmark <fam> --size <8|32> -o eval/lists/<file>.txt -v
    # needs SUPABASE_URL + SUPABASE_SERVICE_ROLE_KEY
--require-priority-list is LOAD-BEARING

--priority-file alone only changes sort order; the filter "skip models not in the list" lives behind --require-priority-list (unified_eval_listener.py ~L978). Without it the listener submits evals for every unevaled model in the lookback window (routinely 700+). Always pass both. If launched without it by accident: kill the listener before it leaves pre-download (submission is after Pre-downloading…), then confirm via squeue/sacct --starttime=now-Nmin. The listener python is a child of the sshd: …@notty session and survives the local ssh client dying — pkill -9 -f unified_eval_listener.py on the cluster to stop it.

Benchmark = a --preset (one, not both with --datasets)

tb2=terminal_bench_2, v2=dev_set_v2, dev=dev_set_71_tasks, swebench, bfcl, aider. --preset swebench is the random-100 subset (DCAgent/swebench_verified_eval_set → swebench-verified-random-100-folders, n_concurrent 32), not the full set.

"ID evals" — launch all three legs

Each leg is a separate listener invocation (different n_concurrent/harbor-config, don't combine into one --datasets):

leg--presetdataset (post-alias)n_concurrent
SWE-bench-verified random-100swebenchswebench-verified-random-100-folders32
dev_set_v2v2DCAgent/dev_set_v2128
terminal_bench_2tb2DCAgent2/terminal_bench_264

"Run the ID evals" = fire one listener per leg (§4) + the §5 infra check on each. (Full SWE-bench-verified and other benchmarks are OOD.) Scoring side (crud-otagent-supabase) uses the same 3-member set; dev_set_v2 is partial-credit → counts toward the ID mean but excluded from the ID SE and model-vs-model ranking.

Re-eval / parity test → --force-eval

By default the listener Skips any model with a Finished+metrics row (reason=job finished) — correct for cohort fill, but blocks a deliberate re-run. --force-eval bypasses that dedup and submits a fresh sandbox_jobs row (doesn't touch the existing row → no metrics-clearing, works across users). Pair with --require-priority-list + a single-model --priority-file so only the intended model is forced. --stale-started-hours does NOT override a Finished row (only re-ages Started). Distinct from --force-reeval (resume-path flag, see eval-agentic-cleanup check 4).

2. Harbor config + timeout multiplier (config-by-size — usually nothing to do)

Do NOT pass --harbor-config for standard terminus-2 evals. The listener selects the canonical config by model size and sets EVAL_HARBOR_CONFIG per-model:

model sizeselected configtimeout multiplier
8B-class (≤ ~14B; 1.5B/7B/14B)hpc/harbor_yaml/eval/dcagent_eval_defaults.yaml2×
32B-class (~28–42B; incl. MoE 30b-a3b)hpc/harbor_yaml/eval/dcagent_eval_defaults_32b.yaml16×
out-of-band (70B/80B) or no size tokenbase default2× + logged note

Size is read from the largest \dB token in the HF name. The multiplier flows as EVAL_TIMEOUT_MULTIPLIER and is recorded in the Pending row so dedup matches what ran. The deprecated eval_ctx*_non_it* / ctx32k_non_it_16x_eval_.yaml configs carry stale *-drop-ei metrics → JobConfig ValidationError.

Resolution order (first wins): (1) explicit --harbor-config / preset harbor_config — overrides size selection for every model (use for 131k context / openhands_* installed-harness); (2) per-model timeout_multiplier: in the registry (for names with no size token, e.g. a Qwen3-8B named laion/GLM-4_7-swesmith-…); (3) size-based table above. For a one-off harbor jobs start, point --config at the 8B/32B file.

3. Pinggy tunnel — installed-harness ONLY (not terminus-2)

Skip for the default terminus-2 agent (every eval_ctx*/*_non_it* config; all --presets). Do NOT pass --pinggy_* / consume a pair.

Installed harnesses (opencode / openhands_*) run in the Daytona sandbox and call back out to the served model over a public pinggy tunnel. The full recipe for the opencode installed-harness + pinggy ID-eval — the config (eval_opencode_ctx32k.yaml), config-delivery mechanism (--config-yaml on TACC vs --harbor-config), the --pinggy_persistent_url/--pinggy_token flags + pairs 8/9/10, the sbatch installed-agent tunnel/routing branch, the vllm/ provider, -Thinking- model specifics, and the pinggy infra checks — lives in .agents/projects/harbor/ops.md → "Agentic ID-eval via the opencode (installed) harness + pinggy". The privileged URL/token bank is in .agents/secret.md / notes/ot-agent/pinggy_bank.md (never inline it). Resume of an installed-harness eval also needs the tunnel — see eval-agentic-cleanup check 4.

4. Launch (in tmux — listener is long-running)

Concurrent-submit guard: ONE listener enqueues many legs; do NOT fire N concurrent --once processes. A multi-leg refill is one invocation that submits each leg internally with a 1s submission_delay (unified_eval_listener.py L3204–3205). Firing N listener processes near-simultaneously on the login node races conda's lazily-imported plugin registry → a circular-import at activation. If multiple listener processes are truly required (incompatible n_concurrent), stagger them ~30–45s apart — never & them together. (Per-job conda activate inside the sbatch runs on independent compute nodes and never races.)

bash
# inside tmux. The front door does the preamble (no manual source/PYTHONPATH).
python -m hpc.launch --job_type eval_listener \
  --cluster-config <cluster-name> \
  --preset <preset> \
  --require-priority-list --priority-file eval/lists/<file>.txt \
  --config-yaml dcagent_eval_config_no_override.yaml \
  [--agent-kwarg 'extra_body={"chat_template_kwargs":{"enable_thinking":true}}'] [--agent-parser json] [--max-output-tokens 16384] \
  [--pre-download] [--force-reeval] [--pinggy_persistent_url <URL> --pinggy_token <TOKEN>] \
  --once --verbose 2>&1 | tee eval/<cluster>/logs/<preset>_listener_$(date +%Y%m%d_%H%M%S).log
# Raw-script fallback (you own the preamble): from repo root, export PYTHONPATH="$PWD:${PYTHONPATH:-}", run python eval/unified_eval_listener.py … with the same flags.
  • --cluster-config takes a bare cluster name (leonardo, tacc) resolved from hpc.hpc's eval_cluster_view (a .yaml path still works as back-compat). Supplies sbatch_script/hardware/conda_envs/paths — so you no longer pass --sbatch-script/--n-concurrent/--gpu-memory-util.
  • Per-model serve config (conda_env, tensor_parallel_size, data_parallel_size, max_model_len, limit_mm_per_prompt, max_output_tokens) comes from the shared registry by default — no flag. The cluster yaml's hardware_profile: (e.g. gh200) selects the per-cluster recipe; a per-cluster intrinsic delta is name@<profile>, a hardware delta is variants: {<profile>: {…}}. Confirm it loaded: listener logs Model-config registry ENABLED + Loaded model registry: N model config(s) + Using conda env '<env>' for <model>. --baseline-model-configs is deprecated (opt-out of the registry).
  • Edit per-model serve config in model_config/<org>/<slug>.yaml, NOT the generated eval/configs/model_configs.yaml (auto-generated, carries a # do NOT hand-edit banner). Regenerate with python scripts/generate_eval_registry.py (drift gate: --check).
  • Thinking is per-model authoritative (sourced from the registry via agent_kwargs: [extra_body={…enable_thinking:true}]); presets never carry thinking; there is no --enable-thinking flag. Override with --agent-kwarg 'extra_body={"chat_template_kwargs":{"enable_thinking":true}}' (precedence: CLI > registry > preset).
Show full SKILL.md (489 more words)Show less

5. VERIFY the launch — 15-min infra sanity check (do NOT trust "RUNNING")

A job can report RUNNING while nothing happens (pinggy locked, launcher missing --pinggy_*, dead vLLM engine). After launching, schedule a 15-min (ScheduleWakeup delaySeconds: 900) check and re-arm each pass until the eval terminates / you have a verdict.

Checks 1–2 are pinggy-path (installed-harness) ONLY — skip for terminus-2. For terminus-2, served-model reachability is proven by check 3. Checks 3–4 apply to every launch.

  1. Pinggy tunnel (installed-harness) — pinggy auth You are authenticated as … + a growing RB:/SB:/TC: traffic counter. Full checks (auth, sandbox api_base = public *.a.pinggy.link/v1 not 10.*, relaunch-on-lock) → .agents/projects/harbor/ops.md § opencode + pinggy.
  2. Daytona → cluster (installed-harness) — a trial's config.json api_base MUST be https://*.a.pinggy.link/v1, NOT 10.*.*.* (see the harbor ops.md opencode+pinggy section).
  3. vLLM serving — POST /v1/chat/completions count grows ≥ a few/min, 200 OK dominates. 400 ratio > 15% → context overflow (VLLMValidationError: input tokens … → lower max_input_tokens/max_output_tokens in the harbor yaml).
  4. Trial progression — count trials with agent/ populated (active) and result.json (done). 30+ min with zero agent/command-0/ (OpenHands) → setup stalled. Completions with n_output_tokens: None and agent_execution.finished_at ≈ started_at (instant-fail) = tunnel not carrying traffic despite a healthy-looking job.

Quick liveness (≈15 min after submit): ssh <cluster> "squeue -u $USER --format='%.18i %.50j %.8T %.10M'" then tail the newest log — vLLM health-check pass, (Leonardo) SSH tunnel up, trial/reward lines, no OOM/repeated DaytonaErrors.

6. Trial directory layout

<run_tag>/<task>__<trial_id>/: config.json (mtime≈start, has api_base), trial.log, result.json (timestamps + verifier_result.rewards.reward + exception_info), exception.txt, agent/trajectory.json, verifier/{reward.txt,detailed_scores.json}. Eval cleanup + manual DB register + trace upload → eval-agentic-cleanup.

Other gotchas

  • PermissionError: [Errno 13] at harbor/job.py … job_dir.mkdir() = a jobs_dir in the harbor config that another user owns. The canonical configs ship no jobs_dir; eval_harbor.sbatch passes --jobs-dir "$EVAL_JOBS_DIR" (per-user …/ot-baf/eval_jobs) which overrides the config. If you see this, confirm the sbatch has the --jobs-dir line; a hand-rolled harbor jobs start will reintroduce it. Resume is unaffected (takes -p $RUN_DIR).

  • A crashed eval leaves a non-terminal DB row blocking resubmission for 24h (reason=job in progress). After a crash the row stays started; the listener only resubmits started rows older than --stale-started-hours (default 24h, EVAL_LISTENER_STALE_HOURS). Pass a small value (e.g. --stale-started-hours 0.05 = 3 min) to force resubmit of the just-crashed attempt. Pending rows use --stale-pending-hours (default 6h, auto-cancels the stale SLURM job).

  • Jupiter: pass --reservation reformo or eval jobs starve behind RL (the reservation holds ~128 nodes while the general booster pool is empty). eval/jupiter/eval_harbor.sbatch sets --account reformo but no #SBATCH --reservation. Check scontrol show reservation for the live name/expiry before relying on it (the flag errors if the reservation is dead). Rescue already-PENDING jobs: scontrol update jobid=<j> reservation=reformo.

  • hosted_vllm/<org>/<model> evals need harbor commit 0f5a6e9e (allows 2-slash org-qualified names — validate_hosted_vllm_model_config in llms/utils.py) and model_info supplied via --agent-kwarg ({"max_input_tokens":…,"max_output_tokens":…,"input_cost_per_token":0,"output_cost_per_token":0}; token limits from the served vLLM max_model_len, costs 0 = self-hosted). Both are wired by default into the eval_harbor.sbatch files via EVAL_VLLM_MAX_MODEL_LEN (default 32768) + EVAL_MAX_OUTPUT_TOKENS (default 16384) (OT-Agent commit d0064011). If org-model evals fast-fail (~9 min, 0 POST 200s, 0 trajectories, all N trials raise identically), confirm those commits are in the cluster's harbor clone.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/eval-agentic-launch of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Eval Agentic Launch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Agentic Launch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Agentic Launch this skillopen-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Hugging Face Community Evalssickn33/agentic-awesome-skills47k1 repos~1.9kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k3 repos~3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Diffusion Perf Optvllm-project/vllm-omni7.1k—~7.5kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Community Evals

    sickn33/agentic-awesome-skills

    Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

    47k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 3 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 10 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about Eval Agentic Launch

What does Eval Agentic Launch do?

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…. Eval Agentic Launch is an agent skill from open-thoughts/OpenThoughts-Agent.py / priority lists), wire the pinggy served-model tunnel, submit with the right preset + flags in tmux, then VERIFY the launch actually works via the 15-min infra sanity check (pinggy auth, Daytona→cluster apibase, vLLM POSTs, trial progression — catches "RUNNING but silently dead" jobs).

When should I use Eval Agentic Launch?

Eval Agentic Launch fits situations like: asked to launch/relaunch agentic evals; eval a model on a benchmark (terminalbench2 / devsetv2 / swebench / bfcl / aider).

How do I install Eval Agentic Launch in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a claude-code`. Or copy the skill folder (.agents/skills/eval-agentic-launch in open-thoughts/OpenThoughts-Agent) into .claude/skills/eval-agentic-launch in your project. Claude Code loads it when a task matches its description.

How do I install Eval Agentic Launch in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a codex`. Or copy the skill folder (.agents/skills/eval-agentic-launch in open-thoughts/OpenThoughts-Agent) into .agents/skills/eval-agentic-launch in your project. Codex loads it when a task matches its description.

Can I use Eval Agentic Launch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-agentic-launch, .gemini/skills/eval-agentic-launch, .github/skills/eval-agentic-launch and .opencode/skills/eval-agentic-launch in your project.

What does Eval Agentic Launch need to run?

Going by SKILL.md and its folder, Eval Agentic Launch needs the command-line tools its instructions call (python, conda and ssh) and credentials named DAYTONA_API_KEY, DAYTONA_DATA_API_KEY and SUPABASE_SERVICE_ROLE_KEY. Our summary lists: Python 3; A credential in DAYTONA_API_KEY; A credential in DAYTONA_DATA_API_KEY.

Does Eval Agentic Launch access the network?

SKILL.md contains no URLs. Its commands use ssh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval Agentic Launch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Agentic Launch use?

Eval Agentic Launch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Agentic Launch use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Agentic Launch?

Skills that share tags, products or a category with Eval Agentic Launch: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Hugging Face Community Evals (sickn33/agentic-awesome-skills, 47k stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Agentic Launch?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.