Hugging Face Local Model Evals
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-agentic-launch .claude/skills/eval-agentic-launch && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-agentic-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch into .claude/skills/eval-agentic-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/eval-agentic-launch .agents/skills/eval-agentic-launch && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-agentic-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch into .agents/skills/eval-agentic-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/eval-agentic-launch .cursor/skills/eval-agentic-launch && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-agentic-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch into .cursor/skills/eval-agentic-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/eval-agentic-launch--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/eval-agentic-launch .gemini/skills/eval-agentic-launch && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-agentic-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch into .gemini/skills/eval-agentic-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/eval-agentic-launch .github/skills/eval-agentic-launch && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-agentic-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch into .github/skills/eval-agentic-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent eval-agentic-launch --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/eval-agentic-launch .opencode/skills/eval-agentic-launch && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-agentic-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/eval-agentic-launch into .opencode/skills/eval-agentic-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-agentic-launch", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-agentic-launchLaunch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…
Eval Agentic Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy served-model tunnel, submit with the right preset + flags in tmux, then VERIFY the launch actually works via the 15-min infra sanity check (pinggy auth, Daytona→cluster apibase, vLLM POSTs, trial progression — catches "RUNNING but silently dead" jobs). Cluster-AGNOSTIC: per-cluster particulars (sbatch script, gpu-mem ceiling…
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation and LLM inference and serving. It works with tmux and vLLM. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythoncondasshFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use ssh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DAYTONA_API_KEYDAYTONA_DATA_API_KEYSUPABASE_SERVICE_ROLE_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Agentic Launch loads about 3.7k tokens when it runs. Until then it costs about 197 tokens; SKILL.md has 1,485 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,485 words, ~3,743 tokens.
.claude/skills/eval-agentic-launch/SKILL.md (or your agent's skills folder).Launch agentic Harbor evals via the unified eval listener (eval/unified_eval_listener.py). Cluster-agnostic; read .agents/ops/<cluster>/ops.md first for the cluster's sbatch script, gpu-mem ceiling, concurrency, cert/tunnel, conda env, paths, Daytona eval-org key, and whether --pre-download is needed.
Front door:
python -m hpc.launch --job_type eval_listener …. Runs the listener'smain()in-process after the launcher preamble (detect_hpc+set_environment→DCFT/EXPERIMENTS_DIR/PYTHONPATH+ hosted-vllm/Supabase keys +chdirto repo root), so no manualsource hpc/dotenv/<cluster>.env/export PYTHONPATH/cdis needed. Forwards the listener's ~50 flags verbatim (strips only--job_type eval_listener). The rawpython eval/unified_eval_listener.py …fallback still works (same public API) but you own the preamble — if you ever seeFATAL: WORKDIR=... is not the OpenThoughts-Agent repo root, you used the raw script from the wrong place; switch to the front door.
⚠ Secrets from
$DC_AGENT_SECRET_ENV, never hardcoded in a script/config/commit. The eval sbatch sources it (~/secrets.env; TACC$SCRATCH/keys.env) and reads the two Daytona eval-org keys:DAYTONA_API_KEY(org1) +DAYTONA_DATA_API_KEY(org2), 3:1-weighted (3/4 org2). Fails loudly (:?) if either is unset. A literaldtn_…key committed anywhere is a leak — rotate/revoke it, don't just fix-forward.
eval/lists/ (models_8b_*.txt, models_32b.txt, models_131k.txt). Launch with --require-priority-list --priority-file eval/lists/<file>.scripts/database/query_unevaled_models.py (resolves benchmark families via the Supabase duplicate_of field, e.g. dev_set_v2 ⊇ DCAgent_dev_set_v2/dev_set_v2_2.0x/openthoughts-tblite):python scripts/database/query_unevaled_models.py --benchmark <fam> --size <8|32> -o eval/lists/<file>.txt -v
# needs SUPABASE_URL + SUPABASE_SERVICE_ROLE_KEY--require-priority-list is LOAD-BEARING--priority-file alone only changes sort order; the filter "skip models not in the list" lives behind --require-priority-list (unified_eval_listener.py ~L978). Without it the listener submits evals for every unevaled model in the lookback window (routinely 700+). Always pass both. If launched without it by accident: kill the listener before it leaves pre-download (submission is after Pre-downloading…), then confirm via squeue/sacct --starttime=now-Nmin. The listener python is a child of the sshd: …@notty session and survives the local ssh client dying — pkill -9 -f unified_eval_listener.py on the cluster to stop it.
--preset (one, not both with --datasets)tb2=terminal_bench_2, v2=dev_set_v2, dev=dev_set_71_tasks, swebench, bfcl, aider. --preset swebench is the random-100 subset (DCAgent/swebench_verified_eval_set → swebench-verified-random-100-folders, n_concurrent 32), not the full set.
Each leg is a separate listener invocation (different n_concurrent/harbor-config, don't combine into one --datasets):
| leg | --preset | dataset (post-alias) | n_concurrent |
|---|---|---|---|
| SWE-bench-verified random-100 | swebench | swebench-verified-random-100-folders | 32 |
| dev_set_v2 | v2 | DCAgent/dev_set_v2 | 128 |
| terminal_bench_2 | tb2 | DCAgent2/terminal_bench_2 | 64 |
"Run the ID evals" = fire one listener per leg (§4) + the §5 infra check on each. (Full SWE-bench-verified and other benchmarks are OOD.) Scoring side (crud-otagent-supabase) uses the same 3-member set; dev_set_v2 is partial-credit → counts toward the ID mean but excluded from the ID SE and model-vs-model ranking.
--force-evalBy default the listener Skips any model with a Finished+metrics row (reason=job finished) — correct for cohort fill, but blocks a deliberate re-run. --force-eval bypasses that dedup and submits a fresh sandbox_jobs row (doesn't touch the existing row → no metrics-clearing, works across users). Pair with --require-priority-list + a single-model --priority-file so only the intended model is forced. --stale-started-hours does NOT override a Finished row (only re-ages Started). Distinct from --force-reeval (resume-path flag, see eval-agentic-cleanup check 4).
Do NOT pass --harbor-config for standard terminus-2 evals. The listener selects the canonical config by model size and sets EVAL_HARBOR_CONFIG per-model:
| model size | selected config | timeout multiplier |
|---|---|---|
| 8B-class (≤ ~14B; 1.5B/7B/14B) | hpc/harbor_yaml/eval/dcagent_eval_defaults.yaml | 2× |
32B-class (~28–42B; incl. MoE 30b-a3b) | hpc/harbor_yaml/eval/dcagent_eval_defaults_32b.yaml | 16× |
| out-of-band (70B/80B) or no size token | base default | 2× + logged note |
Size is read from the largest \dB token in the HF name. The multiplier flows as EVAL_TIMEOUT_MULTIPLIER and is recorded in the Pending row so dedup matches what ran. The deprecated eval_ctx*_non_it* / ctx32k_non_it_16x_eval_.yaml configs carry stale *-drop-ei metrics → JobConfig ValidationError.
Resolution order (first wins): (1) explicit --harbor-config / preset harbor_config — overrides size selection for every model (use for 131k context / openhands_* installed-harness); (2) per-model timeout_multiplier: in the registry (for names with no size token, e.g. a Qwen3-8B named laion/GLM-4_7-swesmith-…); (3) size-based table above. For a one-off harbor jobs start, point --config at the 8B/32B file.
Skip for the default terminus-2 agent (every eval_ctx*/*_non_it* config; all --presets). Do NOT pass --pinggy_* / consume a pair.
Installed harnesses (opencode / openhands_*) run in the Daytona sandbox and call back out to the served model over a public pinggy tunnel. The full recipe for the opencode installed-harness + pinggy ID-eval — the config (eval_opencode_ctx32k.yaml), config-delivery mechanism (--config-yaml on TACC vs --harbor-config), the --pinggy_persistent_url/--pinggy_token flags + pairs 8/9/10, the sbatch installed-agent tunnel/routing branch, the vllm/ provider, -Thinking- model specifics, and the pinggy infra checks — lives in .agents/projects/harbor/ops.md → "Agentic ID-eval via the opencode (installed) harness + pinggy". The privileged URL/token bank is in .agents/secret.md / notes/ot-agent/pinggy_bank.md (never inline it). Resume of an installed-harness eval also needs the tunnel — see eval-agentic-cleanup check 4.
Concurrent-submit guard: ONE listener enqueues many legs; do NOT fire N concurrent
--onceprocesses. A multi-leg refill is one invocation that submits each leg internally with a 1ssubmission_delay(unified_eval_listener.pyL3204–3205). Firing N listener processes near-simultaneously on the login node races conda's lazily-imported plugin registry → a circular-import at activation. If multiple listener processes are truly required (incompatible n_concurrent), stagger them ~30–45s apart — never&them together. (Per-jobconda activateinside the sbatch runs on independent compute nodes and never races.)
# inside tmux. The front door does the preamble (no manual source/PYTHONPATH).
python -m hpc.launch --job_type eval_listener \
--cluster-config <cluster-name> \
--preset <preset> \
--require-priority-list --priority-file eval/lists/<file>.txt \
--config-yaml dcagent_eval_config_no_override.yaml \
[--agent-kwarg 'extra_body={"chat_template_kwargs":{"enable_thinking":true}}'] [--agent-parser json] [--max-output-tokens 16384] \
[--pre-download] [--force-reeval] [--pinggy_persistent_url <URL> --pinggy_token <TOKEN>] \
--once --verbose 2>&1 | tee eval/<cluster>/logs/<preset>_listener_$(date +%Y%m%d_%H%M%S).log
# Raw-script fallback (you own the preamble): from repo root, export PYTHONPATH="$PWD:${PYTHONPATH:-}", run python eval/unified_eval_listener.py … with the same flags.--cluster-config takes a bare cluster name (leonardo, tacc) resolved from hpc.hpc's eval_cluster_view (a .yaml path still works as back-compat). Supplies sbatch_script/hardware/conda_envs/paths — so you no longer pass --sbatch-script/--n-concurrent/--gpu-memory-util.conda_env, tensor_parallel_size, data_parallel_size, max_model_len, limit_mm_per_prompt, max_output_tokens) comes from the shared registry by default — no flag. The cluster yaml's hardware_profile: (e.g. gh200) selects the per-cluster recipe; a per-cluster intrinsic delta is name@<profile>, a hardware delta is variants: {<profile>: {…}}. Confirm it loaded: listener logs Model-config registry ENABLED + Loaded model registry: N model config(s) + Using conda env '<env>' for <model>. --baseline-model-configs is deprecated (opt-out of the registry).model_config/<org>/<slug>.yaml, NOT the generated eval/configs/model_configs.yaml (auto-generated, carries a # do NOT hand-edit banner). Regenerate with python scripts/generate_eval_registry.py (drift gate: --check).agent_kwargs: [extra_body={…enable_thinking:true}]); presets never carry thinking; there is no --enable-thinking flag. Override with --agent-kwarg 'extra_body={"chat_template_kwargs":{"enable_thinking":true}}' (precedence: CLI > registry > preset).A job can report RUNNING while nothing happens (pinggy locked, launcher missing --pinggy_*, dead vLLM engine). After launching, schedule a 15-min (ScheduleWakeup delaySeconds: 900) check and re-arm each pass until the eval terminates / you have a verdict.
Checks 1–2 are pinggy-path (installed-harness) ONLY — skip for terminus-2. For terminus-2, served-model reachability is proven by check 3. Checks 3–4 apply to every launch.
You are authenticated as … + a growing RB:/SB:/TC: traffic counter. Full checks (auth, sandbox api_base = public *.a.pinggy.link/v1 not 10.*, relaunch-on-lock) → .agents/projects/harbor/ops.md § opencode + pinggy.config.json api_base MUST be https://*.a.pinggy.link/v1, NOT 10.*.*.* (see the harbor ops.md opencode+pinggy section).POST /v1/chat/completions count grows ≥ a few/min, 200 OK dominates. 400 ratio > 15% → context overflow (VLLMValidationError: input tokens … → lower max_input_tokens/max_output_tokens in the harbor yaml).agent/ populated (active) and result.json (done). 30+ min with zero agent/command-0/ (OpenHands) → setup stalled. Completions with n_output_tokens: None and agent_execution.finished_at ≈ started_at (instant-fail) = tunnel not carrying traffic despite a healthy-looking job.Quick liveness (≈15 min after submit): ssh <cluster> "squeue -u $USER --format='%.18i %.50j %.8T %.10M'" then tail the newest log — vLLM health-check pass, (Leonardo) SSH tunnel up, trial/reward lines, no OOM/repeated DaytonaErrors.
<run_tag>/<task>__<trial_id>/: config.json (mtime≈start, has api_base), trial.log, result.json (timestamps + verifier_result.rewards.reward + exception_info), exception.txt, agent/trajectory.json, verifier/{reward.txt,detailed_scores.json}. Eval cleanup + manual DB register + trace upload → eval-agentic-cleanup.
PermissionError: [Errno 13] at harbor/job.py … job_dir.mkdir() = a jobs_dir in the harbor config that another user owns. The canonical configs ship no jobs_dir; eval_harbor.sbatch passes --jobs-dir "$EVAL_JOBS_DIR" (per-user …/ot-baf/eval_jobs) which overrides the config. If you see this, confirm the sbatch has the --jobs-dir line; a hand-rolled harbor jobs start will reintroduce it. Resume is unaffected (takes -p $RUN_DIR).
A crashed eval leaves a non-terminal DB row blocking resubmission for 24h (reason=job in progress). After a crash the row stays started; the listener only resubmits started rows older than --stale-started-hours (default 24h, EVAL_LISTENER_STALE_HOURS). Pass a small value (e.g. --stale-started-hours 0.05 = 3 min) to force resubmit of the just-crashed attempt. Pending rows use --stale-pending-hours (default 6h, auto-cancels the stale SLURM job).
Jupiter: pass --reservation reformo or eval jobs starve behind RL (the reservation holds ~128 nodes while the general booster pool is empty). eval/jupiter/eval_harbor.sbatch sets --account reformo but no #SBATCH --reservation. Check scontrol show reservation for the live name/expiry before relying on it (the flag errors if the reservation is dead). Rescue already-PENDING jobs: scontrol update jobid=<j> reservation=reformo.
hosted_vllm/<org>/<model> evals need harbor commit 0f5a6e9e (allows 2-slash org-qualified names — validate_hosted_vllm_model_config in llms/utils.py) and model_info supplied via --agent-kwarg ({"max_input_tokens":…,"max_output_tokens":…,"input_cost_per_token":0,"output_cost_per_token":0}; token limits from the served vLLM max_model_len, costs 0 = self-hosted). Both are wired by default into the eval_harbor.sbatch files via EVAL_VLLM_MAX_MODEL_LEN (default 32768) + EVAL_MAX_OUTPUT_TOKENS (default 16384) (OT-Agent commit d0064011). If org-model evals fast-fail (~9 min, 0 POST 200s, 0 trajectories, all N trials raise identically), confirm those commits are in the cluster's harbor clone.
© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/eval-agentic-launch of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Eval Agentic Launch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Agentic Launch this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Community Evalssickn33/agentic-awesome-skills | 47k | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Aider DelegateamElnagdy/delegate-skills | 2.3k | 3 repos | ~3k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Diffusion Perf Optvllm-project/vllm-omni | 7.1k | — | ~7.5k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
sickn33/agentic-awesome-skills
Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
Categories
Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…. Eval Agentic Launch is an agent skill from open-thoughts/OpenThoughts-Agent.py / priority lists), wire the pinggy served-model tunnel, submit with the right preset + flags in tmux, then VERIFY the launch actually works via the 15-min infra sanity check (pinggy auth, Daytona→cluster apibase, vLLM POSTs, trial progression — catches "RUNNING but silently dead" jobs).
Eval Agentic Launch fits situations like: asked to launch/relaunch agentic evals; eval a model on a benchmark (terminalbench2 / devsetv2 / swebench / bfcl / aider).
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a claude-code`. Or copy the skill folder (.agents/skills/eval-agentic-launch in open-thoughts/OpenThoughts-Agent) into .claude/skills/eval-agentic-launch in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a codex`. Or copy the skill folder (.agents/skills/eval-agentic-launch in open-thoughts/OpenThoughts-Agent) into .agents/skills/eval-agentic-launch in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill eval-agentic-launch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-agentic-launch, .gemini/skills/eval-agentic-launch, .github/skills/eval-agentic-launch and .opencode/skills/eval-agentic-launch in your project.
Going by SKILL.md and its folder, Eval Agentic Launch needs the command-line tools its instructions call (python, conda and ssh) and credentials named DAYTONA_API_KEY, DAYTONA_DATA_API_KEY and SUPABASE_SERVICE_ROLE_KEY. Our summary lists: Python 3; A credential in DAYTONA_API_KEY; A credential in DAYTONA_DATA_API_KEY.
SKILL.md contains no URLs. Its commands use ssh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Agentic Launch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Agentic Launch: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Hugging Face Community Evals (sickn33/agentic-awesome-skills, 47k stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.