Agent skill

Subwave LLM Bench

by perminder-klair in perminder-klair/subwave

Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

MITAuto-check: notesAI & LLM Engineering

Install Subwave LLM Bench

skills CLI
$ npx skills add perminder-klair/subwave --skill subwave-llm-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install perminder-klair/subwave subwave-llm-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/perminder-klair/subwave.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/subwave-llm-bench .claude/skills/subwave-llm-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
subwave-llm-bench
GitHub stars
1.4k
Token cost
~2.4k tokens
SKILL.md length
1,021 words
Files
2 (incl. scripts)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

  • Works in 2 steps: Routing matters as much as the model.… → The stress scenarios are where truth…
  • The user wants to assess
  • SKILL.md covers The two things that surprise…, When NOT to use this, Running llm-bench and Reading llm-bench results, plus 3 more sections
  • Runs Shell scripts from its folder; calls npm, docker and npx; needs OPENROUTER_API_KEY

What it does

Subwave LLM Bench is an agent skill from perminder-klair/subwave. Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using controller/scripts/llm-bench (the matrix harness) or the legacy picker-only controller/scripts/picker-test.mjs. Use this skill whenever the user wants to assess, benchmark, compare, or test which LLM model to run the station on — phrases like "which model should I use", "benchmark the picker / the DJ / this model", "test…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/assess-models.sh`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Ollama. The repository describes itself as: Personal internet radio: Agentic AI DJ. The licence is MIT.

When your agent uses it

  • The user wants to assess
  • Test which LLM model to run the station on — phrases like which model should I use
  • Benchmark the picker / the DJ / this model
  • Test these models

Example prompts

  • “which model should I use”
  • “benchmark the picker / the DJ / this model”
  • “test these models”
  • “/subwave-llm-bench”

Requirements

  • Node.js
  • A Bash shell
  • Docker
  • A credential in OPENROUTER_API_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Routing matters as much as the model. The same model can pass through
  2. The stress scenarios are where truth lives. Easy cells make every model

What it can do on your machine

Read from SKILL.md and the folder at commit 7dbc57c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • npm
    • docker
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, docker and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Subwave LLM Bench loads about 2.4k tokens when it runs. Until then it costs about 257 tokens; SKILL.md has 1,021 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~257
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:68
    from `state/secrets.env` and `controller/.env` automatically

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from perminder-klair/subwave at commit 7dbc57c, republished under its MIT licence (© perminder-klair). 1,021 words, ~2,365 tokens.

Download SKILL.mdSave it as .claude/skills/subwave-llm-bench/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
subwave-llm-bench
description
Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using controller/scripts/llm-bench (the matrix harness) or the legacy picker-only controller/scripts/picker-test.mjs. Use this skill whenever the user wants to assess, benchmark, compare, or test which LLM model to run the station on — phrases like "which model should I use", "benchmark the picker / the DJ / this model", "test these models", "is <model> good enough for the radio", "compare the ollama models", "will a small model work", "run llm-bench", "run picker-test", or when diagnosing slow/failing djAgentPick / djAgentSegment / generate* calls and the model choice is suspect. Trigger it even if the user doesn't name a harness — any request to evaluate model reliability or choose a model for SUB/WAVE belongs here. This skill only measures and recommends; it does NOT change the live station's configured model.

SUB/WAVE LLM model benchmark

Two harnesses, one job: measure how well a provider + model handles the station's real LLM calls before trusting it on air.

  • llm-bench (controller/scripts/llm-bench/, npm run llm-bench) — the primary harness. A matrix runner over every on-air call kind: pool picks (pickNextTrack), agent picks (djAgentPick), pool + agent segments (generateSegment / djAgentSegment), request matching (matchRequest / djAgentRequest), the free-text scripts (intro, link, station ID, hourly), banter, and the programme family (plan, beats, exchanges). Scores reliability + deterministic rule checks, prints a per-model comparison table, writes a diffable JSON report.
  • picker-test.mjs — the legacy picker-only deep-dive. Still useful for high-iteration picker runs comparable with historical results, and it has a bundled orchestration script (see the legacy section at the end).

Both import the live prompts and schemas from src/ (never copies), fake everything external (library, tools, weather/news data), and override provider/model only inside their own short-lived process — the live controller's configured model is never touched.

The two things that surprise people

  1. Routing matters as much as the model. The same model can pass through one provider and fail through another, because each @ai-sdk/* provider translates tools / structured output differently. Canonical case: deepseek-v4-flash scored 0/4 via the deepseek direct provider but 4/4 via openrouter. Always benchmark through the routing you'll actually deploy — "is model X good?" is the wrong question; "is provider+X good?" is the right one.
  2. The stress scenarios are where truth lives. Easy cells make every model look fine. The verdict cells are: djAgentPick/long-context (full prompt + tool loop), generateSegment/dull-weather (the model must decline to air — small models botch the silence encoding; a huge wall-clock here means the structured-output retry rescued a failed first attempt), and the multi-voice JSON kinds (generateBanter, generateProgrammeExchange).

When NOT to use this

  • Changing the live model — that's an admin Settings change, not this skill.
  • Diagnosing general runtime behaviour — subwave-log-analysis.
  • Starting/stopping the stack — subwave-control.

Running llm-bench

Runs on the host from a repo clone (it imports TS source via tsx; prod containers ship only compiled dist/). The running stack is not required — keys resolve from state/secrets.env and controller/.env automatically (existing env vars win). If a key exists only inside a prod container's env, copy it out first: export OPENROUTER_API_KEY=$(docker exec sub-wave-controller printenv OPENROUTER_API_KEY).

bash
cd <repo>/controller
npm run llm-bench -- \
  --models openrouter:google/gemma-4-31b-it,ollama:qwen3:8b \
  --iterations 3

Flags:

  • --models (required) — comma list of provider:model, split on the FIRST colon, so Ollama tags keep theirs (ollama:qwen3:8b). Providers: ollama | openai-compatible | anthropic | openai | google | deepseek | openrouter | requesty | gateway.
  • --kinds — groups (pick,segment,request,scripts,banter,programme) or exact kind names. Default: all.
  • --modes — pool,agent (default both). Mode-independent kinds run once either way. If the operator runs pool mode live (Agentic picker off), --modes pool is the honest benchmark — agent cells they'll never use just add noise and cost.
  • --iterations — per scenario. 1 to smoke, 3 to screen (default), 5+ to confirm a winner. Full matrix ≈ 32 scenario cells, so ~100 calls per model at 3 iterations.
  • --out — report path (default scripts/llm-bench/reports/<ts>.json, gitignored).

A full run takes many minutes (every call is a real model call, agent cells can take 45 s each). Run it in the background with a long timeout and don't poll; OLLAMA_URL=http://localhost:11434 translates the container-internal Ollama address for host runs.

Show full SKILL.md (508 more words)Show less

Reading llm-bench results

Each run is ok, violation (call succeeded, named rule checks failed), or thrown (bucketed: no-object-generated, timeout, unreachable, thrown). The summary prints pass% + p50 latency per cell, then per-model histograms of rule failures and thrown buckets.

Judge in this order:

  1. Thrown rate — a model that can't produce the schema at all is out. no-object-generated on structured kinds = weak structured output; everything landing at ~45 s = too slow for the loop, not incapable. unreachable = provider/config problem, not the model — fix and re-run.
  2. Violations — the named rules tell you the failure character: hallucinated-id (invents track ids — dangerous), banned-phrase:* / stage-direction:* / wrapping-quotes (would be read aloud by TTS), digits-in-spoken-time / clock-leak (hourly/link discipline), opener-repeat (ignores anti-repeat lists), unusable-exchange / single-voice (can't hold multi-voice JSON), variety:same-artist (editorial-pressure miss — informative, not disqualifying).
  3. Latency — p50 per cell; a pick has a whole track of slack so even 20 s works, but segments and links air sooner. Watch for retry-rescued cells (pass with 100×-median wall clock).
  4. Pool vs agent gap — a model that's clean in pool cells but falls apart in agent cells is a "candidate pool" model; that's a valid deployment, just say so in the recommendation.

Reports are JSON with every run record — keep old ones to diff a model across prompt changes or across model updates.

Producing the recommendation

  1. A compact table: one row per provider+model, columns for thrown rate, violation rate, p50 latency, and the dominant failure.
  2. A clear call naming the best provider + model + mode ("gemma via OpenRouter, pool mode — clean structured output, agent cells time out").
  3. Call out disqualifiers and deferrals (5xx throughout = re-test later).

Do not apply the change. The operator sets llm.provider + llm.model (and the Agentic picker toggle) in admin Settings — those are global, their call.

Legacy: picker-test.mjs deep-dive

For picker-only, high-iteration runs comparable with historical results:

bash
# DEV stack (src bind-mounted in container):
docker exec sub-wave-controller npx tsx scripts/picker-test.mjs <provider> <model> [iterations] [short|long]
# Host (works regardless of stack):
( cd <repo>/controller && STATE_DIR=<repo>/state OLLAMA_URL=http://localhost:11434 \
  npx tsx scripts/picker-test.mjs <provider> <model> [iterations] [short|long] )

long is the verdict mode (realistic session window); short is a sanity check. The bundled scripts/assess-models.sh (in this skill's directory) runs one or more models in both modes, auto-detects dev/prod, and summarises event- log failures: assess-models.sh <provider> [iterations] <model>... — with ollama and no models it discovers and tests everything installed.

Failure strings glossary (from the event log, kind: pickerTest):

  • agent did not call the done tool before stopping — model ignored the tool protocol. Check routing before rejecting the model (see surprise #1).
  • Every failure at ~agentTimeoutMs (45000) with tools=0 — a latency failure, not incapability; the deadline aborted it mid-loop.
  • Failed after N attempts … Internal Server Error — provider outage; re-run later, don't reject the model on this.
  • hallucinated-id / no-object-generated — weak structured output.

Cross-check live behaviour in /admin/debug → llm.recentCalls: the via field shows which path produced each result (ai-sdk:agent:native = clean native path; …:recovery = rescued by the done-tool recovery; a thrown djAgentPick falls back to the pool picker, so the station never goes silent).

Notes

  • Both harnesses append telemetry rows to the controller event log (state/logs/events-*.jsonl) and, if LLM_DEBUG_RAW is on, raw request bodies to state/logs/llm-debug.log. Harmless, filterable.
  • The docs page mapping every call kind to its code lives at docs/llm-calls.md; the harness design spec at docs/superpowers/specs/2026-07-09-llm-bench-design.md.

© perminder-klair, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .claude/skills/subwave-llm-bench of perminder-klair/subwave.

  • SKILL.md
  • scripts/assess-models.sh

Open the folder on GitHubat commit 7dbc57c

Compare with similar skills

Subwave LLM Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Subwave LLM Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Subwave LLM Bench this skillperminder-klair/subwave1.4k—~2.4kAutomated safety check: NotesMIT
Pii Safe Documentsdanyuchn/pii-guard249—~2.7kAutomated safety check: PassMIT
Visiongridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0
Domodomo Local AI Maintenancedarknecrocities/DomoDomo---All-in-one-Tool239—~17kAutomated safety check: PassNone
Cc Ollamamathruffian-dot/claude-code-lazy-packs254—~118Automated safety check: PassMIT
Bdistill Knowledge Extractionsickn33/agentic-awesome-skills47k2 repos~926Automated safety check: PassMIT

Similar skills

  • Pii Safe Documents

    danyuchn/pii-guard

    Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

    249 GitHub stars~2.7k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Domodomo Local AI Maintenance

    darknecrocities/DomoDomo---All-in-one-Tool

    Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.

    239 GitHub stars~17k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Cc Ollama

    mathruffian-dot/claude-code-lazy-packs

    Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.

    254 GitHub stars~118 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Bdistill Knowledge Extraction

    sickn33/agentic-awesome-skills

    Extract structured domain knowledge from AI models in-session or from local open-source models via Ollama.

    47k GitHub starsUsed in 2 repos~926 tokens
    AI & LLM EngineeringAuto-check passed
  • Pudu Task Telemetry

    davila7/claude-code-templates

    Measure local AI task latency, token usage, errors and verified outcomes using Pudu AI hardware evidence and installed Ollama models.

    33k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from perminder-klair/subwave

  • Subwave Worktree Dev

    perminder-klair/subwave

    Stage a SUB/WAVE git worktree so the dev stack can run from it, then start it.

    1.4k GitHub stars~2.4k tokensUpdated today
    Auto-check: notes
  • Subwave Control

    perminder-klair/subwave

    Start or stop the SUB/WAVE radio stack (a personal internet radio station) in dev or production mode — no builds, no rebuilds, no config rendering.

    1.4k GitHub stars~1.6k tokensUpdated today
    Auto-check: notes
  • Subwave Discord Release

    perminder-klair/subwave

    Draft a Discord release announcement for SUB/WAVE from a release PR, tag, or version number.

    1.4k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Subwave Release PR

    perminder-klair/subwave

    Open a release pull request from develop to main for SUB/WAVE, and keep develop from drifting behind main.

    1.4k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Subwave News Dispatch

    perminder-klair/subwave

    Write a news post for the SUB/WAVE "Dispatches" page, a short human-friendly tutorial about a feature, fix, or release.

    1.4k GitHub stars~1.6k tokensUpdated today
    Auto-check: warnings
  • Verify

    perminder-klair/subwave

    Drive a controller/admin-UI change end-to-end from a worktree without touching the live station — isolated controller on a spare port + temp STATEDIR, worktree Next dev server, Playwright against…

    1.4k GitHub stars~451 tokensUpdated today
    Auto-check passed

Works with

Questions about Subwave LLM Bench

What does Subwave LLM Bench do?

Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…. Subwave LLM Bench is an agent skill from perminder-klair/subwave.mjs.

When should I use Subwave LLM Bench?

Subwave LLM Bench fits situations like: the user wants to assess; test which LLM model to run the station on — phrases like which model should I use; benchmark the picker / the DJ / this model; test these models.

How do I install Subwave LLM Bench in Claude Code?

Run `npx skills add perminder-klair/subwave --skill subwave-llm-bench -a claude-code`. Or copy the skill folder (.claude/skills/subwave-llm-bench in perminder-klair/subwave) into .claude/skills/subwave-llm-bench in your project. Claude Code loads it when a task matches its description.

How do I install Subwave LLM Bench in Codex?

Run `npx skills add perminder-klair/subwave --skill subwave-llm-bench -a codex`. Or copy the skill folder (.claude/skills/subwave-llm-bench in perminder-klair/subwave) into .agents/skills/subwave-llm-bench in your project. Codex loads it when a task matches its description.

Can I use Subwave LLM Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add perminder-klair/subwave --skill subwave-llm-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/subwave-llm-bench, .gemini/skills/subwave-llm-bench, .github/skills/subwave-llm-bench and .opencode/skills/subwave-llm-bench in your project.

What does Subwave LLM Bench need to run?

Going by SKILL.md and its folder, Subwave LLM Bench needs a shell for the scripts in its folder, the command-line tools its instructions call (npm, docker and npx) and credentials named OPENROUTER_API_KEY. Our summary lists: Node.js; A Bash shell; Docker; A credential in OPENROUTER_API_KEY.

Does Subwave LLM Bench access the network?

SKILL.md contains no URLs. Its commands use npm, docker and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Subwave LLM Bench safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Subwave LLM Bench use?

Subwave LLM Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Subwave LLM Bench use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Subwave LLM Bench?

Skills that share tags, products or a category with Subwave LLM Bench: Pii Safe Documents (danyuchn/pii-guard, 249 stars), Vision (gridaco/grida, 2.7k stars), Domodomo Local AI Maintenance (darknecrocities/DomoDomo---All-in-one-Tool, 239 stars) and Cc Ollama (mathruffian-dot/claude-code-lazy-packs, 254 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Subwave LLM Bench?

perminder-klair (a GitHub user) maintains it in perminder-klair/subwave, which has 1,420 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 10, 2026.

Source: perminder-klair/subwave on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.