Subwave LLM Bench
perminder-klair/subwave
Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…
A skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the…
$ npx skills add heypinchy/pinchy --skill run-model-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install heypinchy/pinchy run-model-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/heypinchy/pinchy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/run-model-eval .claude/skills/run-model-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run-model-eval" agent skill from https://github.com/heypinchy/pinchy/tree/main/.claude/skills/run-model-eval into .claude/skills/run-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-model-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/heypinchy/pinchy/tree/main/.claude/skills/run-model-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add heypinchy/pinchy --skill run-model-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install heypinchy/pinchy run-model-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heypinchy/pinchy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/run-model-eval .agents/skills/run-model-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run-model-eval" agent skill from https://github.com/heypinchy/pinchy/tree/main/.claude/skills/run-model-eval into .agents/skills/run-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-model-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add heypinchy/pinchy --skill run-model-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install heypinchy/pinchy run-model-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heypinchy/pinchy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/run-model-eval .cursor/skills/run-model-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run-model-eval" agent skill from https://github.com/heypinchy/pinchy/tree/main/.claude/skills/run-model-eval into .cursor/skills/run-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-model-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/heypinchy/pinchy.git --path .claude/skills/run-model-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add heypinchy/pinchy --skill run-model-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install heypinchy/pinchy run-model-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heypinchy/pinchy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/run-model-eval .gemini/skills/run-model-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run-model-eval" agent skill from https://github.com/heypinchy/pinchy/tree/main/.claude/skills/run-model-eval into .gemini/skills/run-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-model-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install heypinchy/pinchy run-model-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add heypinchy/pinchy --skill run-model-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/heypinchy/pinchy.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/run-model-eval .github/skills/run-model-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run-model-eval" agent skill from https://github.com/heypinchy/pinchy/tree/main/.claude/skills/run-model-eval into .github/skills/run-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-model-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add heypinchy/pinchy --skill run-model-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install heypinchy/pinchy run-model-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heypinchy/pinchy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/run-model-eval .opencode/skills/run-model-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run-model-eval" agent skill from https://github.com/heypinchy/pinchy/tree/main/.claude/skills/run-model-eval into .opencode/skills/run-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-model-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
run-model-evalA skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the…
Run Model Eval is an agent skill from heypinchy/pinchy. Use when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the published dataset in packages/web/eval/data/, or when a long sweep needs unattended keep-alive (watchdog), stalls, or produced suspicious/contaminated results.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `watchdog.sh`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Ollama. The repository describes itself as: Self-hosted AI agent platform built on OpenClaw. Enterprise-ready, offline-capable, open source. 🦞. The licence is AGPL-3.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5159959. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
pnpmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DB_PASSWORDOLLAMA_CLOUD_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Run Model Eval loads about 2k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 937 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from heypinchy/pinchy at commit 5159959, republished under its AGPL-3.0 licence (© heypinchy). 937 words, ~2,040 tokens.
.claude/skills/run-model-eval/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.State-based agent-reliability benchmark: real models over Ollama Cloud /v1
drive a Pinchy agent against mock email/ERP backends; grading reads the
database back, never the transcript. Harness mechanics live in
packages/web/eval/README.md; the published dataset contract in
packages/web/eval/data/README.md. This skill is the operational runbook:
the ordering, the iron rules, and the gotchas that are NOT recoverable from the
repo alone.
Core principle: probe before you sweep, one sweep per stack, everything resumes from JSONL.
pnpm models:discover
(see the update-ollama-cloud-models skill) and act on the delta. The model
set decays under you: on 2026-07-15 Ollama retired deepseek-v3.2 and
glm-4.7 mid-benchmark, and we only noticed two days later — by accident,
while researching prices. models:discover exits non-zero on REMOVED, so
this is a 30-second check that prevents two expensive failures: a sweep that
burns hours 404-ing on a model that no longer exists, and a published
benchmark whose model set the provider no longer serves. ADDED matters just
as much — a sweep that silently omits the newest models is stale the day it
ships. The skill's own trigger list said "before a release", never "before a
sweep"; that gap is exactly how this bit us.
Retired models are NOT deleted from the dataset: their last measured numbers
stay published and citable, marked as withdrawn from the serving path (see
data/CHANGELOG.md, and the legacy policy in data/README.md).EVAL_N=3, EVAL_CANDIDATE_MODELS=...), then
read the trajectories (results/<label>.trajectories.jsonl) — check
tool calls, final messages, and that failures are model behavior, not
harness artifacts. Probes caught: a false-green phrase-list grader, missing
tool names in the audit collector, id-fidelity false-flags on multi-email
inboxes, a mock that couldn't sum two-step line entries, a stack duplicate
guard masking behavior. A full sweep on a broken grader wastes ~12h and
contaminates the dataset.active-scenario ≠ none), and never two sweeps concurrently —
they share mock state + the agent's model pin and corrupt each other's
state-based grades. Check pgrep -f eval:models first. Contaminated
models show ≠12 runs per cell: delete their rows from BOTH
<label>.jsonl and <label>.trajectories.jsonl, then re-run them.docker-compose.eval.yml build context, no
volume mount). Mock changes need
... up -d --build odoo-mock — and never mid-sweep.PINCHY_VERSION=latest DB_PASSWORD=eval_dev_pw PINCHY_BUILD_SHA=$(git rev-parse HEAD) docker compose -p pinchy-eval -f docker-compose.yml -f docker-compose.e2e.yml -f docker-compose.eval.yml up --build -d. DB_PASSWORD must be non-default (Pinchy rotates pinchy_dev
away). PINCHY_BUILD_SHA is what makes the run fingerprint comparable
(#799): a locally-built image bakes build:"dev", so without this the sweep
can't anchor a cross-version regression baseline — set it to the platform
checkout's commit (a dirty tree still lands comparable:false via the
harness dirty-check, so a stamped-but-dirty run is never a false baseline).
If openclaw won't stabilise with SecretRefResolutionError: stale config
volume — surgically delete /openclaw-config/openclaw.json* in the pinchy
container and restart pinchy+openclaw (never down -v).OLLAMA_CLOUD_API_KEY via env on the first
eval:models run (it lands in the eval DB); later runs and the watchdog
resume keyless. Never write the key to disk.results/ from data/ before topping up, or the
rebuilt scorecards will contain only the new model:
cp packages/web/eval/data/*.jsonl packages/web/eval/data/*.json packages/web/eval/results/watchdog.sh (in this skill dir; launchd + caffeinate, checks every
15 min, stall-kills after 30 min without progress). The Mac must stay
awake and powered; keep EXPECTED_RUNS = models × N in sync.results/<label>{.jsonl,.trajectories.jsonl,.json} → eval/data/, update
the manifest table in data/README.md, commit as
data(eval): <scenario> ... (N models, M runs).TOOL_CAPABLE_OLLAMA_CLOUD_MODELS (use the
update-ollama-cloud-models skill; verify tools via
scripts/verify-ollama-cloud-tools.mjs --only=<id>). Flags come from a live
probe, never from a library page — and a single green probe is a smoke test,
not proof: probe a NEW model several times before trusting it.pnpm -C packages/web eval:selftest green.results/ (rule 7). Probe the new model, N=3, across the two
cheapest discriminators (happy + silent); inspect trajectories (rule 2).MODELS + bump EXPECTED_RUNS in
~/.pinchy-eval-watchdog/watchdog.sh, then per scenario label:
echo <label> > ~/.pinchy-eval-watchdog/active-scenario and
launchctl kickstart gui/$(id -u)/com.pinchy.eval-watchdog. Resume skips
models already at N, so only the new model runs.pnpm -C packages/web tsx eval/regrade.ts <label> --quotes (sanity + evidence quotes), then publish (rule 9).active-scenario to none when done.Pure data module in eval/scenarios/ (reuse fixtures; extra inbox emails need
extraGraphMessages + extraIssued*Handles or id-fidelity false-flags) → new
grading mode only if needed (ExpectedOutcome + dispatch in graders.ts,
unit-test against real captured output, never invented phrasings) → wire into
SWEEP_SCENARIOS (eval-models.spec.ts) AND SCENARIO_BY_LABEL (regrade.ts) →
probe → full sweep → publish.
| Mistake | Consequence |
|---|---|
| Full sweep without probe | ~12h burned on a harness artifact; dataset pollution |
| Manual sweep while watchdog armed | Concurrent sweeps corrupt each other's grades |
Judging a failure from RunResult tags alone | Tags lie when the harness is wrong — read the trajectory |
Editing a grader without re-running regrade.ts on existing trajectories | Published numbers no longer match the grader |
| Grader phrases invented instead of calibrated | False-greens (the original silent grader passed blatant fabrications) |
down -v to fix stack issues | Wipes the seeded key + eval DB |
| Trusting a failure-scenario score without the happy score next to it | Incapacity reads as diligence (mistral "honesty") |
© heypinchy, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in .claude/skills/run-model-eval of heypinchy/pinchy.
Open the folder on GitHubat commit 5159959
Run Model Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Run Model Eval this skillheypinchy/pinchy | 182 | — | ~2k | Automated safety check: Pass | AGPL-3.0 | |
| Subwave LLM Benchperminder-klair/subwave | 1.4k | — | ~2.4k | Automated safety check: Notes | MIT | |
| Pii Safe Documentsdanyuchn/pii-guard | 249 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Visiongridaco/grida | 2.7k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Domodomo Local AI Maintenancedarknecrocities/DomoDomo---All-in-one-Tool | 239 | — | ~17k | Automated safety check: Pass | None | |
| Cc Ollamamathruffian-dot/claude-code-lazy-packs | 254 | — | ~118 | Automated safety check: Pass | MIT |
perminder-klair/subwave
Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…
danyuchn/pii-guard
Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.
gridaco/grida
Query images with a local Ollama vision model without loading the image into the main agent context.
darknecrocities/DomoDomo---All-in-one-Tool
Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.
mathruffian-dot/claude-code-lazy-packs
Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.
davila7/claude-code-templates
Measure local AI task latency, token usage, errors and verified outcomes using Pudu AI hardware evidence and installed Ollama models.
heypinchy/pinchy
Answer questions from the organization's indexed documents using knowledgesearch, and cite every claim back to a retrieved passage.
heypinchy/pinchy
Query and summarize data from a connected Odoo instance with the odoo read tools (describe, count, read, aggregate).
heypinchy/pinchy
Use before opening a PR that changes docs/ or a user-visible surface (an API route, the tool registry, an agent template, the audit event catalogue, the settings navigation, plugin tools), and when…
heypinchy/pinchy
A skill your agent uses when bumping general npm/pnpm dependencies across the Pinchy workspace (root, packages/web, packages/plugins/, docs), when the user asks to "update dependencies," "check for…
heypinchy/pinchy
A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.
heypinchy/pinchy
A skill your agent uses when bumping the pinned OpenClaw core version (openclaw npm package), when preparing a Pinchy release, or when the user asks to "update OpenClaw" / "upgrade OpenClaw" / check…
Works with
Categories
A skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the…. Run Model Eval is an agent skill from heypinchy/pinchy. Use when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the published dataset in packages/web/eval/data/, or when a long sweep needs unattended keep-alive (watchdog), stalls, or produced suspicious/contaminated results.
Run Model Eval fits situations like: running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model; adding scenarios; refreshing the published dataset in packages/web/eval/data/; A long sweep needs unattended keep-alive (watchdog).
Run `npx skills add heypinchy/pinchy --skill run-model-eval -a claude-code`. Or copy the skill folder (.claude/skills/run-model-eval in heypinchy/pinchy) into .claude/skills/run-model-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add heypinchy/pinchy --skill run-model-eval -a codex`. Or copy the skill folder (.claude/skills/run-model-eval in heypinchy/pinchy) into .agents/skills/run-model-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add heypinchy/pinchy --skill run-model-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-model-eval, .gemini/skills/run-model-eval, .github/skills/run-model-eval and .opencode/skills/run-model-eval in your project.
Going by SKILL.md and its folder, Run Model Eval needs a shell for the scripts in its folder, the command-line tools its instructions call (pnpm) and credentials named DB_PASSWORD and OLLAMA_CLOUD_API_KEY. Our summary lists: A Bash shell; Docker; A credential in OLLAMA_CLOUD_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Run Model Eval is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Run Model Eval: Subwave LLM Bench (perminder-klair/subwave, 1.4k stars), Pii Safe Documents (danyuchn/pii-guard, 249 stars), Vision (gridaco/grida, 2.7k stars) and Domodomo Local AI Maintenance (darknecrocities/DomoDomo---All-in-one-Tool, 239 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
heypinchy (a GitHub organization) maintains it in heypinchy/pinchy, which has 182 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on September 21, 2026.
Source: heypinchy/pinchy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.