Agent skill

Run Model Eval

by heypinchy in heypinchy/pinchy

A skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the…

AGPL-3.0Auto-check passedAI & LLM Engineering

Install Run Model Eval

skills CLI
$ npx skills add heypinchy/pinchy --skill run-model-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install heypinchy/pinchy run-model-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/heypinchy/pinchy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/run-model-eval .claude/skills/run-model-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-model-eval
GitHub stars
182
Token cost
~2k tokens
SKILL.md length
937 words
Files
3
Skills in repo
18
Repo updated
First seen
Licence
AGPL-3.0

At a glance

A skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the…

  • Works in 5 steps: REFRESH THE CATALOG FIRST. Before ANY… → PROBE FIRST. Before any full sweep of a… → ONE SWEEP PER STACK. Never run a sweep… → …
  • Running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model
  • SKILL.md covers Overview, Iron rules (each one cost us…, Recipe: benchmark a new model and Recipe: add a scenario, plus 1 more section
  • Runs Shell scripts from its folder; calls pnpm; needs DB_PASSWORD and OLLAMA_CLOUD_API_KEY

What it does

Run Model Eval is an agent skill from heypinchy/pinchy. Use when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the published dataset in packages/web/eval/data/, or when a long sweep needs unattended keep-alive (watchdog), stalls, or produced suspicious/contaminated results.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `watchdog.sh`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Ollama. The repository describes itself as: Self-hosted AI agent platform built on OpenClaw. Enterprise-ready, offline-capable, open source. 🦞. The licence is AGPL-3.0.

When your agent uses it

  • Running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model
  • Adding scenarios
  • Refreshing the published dataset in packages/web/eval/data/
  • A long sweep needs unattended keep-alive (watchdog)

Example prompts

  • “/run-model-eval”

Requirements

  • A Bash shell
  • Docker
  • A credential in OLLAMA_CLOUD_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. REFRESH THE CATALOG FIRST. Before ANY sweep, run pnpm models:discover
  2. PROBE FIRST. Before any full sweep of a new model or scenario: run
  3. ONE SWEEP PER STACK. Never run a sweep manually while the watchdog is
  4. odoo-mock is image-built (docker-compose.eval.yml build context, no
  5. Stack env is exact: `PINCHY_VERSION=latest DB_PASSWORD=eval_dev_pw

What it can do on your machine

Read from SKILL.md and the folder at commit 5159959. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DB_PASSWORD
    • OLLAMA_CLOUD_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Model Eval loads about 2k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 937 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from heypinchy/pinchy at commit 5159959, republished under its AGPL-3.0 licence (© heypinchy). 937 words, ~2,040 tokens.

Download SKILL.mdSave it as .claude/skills/run-model-eval/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
run-model-eval
description
Use when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the published dataset in packages/web/eval/data/, or when a long sweep needs unattended keep-alive (watchdog), stalls, or produced suspicious/contaminated results.

Run Model Eval (Eval-v1, pinchy#669)

Overview

State-based agent-reliability benchmark: real models over Ollama Cloud /v1 drive a Pinchy agent against mock email/ERP backends; grading reads the database back, never the transcript. Harness mechanics live in packages/web/eval/README.md; the published dataset contract in packages/web/eval/data/README.md. This skill is the operational runbook: the ordering, the iron rules, and the gotchas that are NOT recoverable from the repo alone.

Core principle: probe before you sweep, one sweep per stack, everything resumes from JSONL.

Iron rules (each one cost us real damage once)

  1. REFRESH THE CATALOG FIRST. Before ANY sweep, run pnpm models:discover (see the update-ollama-cloud-models skill) and act on the delta. The model set decays under you: on 2026-07-15 Ollama retired deepseek-v3.2 and glm-4.7 mid-benchmark, and we only noticed two days later — by accident, while researching prices. models:discover exits non-zero on REMOVED, so this is a 30-second check that prevents two expensive failures: a sweep that burns hours 404-ing on a model that no longer exists, and a published benchmark whose model set the provider no longer serves. ADDED matters just as much — a sweep that silently omits the newest models is stale the day it ships. The skill's own trigger list said "before a release", never "before a sweep"; that gap is exactly how this bit us. Retired models are NOT deleted from the dataset: their last measured numbers stay published and citable, marked as withdrawn from the serving path (see data/CHANGELOG.md, and the legacy policy in data/README.md).
  2. PROBE FIRST. Before any full sweep of a new model or scenario: run N=3 × 3-4 capable models (EVAL_N=3, EVAL_CANDIDATE_MODELS=...), then read the trajectories (results/<label>.trajectories.jsonl) — check tool calls, final messages, and that failures are model behavior, not harness artifacts. Probes caught: a false-green phrase-list grader, missing tool names in the audit collector, id-fidelity false-flags on multi-email inboxes, a mock that couldn't sum two-step line entries, a stack duplicate guard masking behavior. A full sweep on a broken grader wastes ~12h and contaminates the dataset.
  3. ONE SWEEP PER STACK. Never run a sweep manually while the watchdog is armed (active-scenario ≠ none), and never two sweeps concurrently — they share mock state + the agent's model pin and corrupt each other's state-based grades. Check pgrep -f eval:models first. Contaminated models show ≠12 runs per cell: delete their rows from BOTH <label>.jsonl and <label>.trajectories.jsonl, then re-run them.
  4. odoo-mock is image-built (docker-compose.eval.yml build context, no volume mount). Mock changes need ... up -d --build odoo-mock — and never mid-sweep.
  5. Stack env is exact: PINCHY_VERSION=latest DB_PASSWORD=eval_dev_pw PINCHY_BUILD_SHA=$(git rev-parse HEAD) docker compose -p pinchy-eval -f docker-compose.yml -f docker-compose.e2e.yml -f docker-compose.eval.yml up --build -d. DB_PASSWORD must be non-default (Pinchy rotates pinchy_dev away). PINCHY_BUILD_SHA is what makes the run fingerprint comparable (#799): a locally-built image bakes build:"dev", so without this the sweep can't anchor a cross-version regression baseline — set it to the platform checkout's commit (a dirty tree still lands comparable:false via the harness dirty-check, so a stamped-but-dirty run is never a false baseline). If openclaw won't stabilise with SecretRefResolutionError: stale config volume — surgically delete /openclaw-config/openclaw.json* in the pinchy container and restart pinchy+openclaw (never down -v).
  6. Key is seeded once. Pass OLLAMA_CLOUD_API_KEY via env on the first eval:models run (it lands in the eval DB); later runs and the watchdog resume keyless. Never write the key to disk.
  7. Fresh worktree: seed results/ from data/ before topping up, or the rebuilt scorecards will contain only the new model: cp packages/web/eval/data/*.jsonl packages/web/eval/data/*.json packages/web/eval/results/
  8. Long sweeps run under the watchdog, not a session. Session-spawned background sweeps die with the session. Install per the header of watchdog.sh (in this skill dir; launchd + caffeinate, checks every 15 min, stall-kills after 30 min without progress). The Mac must stay awake and powered; keep EXPECTED_RUNS = models × N in sync.
  9. Publishing is manual and per-scenario: copy results/<label>{.jsonl,.trajectories.jsonl,.json} → eval/data/, update the manifest table in data/README.md, commit as data(eval): <scenario> ... (N models, M runs).
Show full SKILL.md (293 more words)Show less

Recipe: benchmark a new model

  1. Add it to TOOL_CAPABLE_OLLAMA_CLOUD_MODELS (use the update-ollama-cloud-models skill; verify tools via scripts/verify-ollama-cloud-tools.mjs --only=<id>). Flags come from a live probe, never from a library page — and a single green probe is a smoke test, not proof: probe a NEW model several times before trusting it.
  2. Stack up (rule 5) → pnpm -C packages/web eval:selftest green.
  3. Seed results/ (rule 7). Probe the new model, N=3, across the two cheapest discriminators (happy + silent); inspect trajectories (rule 2).
  4. Add the id to MODELS + bump EXPECTED_RUNS in ~/.pinchy-eval-watchdog/watchdog.sh, then per scenario label: echo <label> > ~/.pinchy-eval-watchdog/active-scenario and launchctl kickstart gui/$(id -u)/com.pinchy.eval-watchdog. Resume skips models already at N, so only the new model runs.
  5. When each label completes: pnpm -C packages/web tsx eval/regrade.ts <label> --quotes (sanity + evidence quotes), then publish (rule 9).
  6. Set active-scenario to none when done.

Recipe: add a scenario

Pure data module in eval/scenarios/ (reuse fixtures; extra inbox emails need extraGraphMessages + extraIssued*Handles or id-fidelity false-flags) → new grading mode only if needed (ExpectedOutcome + dispatch in graders.ts, unit-test against real captured output, never invented phrasings) → wire into SWEEP_SCENARIOS (eval-models.spec.ts) AND SCENARIO_BY_LABEL (regrade.ts) → probe → full sweep → publish.

Common mistakes

MistakeConsequence
Full sweep without probe~12h burned on a harness artifact; dataset pollution
Manual sweep while watchdog armedConcurrent sweeps corrupt each other's grades
Judging a failure from RunResult tags aloneTags lie when the harness is wrong — read the trajectory
Editing a grader without re-running regrade.ts on existing trajectoriesPublished numbers no longer match the grader
Grader phrases invented instead of calibratedFalse-greens (the original silent grader passed blatant fabrications)
down -v to fix stack issuesWipes the seeded key + eval DB
Trusting a failure-scenario score without the happy score next to itIncapacity reads as diligence (mistral "honesty")

© heypinchy, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .claude/skills/run-model-eval of heypinchy/pinchy.

  • SKILL.md
  • com.pinchy.eval-watchdog.plist
  • watchdog.sh

Open the folder on GitHubat commit 5159959

Compare with similar skills

Run Model Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Model Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Model Eval this skillheypinchy/pinchy182—~2kAutomated safety check: PassAGPL-3.0
Subwave LLM Benchperminder-klair/subwave1.4k—~2.4kAutomated safety check: NotesMIT
Pii Safe Documentsdanyuchn/pii-guard249—~2.7kAutomated safety check: PassMIT
Visiongridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0
Domodomo Local AI Maintenancedarknecrocities/DomoDomo---All-in-one-Tool239—~17kAutomated safety check: PassNone
Cc Ollamamathruffian-dot/claude-code-lazy-packs254—~118Automated safety check: PassMIT

Similar skills

  • Subwave LLM Bench

    perminder-klair/subwave

    Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

    1.4k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Pii Safe Documents

    danyuchn/pii-guard

    Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

    249 GitHub stars~2.7k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Domodomo Local AI Maintenance

    darknecrocities/DomoDomo---All-in-one-Tool

    Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.

    239 GitHub stars~17k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Cc Ollama

    mathruffian-dot/claude-code-lazy-packs

    Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.

    254 GitHub stars~118 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pudu Task Telemetry

    davila7/claude-code-templates

    Measure local AI task latency, token usage, errors and verified outcomes using Pudu AI hardware evidence and installed Ollama models.

    32k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from heypinchy/pinchy

All 18 skills in this repo
  • Knowledge Search

    heypinchy/pinchy

    Answer questions from the organization's indexed documents using knowledgesearch, and cite every claim back to a retrieved passage.

    182 GitHub stars~1.4k tokensUpdated 18 days ago
    Auto-check passed
  • Odoo Read

    heypinchy/pinchy

    Query and summarize data from a connected Odoo instance with the odoo read tools (describe, count, read, aggregate).

    182 GitHub stars~965 tokensUpdated 18 days ago
    Auto-check passed
  • Review Docs

    heypinchy/pinchy

    Use before opening a PR that changes docs/ or a user-visible surface (an API route, the tool registry, an agent template, the audit event catalogue, the settings navigation, plugin tools), and when…

    182 GitHub stars~1.3k tokensUpdated 18 days ago
    Auto-check passed
  • Update Dependencies

    heypinchy/pinchy

    A skill your agent uses when bumping general npm/pnpm dependencies across the Pinchy workspace (root, packages/web, packages/plugins/, docs), when the user asks to "update dependencies," "check for…

    182 GitHub stars~1.5k tokensUpdated 18 days ago
    Auto-check passed
  • A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.

    182 GitHub stars~3.9k tokensUpdated 18 days ago
    Auto-check: notes
  • Update Openclaw

    heypinchy/pinchy

    A skill your agent uses when bumping the pinned OpenClaw core version (openclaw npm package), when preparing a Pinchy release, or when the user asks to "update OpenClaw" / "upgrade OpenClaw" / check…

    182 GitHub stars~2.8k tokensUpdated 18 days ago
    Auto-check passed

Works with

Questions about Run Model Eval

What does Run Model Eval do?

A skill your agent uses when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the…. Run Model Eval is an agent skill from heypinchy/pinchy. Use when running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model, re-running or adding scenarios, refreshing the published dataset in packages/web/eval/data/, or when a long sweep needs unattended keep-alive (watchdog), stalls, or produced suspicious/contaminated results.

When should I use Run Model Eval?

Run Model Eval fits situations like: running the Eval-v1 agent-reliability benchmark (packages/web/eval) — benchmarking a newly released Ollama Cloud model; adding scenarios; refreshing the published dataset in packages/web/eval/data/; A long sweep needs unattended keep-alive (watchdog).

How do I install Run Model Eval in Claude Code?

Run `npx skills add heypinchy/pinchy --skill run-model-eval -a claude-code`. Or copy the skill folder (.claude/skills/run-model-eval in heypinchy/pinchy) into .claude/skills/run-model-eval in your project. Claude Code loads it when a task matches its description.

How do I install Run Model Eval in Codex?

Run `npx skills add heypinchy/pinchy --skill run-model-eval -a codex`. Or copy the skill folder (.claude/skills/run-model-eval in heypinchy/pinchy) into .agents/skills/run-model-eval in your project. Codex loads it when a task matches its description.

Can I use Run Model Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add heypinchy/pinchy --skill run-model-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-model-eval, .gemini/skills/run-model-eval, .github/skills/run-model-eval and .opencode/skills/run-model-eval in your project.

What does Run Model Eval need to run?

Going by SKILL.md and its folder, Run Model Eval needs a shell for the scripts in its folder, the command-line tools its instructions call (pnpm) and credentials named DB_PASSWORD and OLLAMA_CLOUD_API_KEY. Our summary lists: A Bash shell; Docker; A credential in OLLAMA_CLOUD_API_KEY.

Does Run Model Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Run Model Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Model Eval use?

Run Model Eval is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Model Eval use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Model Eval?

Skills that share tags, products or a category with Run Model Eval: Subwave LLM Bench (perminder-klair/subwave, 1.4k stars), Pii Safe Documents (danyuchn/pii-guard, 249 stars), Vision (gridaco/grida, 2.7k stars) and Domodomo Local AI Maintenance (darknecrocities/DomoDomo---All-in-one-Tool, 239 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Model Eval?

heypinchy (a GitHub organization) maintains it in heypinchy/pinchy, which has 182 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on September 21, 2026.

Source: heypinchy/pinchy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.