Agent skill

Skippy Bench

by Mesh-LLM in Mesh-LLM/mesh-llm

A skill your agent uses when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Skippy Bench

skills CLI
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mesh-LLM/mesh-llm skippy-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/skippy-bench .claude/skills/skippy-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skippy-bench
GitHub stars
3.5k
Token cost
~2.2k tokens
SKILL.md length
976 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.

  • Running benchmark orchestration
  • SKILL.md covers Current Repo Shape and External Agent Evals
  • Calls cargo, just and jq; reaches pypi.org; needs EVAL_LLM_API_KEY
  • Local single-stage

What it does

Skippy Bench is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. The repository describes itself as: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. The licence is Apache-2.0.

When your agent uses it

  • Running benchmark orchestration
  • Local single-stage
  • Split benchmarks
  • Benchmark report flow

Example prompts

  • “/skippy-bench”

Requirements

  • Python 3
  • Docker
  • A credential in EVAL_LLM_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit aaf5a6c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo
    • just
    • jq
    • python
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pypi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • EVAL_LLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skippy Bench loads about 2.2k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 976 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mesh-LLM/mesh-llm at commit aaf5a6c, republished under its Apache-2.0 licence (© Mesh-LLM). 976 words, ~2,168 tokens.

Download SKILL.mdSave it as .claude/skills/skippy-bench/SKILL.md (or your agent's skills folder).
name
skippy-bench
description
Use this skill when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.
metadata.short-description
Benchmark skippy stage runtime

skippy-bench

Use this skill for performance, orchestration, and report-oriented checks. Use skippy-correctness when the question is pass/fail exactness. All reportable benchmark runs need metrics-server. run, focused-runtime, and local-single start a collector by default; endpoint-driving commands such as chat-corpus and eval run require --metrics-http to point at an already-running metrics-server and should use --metrics-run-id matching the target endpoint's Skippy run id.

Benchmark-managed Skippy server runs must use a release skippy-serving build. Run just release-build before run, focused-runtime, local-single, or local split binary benchmarks, and use target/release/skippy-server (the SkippyBench default). Do not use target/debug/skippy-server for performance or full-corpus validation; SkippyBench rejects that path because debug builds can create false timeout and throughput failures.

Current Repo Shape

Standalone skippy-bench may not be present in this mesh checkout yet. Confirm available packages before using old source-repo commands:

bash
cargo metadata --no-deps --format-version 1 | jq -r '.packages[].name' | sort

Useful current checks:

bash
cargo test -p skippy-serving --lib
cargo test -p mesh-llm-host-runtime --lib inference::skippy

When benchmark harnesses are imported, keep reporting separate from request-path serving. Stage runtimes emit telemetry; benchmark/report tooling owns reports.

External Agent Evals

Use skippy-bench eval for external agent/coding benchmark harnesses. The local SkippyBench corpora are for runtime behavior, cache behavior, transport stress, and perf regression; they are not the source of agent benchmark claims.

Core pack:

bash
skippy-bench eval list
skippy-bench eval info terminal-bench
skippy-bench eval sync --pack core
skippy-bench eval doctor
skippy-bench eval run speed-bench \
  --base-url http://127.0.0.1:9337/v1 \
  --model org/repo:Q4_K_M \
  --endpoint-concurrency 1 \
  --metrics-http http://127.0.0.1:18080 \
  --metrics-run-id run-local-qwen

--timeout-secs is passed to the native harness as its request/task timeout where supported. It is not a full-run dataset limit. Use --harness-timeout-secs only when you need a hard wall-clock cap for an operator/debug run; omit it for canonical full-dataset validation. --endpoint-concurrency must match the target endpoint's serve-openai --generation-concurrency value. SkippyBench keeps native harness request concurrency equal to that value; adapter-specific request concurrency overrides such as SWE_BENCH_PRO_NUM_WORKERS and MCP_ATLAS_COMPLETION_CONCURRENCY must match it or eval run fails before starting the upstream harness. Do not run multiple LLM workers against a single-lane Skippy endpoint when validating full corpora.

Core eval ids:

  • speed-bench — llama.cpp SPEED-Bench client for OpenAI-compatible serving latency/throughput. Run the upstream qualitative benchmark across all categories with no Skippy-owned sample limit.
  • terminal-bench — Terminal-Bench CLI via terminal-bench-core==0.1.1.
  • swe-bench-pro — Scale SWE-Bench Pro OS repo; uses the upstream data and SWE-agent patch generation/evaluation flow rather than a Skippy-owned mini benchmark.
  • mcp-atlas — Scale MCP-Atlas native harness. eval run starts the MCP agent environment and completion service when their localhost ports are not already live, then runs the upstream completion script with --no-filter so all Hugging Face dataset rows are attempted, plus the upstream scoring path, without Skippy-specific task limits or tool_choice overrides.

Use-case routing:

NeedEvalWhy
OpenAI-compatible serving latency, tok/s, and full SPEED-Bench trafficspeed-benchNative SPEED-Bench client over the upstream dataset selection.
Terminal agent behavior, shell/task execution, Docker sandbox readinessterminal-benchExercises an agent loop that has to operate in a real terminal task environment.
Coding-agent patch generation and issue-resolution style promptsswe-bench-proUses upstream SWE-agent instance generation, patch gathering, and swe_bench_pro_eval.py.
MCP tool-use benchmark flowmcp-atlasUses upstream MCP-Atlas completion and scoring scripts with the full Hugging Face dataset.
Cache, runtime, transport, split, or mesh performance regressionBuilt-in SkippyBench run, focused-runtime, local-single, or chat-corpusThese are Skippy/runtime benchmarks, not external agent-quality claims.

Optional future packs are intentionally not wired yet:

  • repo-generation: NL2RepoBench.
  • tool-expanded: Toolathlon / Tool-Decathlon.

Keep sync/install opt-in. Do not make normal just build or cargo build download external harnesses, datasets, or Docker images. eval sync checks out the fetched upstream ref directly, and eval run records the resolved harness SHA as harness_commit in run.json. Preserve both behaviors so benchmark evidence remains reproducible even when definitions use floating upstream refs.

Terminal-Bench should be installed with uv tool install --python 3.12 terminal-bench; Python 3.14 currently breaks the tb Typer CLI. Treat Docker as ready only when skippy-bench eval doctor reports that the daemon can start a container; docker info alone is insufficient. skippy-bench eval run performs the same prerequisite checks before launching a native harness. Do not add Skippy-owned task filters, dataset limits, compatibility shims, response-format substitutions, or tool_choice overrides to external evals unless the user explicitly asks for a noncanonical experiment.

Show full SKILL.md (336 more words)Show less

For MCP-Atlas scoring, the wrapper defaults EVAL_LLM_MODEL, EVAL_LLM_BASE_URL, and EVAL_LLM_API_KEY to the same local endpoint/model used for completion, while preserving caller-provided EVAL_LLM_* overrides for judge-model runs. When validating with a very small local Skippy model, run completion against the normal compatibility endpoint and point EVAL_LLM_* at a separate strict structured-output scorer endpoint, for example a second skippy-serving serve-openai --guardrails enforce process; do not patch or post-process the MCP scorer. For resumed operator runs, set MCP_ATLAS_COMPLETION_OUTPUT_NAME to an existing upstream completion_results/*.csv basename so the native completion script can reuse its own processed-row skip behavior, and use MCP_ATLAS_SCORE_CONCURRENCY for the scorer's native --concurrency setting.

For SWE-Bench Pro, the wrapper defaults to the official Docker image namespace with local Docker deployment and local Docker evaluation so the core pack can run without Modal credentials. It still runs upstream helper_code/generate_sweagent_instances.py for the full dataset, then supplies SWE-agent with a native expert_file instance file for local Docker platform, entrypoint settings and SWE-agent's standalone Python/SWE-Rex Docker runtime. Local Docker runs install SWE-agent into a dedicated venv and default SWE_BENCH_PRO_SWEREX_SPEC to swe-rex[modal]==1.4.0, which keeps the native SWE-ReX Docker runtime but includes the upstream python:3.11.9-slim-bookworm builder fix. Modal remains an explicit environment override and uses the Scale SWE-ReX patch flow. Some official SWE-Pro base images point pip at an unavailable localhost package mirror; local Docker runs default SWE_BENCH_PRO_SWEREX_PIP_INDEX_URL to https://pypi.org/simple for the derived-image SWE-ReX install. Use SWE_BENCH_PRO_PARSE_FUNCTION=thought_action for local OpenAI-compatible models that do not emit OpenAI tool calls; this is the upstream SWE-agent local-model path, not a Skippy dataset or harness rewrite.

For TTFT/FTTT, use metrics-server correlation rather than harness-only timing. skippy-bench eval run and skippy-bench chat-corpus create/finalize a metrics-server run. eval run keeps harness success independent from a finalization/export failure and records telemetry as unavailable; chat-corpus still fails when its metrics report cannot be exported. The target endpoint must be emitting OTLP for the same run id. Debug telemetry is required for per-token spans such as stage.openai_decode_token; without it, the JSON report will still include a telemetry block explaining why TTFT/FTTT was unavailable.

© Mesh-LLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/skippy-bench of Mesh-LLM/mesh-llm.

Open the folder on GitHubat commit aaf5a6c

Compare with similar skills

Skippy Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skippy Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skippy Bench this skillMesh-LLM/mesh-llm3.5k—~2.2kAutomated safety check: PassApache-2.0
Agent BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.9k14 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 14 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed

More from Mesh-LLM/mesh-llm

All 25 skills in this repo
  • Release Validation

    Mesh-LLM/mesh-llm

    A skill your agent uses when validating a MeshLLM release candidate or current HEAD against the last GitHub release, assembling the canonical feature/fix/modification inventory, testing locally…

    3.5k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Benchmark Tune

    Mesh-LLM/mesh-llm

    A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

    3.5k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when adding, renaming, removing, validating, or exposing mesh-llm config settings, including built-in settings, plugin config schemas, owner-control apply behavior, CLI…

    3.5k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Connect Agents

    Mesh-LLM/mesh-llm

    A skill your agent uses when connecting agent tools or OpenAI clients to mesh-llm — launching or configuring Goose, Claude Code, OpenCode, Pi, curl, or any OpenAI-compatible client against a local…

    3.5k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when converting Hugging Face SafeTensors checkpoints into split BF16 GGUF model repos with skippy-quantize on Hugging Face Jobs or a local machine, then publishing the…

    3.5k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Hf Gguf Quant Jobs

    Mesh-LLM/mesh-llm

    A skill your agent uses when creating, monitoring, validating, or documenting low-memory Hugging Face Jobs or local runs that quantize split BF16/FP16 GGUF model repos into custom quant GGUF repos…

    3.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Questions about Skippy Bench

What does Skippy Bench do?

A skill your agent uses when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks. Skippy Bench is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.

When should I use Skippy Bench?

Skippy Bench fits situations like: running benchmark orchestration; local single-stage; split benchmarks; benchmark report flow.

How do I install Skippy Bench in Claude Code?

Run `npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a claude-code`. Or copy the skill folder (.agents/skills/skippy-bench in Mesh-LLM/mesh-llm) into .claude/skills/skippy-bench in your project. Claude Code loads it when a task matches its description.

How do I install Skippy Bench in Codex?

Run `npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a codex`. Or copy the skill folder (.agents/skills/skippy-bench in Mesh-LLM/mesh-llm) into .agents/skills/skippy-bench in your project. Codex loads it when a task matches its description.

Can I use Skippy Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skippy-bench, .gemini/skills/skippy-bench, .github/skills/skippy-bench and .opencode/skills/skippy-bench in your project.

What does Skippy Bench need to run?

Going by SKILL.md and its folder, Skippy Bench needs the command-line tools its instructions call (cargo, just, jq, python and docker) and credentials named EVAL_LLM_API_KEY. Our summary lists: Python 3; Docker; A credential in EVAL_LLM_API_KEY.

Does Skippy Bench access the network?

SKILL.md names 1 domain. In commands or code: pypi.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Skippy Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skippy Bench use?

Skippy Bench is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skippy Bench use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skippy Bench?

Skills that share tags, products or a category with Skippy Bench: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skippy Bench?

Mesh-LLM (a GitHub organization) maintains it in Mesh-LLM/mesh-llm, which has 3,487 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.

Source: Mesh-LLM/mesh-llm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.