Agent skill

Benchmark Tune

by Mesh-LLM in Mesh-LLM/mesh-llm

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Benchmark Tune

skills CLI
$ npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mesh-LLM/mesh-llm benchmark-tune --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchmark-tune .claude/skills/benchmark-tune && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-tune
GitHub stars
3.5k
Token cost
~1.6k tokens
SKILL.md length
654 words
Files
2
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

  • Documenting mesh-llm benchmark tune model-serving throughput trials
  • SKILL.md covers Preflight, Targets, Candidate Sweep and Evidence, plus 1 more section
  • Calls just
  • Including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps

What it does

Benchmark Tune is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps, running benchmark tune on local or SSH hosts, collecting JSON evidence, and applying tolerance-aware recommendations. Trigger for requests mentioning benchmark tune, tuning tok/s, ctxsize tradeoffs, mmap or mlock tuning, speculative decoding, MTP, ngram, draft models, or replacing old gpu tune usage.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with CUDA. The repository describes itself as: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. The licence is Apache-2.0.

When your agent uses it

  • Documenting mesh-llm benchmark tune model-serving throughput trials
  • Including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps
  • Running benchmark tune on local
  • Collecting JSON evidence

Example prompts

  • “/benchmark-tune”

What it can do on your machine

Read from SKILL.md and the folder at commit aaf5a6c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • just

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Tune loads about 1.6k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 654 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mesh-LLM/mesh-llm at commit aaf5a6c, republished under its Apache-2.0 licence (© Mesh-LLM). 654 words, ~1,621 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-tune/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
benchmark-tune
description
Use this skill when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps, running benchmark tune on local or SSH hosts, collecting JSON evidence, and applying tolerance-aware recommendations. Trigger for requests mentioning benchmark tune, tuning tok/s, ctx_size tradeoffs, mmap or mlock tuning, speculative decoding, MTP, ngram, draft models, or replacing old gpu tune usage.

Benchmark Tune

Use mesh-llm benchmark tune for model-serving throughput tuning. Do not use mesh-llm gpu tune or mesh-llm gpus tune; the GPU namespace is for hardware inventory and raw fingerprinting (mesh-llm gpus, mesh-llm gpus detect, and hidden gpus run-benchmark).

Preflight

Verify the command surface from the current checkout before long runs:

bash
target/release/mesh-llm benchmark --help
target/release/mesh-llm benchmark tune --help
target/release/mesh-llm gpus --help

For performance work, use a release build on the target host:

bash
just release-build

On NVIDIA remote hosts, verify that the selected CUDA native runtime is actually in use before recording performance results. For Jetson/Orin-style aarch64 CUDA hosts, build the normal backend-neutral product with a CUDA runtime, for example just build backend=cuda cuda_arch=87, with the CUDA toolkit paths exported as needed. The host itself must remain backend-neutral; a generic CPU runtime is not valid performance evidence for GPU tune work. Confirm mesh-llm runtime list selects the intended CUDA runtime before benchmarking.

If the run is on a remote node over SSH and will take time, use the remote-observable-process skill. Prefer a TTY/login shell and tee logs over detached first attempts.

Targets

Benchmark tune accepts already-downloaded local/configured model targets only. It will not fetch remote-only refs. If no explicit target is passed, it uses configured local models from ~/.mesh-llm/config.toml.

Use one of:

bash
mesh-llm benchmark tune --model /models/model.gguf
mesh-llm benchmark tune --models /models/a.gguf,/models/b.gguf
mesh-llm benchmark tune

Candidate Sweep

Start with a bounded sweep, then expand around promising values:

bash
mesh-llm benchmark tune \
  --model /models/model.gguf \
  --ctx-sizes 8192,32768,131072,262144 \
  --batch-sizes 512,1024,2048 \
  --ubatch-sizes 256,512,1024 \
  --mmap-values auto,true,false \
  --mlock-values false,true \
  --speculative-types auto \
  --throughput-tolerance-pct 10 \
  --max-tokens 128 \
  --debug-telemetry \
  --json

Rules:

  • ubatch must be less than or equal to batch; invalid pairs are skipped.
  • mmap and mlock are separate controls. Sweep them independently when diagnosing load/runtime behavior.
  • If --mmap-values is omitted, tune tries auto, true, and false.
  • If --mlock-values is omitted, tune tries false and only tries true when the current mlock probe says the evaluated budget can be locked.
  • If --speculative-types is omitted, tune uses auto: it tries mtp first when the model target looks like an MTP model, tries discovered local draft-model candidates when available, tries ngram candidates as a model-free fallback, then includes a disabled baseline.
  • Use --no-speculative-tune when you need to reproduce the older fit-only/disabled-speculation behavior or isolate non-speculative regressions.
  • Use --speculative-types mtp,draft,ngram,disabled to force an explicit speculative sweep. draft requires either --spec-draft-models, a configured draft_model_path, or a local sibling GGUF whose filename looks like a draft/EAGLE model for the target.
  • MTP and draft sweeps use --spec-draft-max-tokens and --spec-draft-min-tokens. Ngram sweeps use --spec-ngram-min and --spec-ngram-max.
  • Use longer --max-tokens when decode throughput is noisy; use shorter values only for smoke checks.
  • Keep --throughput-tolerance-pct near the default 10 unless the user asks for stricter raw throughput optimization.
  • Add --debug-telemetry when you need proof that speculative decoding is actually active. It runs trial children with Skippy debug telemetry mirrored into target/gpu-tune/.../serve.log.
Show full SKILL.md (231 more words)Show less

Evidence

Capture machine-readable output and trial logs:

bash
mkdir -p target/benchmark-tune
mesh-llm benchmark tune ... --json \
  | tee target/benchmark-tune/$(hostname)-$(date +%Y%m%d-%H%M%S).json

For remote hosts, include host, branch, commit, binary path, command, and output path in the final report. Benchmark tune keeps per-trial logs under target/gpu-tune/; inspect those logs when a trial fails or startup readiness is slow.

Useful JSON fields:

  • benchmarks[].best: tolerance-aware recommendation.
  • benchmarks[].raw_best: highest observed decode tok/s.
  • benchmarks[].pareto_frontier: tradeoff set for decode tok/s vs ctx_size.
  • benchmarks[].trials[].decode_tok_s: measured decode throughput.
  • benchmarks[].trials[].candidate.speculative: speculative mode and settings used for that isolated trial.
  • benchmarks[].trials[].timings: lifecycle timing stats: setup_ms, readiness_ms, request_ms, shutdown_ms, total_ms, and readiness_attempts.
  • benchmarks[].trials[].error and log_path: first stop for failures.

Interpretation

Report both raw best and recommended settings. The recommendation is tolerance-aware: candidates within --throughput-tolerance-pct of raw best are treated as throughput-equivalent, then larger ctx_size is preferred.

Call out tradeoffs explicitly:

  • If raw best and recommended differ, explain the tok/s delta and context gain.
  • If mmap or mlock changes the winner, report those controls separately.
  • If speculative decoding changes the winner, report both tok/s and the active speculative candidate. For MTP, inspect trial logs/telemetry for llama_stage.native_mtp.enabled, drafted/accepted/rejected counts, and accept rate before concluding it is helping. Use --debug-telemetry if those attributes are not present in the trial log.
  • If all trials fail, summarize the shared failure reason and link the trial log paths rather than claiming no viable configuration exists.
  • If results are close, avoid overfitting decimals; prefer the setting with the better context or operational posture.

© Mesh-LLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/benchmark-tune of Mesh-LLM/mesh-llm.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit aaf5a6c

Compare with similar skills

Benchmark Tune next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Tune compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Tune this skillMesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0
Add Inference Backendintel/auto-round1.6k—~2.3kAutomated safety check: PassApache-2.0
Quark Onnx Quant Planamd/Quark181—~4.8kAutomated safety check: PassMIT
Kernel Microbenchmarkguqiong96/Lvllm4652 repos~1.5kAutomated safety check: PassApache-2.0
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT

Similar skills

  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Add Inference Backend

    intel/auto-round

    Official

    Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

    1.6k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.

    181 GitHub stars~4.8k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Kernel Microbenchmark

    guqiong96/Lvllm

    Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

    465 GitHub starsUsed in 2 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from Mesh-LLM/mesh-llm

All 25 skills in this repo
  • Release Validation

    Mesh-LLM/mesh-llm

    A skill your agent uses when validating a MeshLLM release candidate or current HEAD against the last GitHub release, assembling the canonical feature/fix/modification inventory, testing locally…

    3.5k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when adding, renaming, removing, validating, or exposing mesh-llm config settings, including built-in settings, plugin config schemas, owner-control apply behavior, CLI…

    3.5k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Connect Agents

    Mesh-LLM/mesh-llm

    A skill your agent uses when connecting agent tools or OpenAI clients to mesh-llm — launching or configuring Goose, Claude Code, OpenCode, Pi, curl, or any OpenAI-compatible client against a local…

    3.5k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when converting Hugging Face SafeTensors checkpoints into split BF16 GGUF model repos with skippy-quantize on Hugging Face Jobs or a local machine, then publishing the…

    3.5k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Hf Gguf Quant Jobs

    Mesh-LLM/mesh-llm

    A skill your agent uses when creating, monitoring, validating, or documenting low-memory Hugging Face Jobs or local runs that quantize split BF16/FP16 GGUF model repos into custom quant GGUF repos…

    3.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Hf Layer Package Jobs

    Mesh-LLM/mesh-llm

    A skill your agent uses when changing mesh-llm automation or CLI flows that discover Hugging Face GGUF models, plan CPU Hugging Face Jobs for layer-package splitting, estimate max cost, or publish…

    3.5k GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Questions about Benchmark Tune

What does Benchmark Tune do?

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…. Benchmark Tune is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps, running benchmark tune on local or SSH hosts, collecting JSON evidence, and applying tolerance-aware recommendations.

When should I use Benchmark Tune?

Benchmark Tune fits situations like: documenting mesh-llm benchmark tune model-serving throughput trials; including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps; running benchmark tune on local; collecting JSON evidence.

How do I install Benchmark Tune in Claude Code?

Run `npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a claude-code`. Or copy the skill folder (.agents/skills/benchmark-tune in Mesh-LLM/mesh-llm) into .claude/skills/benchmark-tune in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Tune in Codex?

Run `npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a codex`. Or copy the skill folder (.agents/skills/benchmark-tune in Mesh-LLM/mesh-llm) into .agents/skills/benchmark-tune in your project. Codex loads it when a task matches its description.

Can I use Benchmark Tune in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-tune, .gemini/skills/benchmark-tune, .github/skills/benchmark-tune and .opencode/skills/benchmark-tune in your project.

What does Benchmark Tune need to run?

Going by SKILL.md and its folder, Benchmark Tune needs the command-line tools its instructions call (just).

Does Benchmark Tune access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Tune safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Tune use?

Benchmark Tune is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Tune use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Tune?

Skills that share tags, products or a category with Benchmark Tune: Hugging Face Local Models (huggingface/skills, 11k stars), Add Inference Backend (intel/auto-round, 1.6k stars), Quark Onnx Quant Plan (amd/Quark, 181 stars) and Kernel Microbenchmark (guqiong96/Lvllm, 465 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Tune?

Mesh-LLM (a GitHub organization) maintains it in Mesh-LLM/mesh-llm, which has 3,487 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.

Source: Mesh-LLM/mesh-llm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.