Hugging Face Local Models
huggingface/skills
Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.
A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…
$ npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Mesh-LLM/mesh-llm benchmark-tune --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchmark-tune .claude/skills/benchmark-tune && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-tune" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/benchmark-tune into .claude/skills/benchmark-tune/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-tune", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/benchmark-tuneType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Mesh-LLM/mesh-llm benchmark-tune --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/benchmark-tune .agents/skills/benchmark-tune && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-tune" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/benchmark-tune into .agents/skills/benchmark-tune/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-tune", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Mesh-LLM/mesh-llm benchmark-tune --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/benchmark-tune .cursor/skills/benchmark-tune && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-tune" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/benchmark-tune into .cursor/skills/benchmark-tune/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-tune", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Mesh-LLM/mesh-llm.git --path .agents/skills/benchmark-tune--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Mesh-LLM/mesh-llm benchmark-tune --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/benchmark-tune .gemini/skills/benchmark-tune && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-tune" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/benchmark-tune into .gemini/skills/benchmark-tune/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-tune", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Mesh-LLM/mesh-llm benchmark-tuneInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/benchmark-tune .github/skills/benchmark-tune && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-tune" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/benchmark-tune into .github/skills/benchmark-tune/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-tune", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Mesh-LLM/mesh-llm benchmark-tune --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/benchmark-tune .opencode/skills/benchmark-tune && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-tune" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/benchmark-tune into .opencode/skills/benchmark-tune/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-tune", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-tuneA skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…
Benchmark Tune is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps, running benchmark tune on local or SSH hosts, collecting JSON evidence, and applying tolerance-aware recommendations. Trigger for requests mentioning benchmark tune, tuning tok/s, ctxsize tradeoffs, mmap or mlock tuning, speculative decoding, MTP, ngram, draft models, or replacing old gpu tune usage.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with CUDA. The repository describes itself as: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit aaf5a6c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
justFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Tune loads about 1.6k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 654 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Mesh-LLM/mesh-llm at commit aaf5a6c, republished under its Apache-2.0 licence (© Mesh-LLM). 654 words, ~1,621 tokens.
.claude/skills/benchmark-tune/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use mesh-llm benchmark tune for model-serving throughput tuning. Do not use
mesh-llm gpu tune or mesh-llm gpus tune; the GPU namespace is for hardware
inventory and raw fingerprinting (mesh-llm gpus, mesh-llm gpus detect, and
hidden gpus run-benchmark).
Verify the command surface from the current checkout before long runs:
target/release/mesh-llm benchmark --help
target/release/mesh-llm benchmark tune --help
target/release/mesh-llm gpus --helpFor performance work, use a release build on the target host:
just release-buildOn NVIDIA remote hosts, verify that the selected CUDA native runtime is actually
in use before recording performance results. For Jetson/Orin-style aarch64 CUDA
hosts, build the normal backend-neutral product with a CUDA runtime, for example
just build backend=cuda cuda_arch=87, with the CUDA toolkit paths exported as
needed. The host itself must remain backend-neutral; a generic CPU runtime is
not valid performance evidence for GPU tune work. Confirm mesh-llm runtime list selects the intended CUDA runtime before benchmarking.
If the run is on a remote node over SSH and will take time, use the
remote-observable-process skill. Prefer a TTY/login shell and tee logs over
detached first attempts.
Benchmark tune accepts already-downloaded local/configured model targets only.
It will not fetch remote-only refs. If no explicit target is passed, it uses
configured local models from ~/.mesh-llm/config.toml.
Use one of:
mesh-llm benchmark tune --model /models/model.gguf
mesh-llm benchmark tune --models /models/a.gguf,/models/b.gguf
mesh-llm benchmark tuneStart with a bounded sweep, then expand around promising values:
mesh-llm benchmark tune \
--model /models/model.gguf \
--ctx-sizes 8192,32768,131072,262144 \
--batch-sizes 512,1024,2048 \
--ubatch-sizes 256,512,1024 \
--mmap-values auto,true,false \
--mlock-values false,true \
--speculative-types auto \
--throughput-tolerance-pct 10 \
--max-tokens 128 \
--debug-telemetry \
--jsonRules:
ubatch must be less than or equal to batch; invalid pairs are skipped.mmap and mlock are separate controls. Sweep them independently when
diagnosing load/runtime behavior.--mmap-values is omitted, tune tries auto, true, and false.--mlock-values is omitted, tune tries false and only tries true when
the current mlock probe says the evaluated budget can be locked.--speculative-types is omitted, tune uses auto: it tries
mtp first when the model target looks like an MTP model, tries
discovered local draft-model candidates when available, tries ngram
candidates as a model-free fallback, then includes a disabled baseline.--no-speculative-tune when you need to reproduce the older
fit-only/disabled-speculation behavior or isolate non-speculative regressions.--speculative-types mtp,draft,ngram,disabled to force an
explicit speculative sweep. draft requires either --spec-draft-models, a
configured draft_model_path, or a local sibling GGUF whose filename looks
like a draft/EAGLE model for the target.--spec-draft-max-tokens and
--spec-draft-min-tokens. Ngram sweeps use --spec-ngram-min and
--spec-ngram-max.--max-tokens when decode throughput is noisy; use shorter values
only for smoke checks.--throughput-tolerance-pct near the default 10 unless the user asks
for stricter raw throughput optimization.--debug-telemetry when you need proof that speculative decoding is
actually active. It runs trial children with Skippy debug telemetry mirrored
into target/gpu-tune/.../serve.log.Capture machine-readable output and trial logs:
mkdir -p target/benchmark-tune
mesh-llm benchmark tune ... --json \
| tee target/benchmark-tune/$(hostname)-$(date +%Y%m%d-%H%M%S).jsonFor remote hosts, include host, branch, commit, binary path, command, and output
path in the final report. Benchmark tune keeps per-trial logs under
target/gpu-tune/; inspect those logs when a trial fails or startup readiness
is slow.
Useful JSON fields:
benchmarks[].best: tolerance-aware recommendation.benchmarks[].raw_best: highest observed decode tok/s.benchmarks[].pareto_frontier: tradeoff set for decode tok/s vs ctx_size.benchmarks[].trials[].decode_tok_s: measured decode throughput.benchmarks[].trials[].candidate.speculative: speculative mode and settings
used for that isolated trial.benchmarks[].trials[].timings: lifecycle timing stats: setup_ms,
readiness_ms, request_ms, shutdown_ms, total_ms, and
readiness_attempts.benchmarks[].trials[].error and log_path: first stop for failures.Report both raw best and recommended settings. The recommendation is
tolerance-aware: candidates within --throughput-tolerance-pct of raw best are
treated as throughput-equivalent, then larger ctx_size is preferred.
Call out tradeoffs explicitly:
mmap or mlock changes the winner, report those controls separately.llama_stage.native_mtp.enabled, drafted/accepted/rejected counts, and
accept rate before concluding it is helping. Use --debug-telemetry if those
attributes are not present in the trial log.© Mesh-LLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .agents/skills/benchmark-tune of Mesh-LLM/mesh-llm.
Open the folder on GitHubat commit aaf5a6c
Benchmark Tune next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Tune this skillMesh-LLM/mesh-llm | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Modelshuggingface/skills | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | |
| Add Inference Backendintel/auto-round | 1.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Quark Onnx Quant Planamd/Quark | 181 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Kernel Microbenchmarkguqiong96/Lvllm | 465 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT |
huggingface/skills
Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.
intel/auto-round
Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).
amd/Quark
Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.
guqiong96/Lvllm
Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Mesh-LLM/mesh-llm
A skill your agent uses when validating a MeshLLM release candidate or current HEAD against the last GitHub release, assembling the canonical feature/fix/modification inventory, testing locally…
Mesh-LLM/mesh-llm
A skill your agent uses when adding, renaming, removing, validating, or exposing mesh-llm config settings, including built-in settings, plugin config schemas, owner-control apply behavior, CLI…
Mesh-LLM/mesh-llm
A skill your agent uses when connecting agent tools or OpenAI clients to mesh-llm — launching or configuring Goose, Claude Code, OpenCode, Pi, curl, or any OpenAI-compatible client against a local…
Mesh-LLM/mesh-llm
A skill your agent uses when converting Hugging Face SafeTensors checkpoints into split BF16 GGUF model repos with skippy-quantize on Hugging Face Jobs or a local machine, then publishing the…
Mesh-LLM/mesh-llm
A skill your agent uses when creating, monitoring, validating, or documenting low-memory Hugging Face Jobs or local runs that quantize split BF16/FP16 GGUF model repos into custom quant GGUF repos…
Mesh-LLM/mesh-llm
A skill your agent uses when changing mesh-llm automation or CLI flows that discover Hugging Face GGUF models, plan CPU Hugging Face Jobs for layer-package splitting, estimate max cost, or publish…
Works with
Categories
A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…. Benchmark Tune is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps, running benchmark tune on local or SSH hosts, collecting JSON evidence, and applying tolerance-aware recommendations.
Benchmark Tune fits situations like: documenting mesh-llm benchmark tune model-serving throughput trials; including choosing ctx/batch/ubatch/mmap/mlock/speculative-decoding sweeps; running benchmark tune on local; collecting JSON evidence.
Run `npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a claude-code`. Or copy the skill folder (.agents/skills/benchmark-tune in Mesh-LLM/mesh-llm) into .claude/skills/benchmark-tune in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a codex`. Or copy the skill folder (.agents/skills/benchmark-tune in Mesh-LLM/mesh-llm) into .agents/skills/benchmark-tune in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mesh-LLM/mesh-llm --skill benchmark-tune -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-tune, .gemini/skills/benchmark-tune, .github/skills/benchmark-tune and .opencode/skills/benchmark-tune in your project.
Going by SKILL.md and its folder, Benchmark Tune needs the command-line tools its instructions call (just).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Benchmark Tune is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmark Tune: Hugging Face Local Models (huggingface/skills, 11k stars), Add Inference Backend (intel/auto-round, 1.6k stars), Quark Onnx Quant Plan (amd/Quark, 181 stars) and Kernel Microbenchmark (guqiong96/Lvllm, 465 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Mesh-LLM (a GitHub organization) maintains it in Mesh-LLM/mesh-llm, which has 3,487 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.
Source: Mesh-LLM/mesh-llm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.