Agent skill

Skippy Cache Family Bench

by Mesh-LLM in Mesh-LLM/mesh-llm

A skill your agent uses when benchmarking Skippy exact-prefix cache across model families, comparing Skippy against llama-server, producing README benchmark tables, updating…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Skippy Cache Family Bench

skills CLI
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-cache-family-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mesh-LLM/mesh-llm skippy-cache-family-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/skippy-cache-family-bench .claude/skills/skippy-cache-family-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skippy-cache-family-bench
GitHub stars
3.5k
Token cost
~764 tokens
SKILL.md length
207 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when benchmarking Skippy exact-prefix cache across model families, comparing Skippy against llama-server, producing README benchmark tables, updating…

  • Works in 5 steps: Run the reproducible wrapper. It builds… → For fast iteration after a build, skip… → If running the pieces manually, keep… → …
  • Benchmarking Skippy exact-prefix cache across model families
  • SKILL.md covers Workflow and Reporting Rules
  • Calls python3

What it does

Skippy Cache Family Bench is an agent skill from Mesh-LLM/mesh-llm. Use this skill when benchmarking Skippy exact-prefix cache across model families, comparing Skippy against llama-server, producing README benchmark tables, updating skippy/crates/skippy-cache/README.md evidence, or diagnosing cache benchmark gaps by family or Hugging Face use case.

Its SKILL.md is about 760 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets and Technical documentation. It works with Hugging Face and llama.cpp. The repository describes itself as: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. The licence is Apache-2.0.

When your agent uses it

  • Benchmarking Skippy exact-prefix cache across model families
  • Comparing Skippy against llama-server
  • Producing README benchmark tables
  • Updating skippy/crates/skippy-cache/README.md evidence

Example prompts

  • “/skippy-cache-family-bench”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Run the reproducible wrapper. It builds by default, then runs full-GGUF
  2. For fast iteration after a build, skip the build step
  3. If running the pieces manually, keep these matched settings for full-GGUF
  4. Run the use-case benchmark matrix against the same family set
  5. Render README-ready tables from the combined JSON outputs

What it can do on your machine

Read from SKILL.md and the folder at commit 1b9f0cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skippy Cache Family Bench loads about 764 tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 207 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~764

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mesh-LLM/mesh-llm at commit 1b9f0cf, republished under its Apache-2.0 licence (© Mesh-LLM). 207 words, ~764 tokens.

Download SKILL.mdSave it as .claude/skills/skippy-cache-family-bench/SKILL.md (or your agent's skills folder).
name
skippy-cache-family-bench
description
Use this skill when benchmarking Skippy exact-prefix cache across model families, comparing Skippy against llama-server, producing README benchmark tables, updating skippy/crates/skippy-cache/README.md evidence, or diagnosing cache benchmark gaps by family or Hugging Face use case.
metadata.short-description
Benchmark Skippy cache by family

skippy-cache-family-bench

Use this skill for reproducible Skippy cache benchmark evidence. The goal is to compare production cache payloads only: ResidentKv for dense families and KvRecurrent for recurrent/hybrid families. Do not report FullState as a production cache mode.

Workflow

  1. Run the reproducible wrapper. It builds by default, then runs full-GGUF baselines, Hugging Face use-case prompts, and the README report renderer:

    bash
    skippy/evals/skippy-cache-family-bench.sh /tmp/skippy-cache-family-bench
  2. For fast iteration after a build, skip the build step:

    bash
    SKIPPY_CACHE_SKIP_BUILD=1 skippy/evals/skippy-cache-family-bench.sh /tmp/skippy-cache-family-bench
  3. If running the pieces manually, keep these matched settings for full-GGUF family baselines:

    bash
    LLAMA_STAGE_BUILD_DIR=.deps/llama-build/build-stage-abi-cpu \
      python3 skippy/evals/skippy-cache-production-bench.py \
        --output-dir /tmp/skippy-cache-family-bench/full-gguf \
        --runtime-lane-count 1 \
        --llama-parallel 1 \
        --prefix-tokens 128
  4. Run the use-case benchmark matrix against the same family set:

    bash
    LLAMA_STAGE_BUILD_DIR=.deps/llama-build/build-stage-abi-cpu \
      python3 skippy/evals/skippy-cache-production-bench.py \
        --output-dir /tmp/skippy-cache-family-bench/use-cases \
        --runtime-lane-count 1 \
        --llama-parallel 1 \
        --prefix-tokens 128 \
        --use-case all
  5. Render README-ready tables from the combined JSON outputs:

    bash
    python3 skippy/evals/skippy-cache-family-report.py \
      --input /tmp/skippy-cache-family-bench/full-gguf/production-cache-bench.json \
      --input /tmp/skippy-cache-family-bench/use-cases/production-cache-bench.json \
      --output /tmp/skippy-cache-family-bench/readme-tables.md

Reporting Rules

  • Keep rows and columns ordered by related family: Qwen3Next, Falcon-H1, Llama, Qwen3 dense, DeepSeek2, GLM-4.7 Flash, GLM4, Gemma4 A4B, Gemma4 E4B, Gemma3, Gemma2, OLMo, MiniMax M2.7.
  • Always report Skippy versus llama-server for full-GGUF rows.
  • Keep DeepSeek3 in package-only evidence unless a machine can run a monolithic full-GGUF llama-server baseline.
  • Use one generated token and matched prefix tokens for apples-to-apples rows.
  • If a family fails correctness, leave the benchmark row out of promoted README evidence and explain the failure in skippy/crates/skippy-cache/TODO.md.
  • Preserve raw outputs under /tmp/... or another explicit run directory; do not paste ad hoc numbers without the backing production-cache-bench.json.

© Mesh-LLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/skippy-cache-family-bench of Mesh-LLM/mesh-llm.

Open the folder on GitHubat commit 1b9f0cf

Compare with similar skills

Skippy Cache Family Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skippy Cache Family Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skippy Cache Family Bench this skillMesh-LLM/mesh-llm3.5k—~764Automated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Huggingface LLM Trainerwaybarrios/opencode-power-pack534—~3kAutomated safety check: PassApache-2.0

Similar skills

  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Huggingface LLM Trainer

    waybarrios/opencode-power-pack

    Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

    534 GitHub stars~3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from Mesh-LLM/mesh-llm

All 25 skills in this repo
  • Release Validation

    Mesh-LLM/mesh-llm

    A skill your agent uses when validating a MeshLLM release candidate or current HEAD against the last GitHub release, assembling the canonical feature/fix/modification inventory, testing locally…

    3.5k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Benchmark Tune

    Mesh-LLM/mesh-llm

    A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

    3.5k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when adding, renaming, removing, validating, or exposing mesh-llm config settings, including built-in settings, plugin config schemas, owner-control apply behavior, CLI…

    3.5k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Connect Agents

    Mesh-LLM/mesh-llm

    A skill your agent uses when connecting agent tools or OpenAI clients to mesh-llm — launching or configuring Goose, Claude Code, OpenCode, Pi, curl, or any OpenAI-compatible client against a local…

    3.5k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when converting Hugging Face SafeTensors checkpoints into split BF16 GGUF model repos with skippy-quantize on Hugging Face Jobs or a local machine, then publishing the…

    3.5k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Hf Gguf Quant Jobs

    Mesh-LLM/mesh-llm

    A skill your agent uses when creating, monitoring, validating, or documenting low-memory Hugging Face Jobs or local runs that quantize split BF16/FP16 GGUF model repos into custom quant GGUF repos…

    3.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Skippy Cache Family Bench

What does Skippy Cache Family Bench do?

A skill your agent uses when benchmarking Skippy exact-prefix cache across model families, comparing Skippy against llama-server, producing README benchmark tables, updating…. Skippy Cache Family Bench is an agent skill from Mesh-LLM/mesh-llm.md evidence, or diagnosing cache benchmark gaps by family or Hugging Face use case.

When should I use Skippy Cache Family Bench?

Skippy Cache Family Bench fits situations like: benchmarking Skippy exact-prefix cache across model families; comparing Skippy against llama-server; producing README benchmark tables; updating skippy/crates/skippy-cache/README.md evidence.

How do I install Skippy Cache Family Bench in Claude Code?

Run `npx skills add Mesh-LLM/mesh-llm --skill skippy-cache-family-bench -a claude-code`. Or copy the skill folder (.agents/skills/skippy-cache-family-bench in Mesh-LLM/mesh-llm) into .claude/skills/skippy-cache-family-bench in your project. Claude Code loads it when a task matches its description.

How do I install Skippy Cache Family Bench in Codex?

Run `npx skills add Mesh-LLM/mesh-llm --skill skippy-cache-family-bench -a codex`. Or copy the skill folder (.agents/skills/skippy-cache-family-bench in Mesh-LLM/mesh-llm) into .agents/skills/skippy-cache-family-bench in your project. Codex loads it when a task matches its description.

Can I use Skippy Cache Family Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mesh-LLM/mesh-llm --skill skippy-cache-family-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skippy-cache-family-bench, .gemini/skills/skippy-cache-family-bench, .github/skills/skippy-cache-family-bench and .opencode/skills/skippy-cache-family-bench in your project.

What does Skippy Cache Family Bench need to run?

Going by SKILL.md and its folder, Skippy Cache Family Bench needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Skippy Cache Family Bench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skippy Cache Family Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skippy Cache Family Bench use?

Skippy Cache Family Bench is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skippy Cache Family Bench use?

About 764 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skippy Cache Family Bench?

Skills that share tags, products or a category with Skippy Cache Family Bench: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Hugging Face Local Models (huggingface/skills, 11k stars) and Add Model (guoqingbao/xinfer, 334 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skippy Cache Family Bench?

Mesh-LLM (a GitHub organization) maintains it in Mesh-LLM/mesh-llm, which has 3,495 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 11, 2026.

Source: Mesh-LLM/mesh-llm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.