Agent skill

Vllm Prefix Cache Bench

by vllm-project in vllm-project/vllm-skills

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Prefix Cache Bench

skills CLI
$ npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-skills vllm-prefix-cache-bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-prefix-cache-bench .claude/skills/vllm-prefix-cache-bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-prefix-cache-bench
GitHub stars
103
Token cost
~1.4k tokens
SKILL.md length
507 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

  • The user asks to benchmark prefix caching hit rate
  • SKILL.md covers When to use, Option 1 (default). Fixed…, Option 2. ShareGPT Dataset… and Option 3. Prefix Repetition…, plus 3 more sections
  • Calls python3, pip and wget; reaches huggingface.co and github.com; needs HF_TOKEN
  • Caching efficiency

What it does

Vllm Prefix Cache Bench is an agent skill from vllm-project/vllm-skills. This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or repeated-prompt performance in vLLM.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Caching. It works with vLLM. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.

When your agent uses it

  • The user asks to benchmark prefix caching hit rate
  • Caching efficiency
  • Repeated-prompt performance in vLLM

Example prompts

  • “/vllm-prefix-cache-bench”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • pip
    • wget
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Prefix Cache Bench loads about 1.4k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 507 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 507 words, ~1,385 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-prefix-cache-bench/SKILL.md (or your agent's skills folder).
name
vllm-prefix-cache-bench
description
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or repeated-prompt performance in vLLM.

vLLM Prefix Caching Benchmark

Benchmark the efficiency of vLLM's automatic prefix caching (APC) feature. The offline script benchmarks/benchmark_prefix_caching.py runs directly against the vLLM engine (no server required). For online/serving tests, use vllm bench serve with the prefix_repetition dataset.

When to use

  • User wants to measure the performance impact of prefix caching for repeated or partially-shared prompts.
  • User wants to compare throughput/latency with and without --enable-prefix-caching.
  • User wants to test prefix caching using a fixed synthetic prompt, a real dataset (e.g. ShareGPT), or a synthetic prefix/suffix repetition pattern.

Option 1 (default). Fixed Prompt with Prefix Caching

Runs a synthetic benchmark with a fixed prompt repeated multiple times to directly measure cache hit efficiency. No dataset download required.

bash
python3 benchmarks/benchmark_prefix_caching.py \
  --model Qwen/Qwen3-8B \
  --enable-prefix-caching \
  --num-prompts 1 \
  --repeat-count 100 \
  --input-length-range 128:256

To compare against the baseline without caching:

bash
python3 benchmarks/benchmark_prefix_caching.py \
  --model Qwen/Qwen3-8B \
  --no-enable-prefix-caching \
  --num-prompts 1 \
  --repeat-count 100 \
  --input-length-range 128:256

Option 2. ShareGPT Dataset with Prefix Caching

Uses real-world conversational data from ShareGPT to evaluate prefix caching with naturally occurring prompt sharing.

First, download the dataset:

bash
wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json

Then run the benchmark:

bash
python3 benchmarks/benchmark_prefix_caching.py \
  --model Qwen/Qwen3-8B \
  --dataset-path ShareGPT_V3_unfiltered_cleaned_split.json \
  --enable-prefix-caching \
  --num-prompts 20 \
  --repeat-count 5 \
  --input-length-range 128:256

Option 3. Prefix Repetition Dataset (Online)

Uses vllm bench serve with the synthetic prefix_repetition dataset to test caching via the serving API. This requires a running vLLM server.

First, start the server:

bash
vllm serve Qwen/Qwen3-8B

Then run the benchmark:

bash
vllm bench serve \
  --backend openai \
  --model Qwen/Qwen3-8B \
  --dataset-name prefix_repetition \
  --num-prompts 100 \
  --prefix-repetition-prefix-len 512 \
  --prefix-repetition-suffix-len 128 \
  --prefix-repetition-num-prefixes 5 \
  --prefix-repetition-output-len 128

Key parameters for prefix_repetition:

ParameterDescription
--prefix-repetition-prefix-lenNumber of tokens in the shared prefix portion
--prefix-repetition-suffix-lenNumber of tokens in the unique suffix portion
--prefix-repetition-num-prefixesNumber of distinct prefixes to cycle through
--prefix-repetition-output-lenNumber of output tokens to generate per request

Notes

  • Run all commands from the root of the vLLM repository (cd vllm).
  • Keep the default model (Qwen/Qwen3-8B) unless the user specifies a different one or the model is unavailable; change only --model.
  • --repeat-count in Option 1 and 2 controls how many times each sampled prompt is replayed; higher values increase cache hit rate.
  • --input-length-range accepts a min:max token range, e.g. 128:256.
  • For multi-GPU setups, add --tensor-parallel-size <N>.
  • To test different hash algorithms for prefix caching internals, use --prefix-caching-hash-algo xxhash (requires pip install xxhash).
Show full SKILL.md (193 more words)Show less

Arguments for benchmark_prefix_caching.py

ArgumentRequiredDescription
--modelYesModel name or path (HuggingFace ID or local path)
--num-promptsYesNumber of prompts to process
--input-length-rangeYesToken length range for inputs, e.g. 128:256
--repeat-countNoNumber of times each prompt is repeated (default: 1)
--dataset-pathNoPath to a dataset file (e.g. ShareGPT JSON). Omit for synthetic fixed-prompt mode
--prefix-lenNoFixed prefix token length to prepend to every prompt
--output-lenNoNumber of output tokens to generate per request
--sortNoSort prompts by length before benchmarking
--enable-prefix-caching / --no-enable-prefix-cachingNoToggle APC (recommended: enable to test caching)
--prefix-caching-hash-algoNoHash algorithm: sha256, sha256_cbor, xxhash, xxhash_cbor
--tensor-parallel-sizeNoNumber of GPUs for tensor parallelism
--disable-detokenizeNoSkip detokenization to reduce overhead

Troubleshooting

  • If python3 benchmarks/*.py reports file not found, locate your local vLLM repository first and run the command from that repo root.
  • If you do not have the repository yet, clone it and continue:
bash
git clone https://github.com/vllm-project/vllm
cd vllm
  • If HuggingFace model download fails due to access restrictions, set your token: export HF_TOKEN=<your_token> or pass --hf-token <your_token>.
  • If xxhash or cbor2 is not installed and you use those hash algorithms, install them first: pip install xxhash cbor2.

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/vllm-skills/skills/vllm-prefix-cache-bench of vllm-project/vllm-skills.

Open the folder on GitHubat commit c996234

Compare with similar skills

Vllm Prefix Cache Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Prefix Cache Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Prefix Cache Bench this skillvllm-project/vllm-skills103—~1.4kAutomated safety check: PassApache-2.0
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Prefix Cache Replaybenchflow-ai/skillsbench1.8k—~2.3kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.0

Similar skills

  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Prefix Cache Replay

    benchflow-ai/skillsbench

    Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics.

    1.8k GitHub stars~2.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-skills

  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    103 GitHub stars~1.5k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    Auto-check: notes

Works with

Questions about Vllm Prefix Cache Bench

What does Vllm Prefix Cache Bench do?

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Vllm Prefix Cache Bench is an agent skill from vllm-project/vllm-skills. This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

When should I use Vllm Prefix Cache Bench?

Vllm Prefix Cache Bench fits situations like: the user asks to benchmark prefix caching hit rate; caching efficiency; repeated-prompt performance in vLLM.

How do I install Vllm Prefix Cache Bench in Claude Code?

Run `npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-prefix-cache-bench in vllm-project/vllm-skills) into .claude/skills/vllm-prefix-cache-bench in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Prefix Cache Bench in Codex?

Run `npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-prefix-cache-bench in vllm-project/vllm-skills) into .agents/skills/vllm-prefix-cache-bench in your project. Codex loads it when a task matches its description.

Can I use Vllm Prefix Cache Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-prefix-cache-bench, .gemini/skills/vllm-prefix-cache-bench, .github/skills/vllm-prefix-cache-bench and .opencode/skills/vllm-prefix-cache-bench in your project.

What does Vllm Prefix Cache Bench need to run?

Going by SKILL.md and its folder, Vllm Prefix Cache Bench needs the command-line tools its instructions call (python3, pip, wget and git) and credentials named HF_TOKEN. Our summary lists: Python 3.

Does Vllm Prefix Cache Bench access the network?

SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Vllm Prefix Cache Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Prefix Cache Bench use?

Vllm Prefix Cache Bench is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Prefix Cache Bench use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vllm Prefix Cache Bench?

Skills that share tags, products or a category with Vllm Prefix Cache Bench: SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Prefix Cache Replay (benchflow-ai/skillsbench, 1.8k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Prefix Cache Bench?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 103 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.

Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.