SGLang Structured Serving
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.
$ npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-skills vllm-prefix-cache-bench --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-prefix-cache-bench .claude/skills/vllm-prefix-cache-bench && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-prefix-cache-bench" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-prefix-cache-bench into .claude/skills/vllm-prefix-cache-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-prefix-cache-bench", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-prefix-cache-benchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-skills vllm-prefix-cache-bench --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-prefix-cache-bench .agents/skills/vllm-prefix-cache-bench && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-prefix-cache-bench" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-prefix-cache-bench into .agents/skills/vllm-prefix-cache-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-prefix-cache-bench", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-skills vllm-prefix-cache-bench --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-prefix-cache-bench .cursor/skills/vllm-prefix-cache-bench && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-prefix-cache-bench" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-prefix-cache-bench into .cursor/skills/vllm-prefix-cache-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-prefix-cache-bench", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-skills.git --path plugins/vllm-skills/skills/vllm-prefix-cache-bench--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-skills vllm-prefix-cache-bench --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-prefix-cache-bench .gemini/skills/vllm-prefix-cache-bench && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-prefix-cache-bench" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-prefix-cache-bench into .gemini/skills/vllm-prefix-cache-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-prefix-cache-bench", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-skills vllm-prefix-cache-benchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-prefix-cache-bench .github/skills/vllm-prefix-cache-bench && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-prefix-cache-bench" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-prefix-cache-bench into .github/skills/vllm-prefix-cache-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-prefix-cache-bench", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-skills vllm-prefix-cache-bench --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-prefix-cache-bench .opencode/skills/vllm-prefix-cache-bench && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-prefix-cache-bench" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-prefix-cache-bench into .opencode/skills/vllm-prefix-cache-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-prefix-cache-bench", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-prefix-cache-benchThis is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.
Vllm Prefix Cache Bench is an agent skill from vllm-project/vllm-skills. This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or repeated-prompt performance in vLLM.
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Caching. It works with vLLM. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3pipwgetgitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.cogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm Prefix Cache Bench loads about 1.4k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 507 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 507 words, ~1,385 tokens.
.claude/skills/vllm-prefix-cache-bench/SKILL.md (or your agent's skills folder).Benchmark the efficiency of vLLM's automatic prefix caching (APC) feature. The offline script benchmarks/benchmark_prefix_caching.py runs directly against the vLLM engine (no server required). For online/serving tests, use vllm bench serve with the prefix_repetition dataset.
--enable-prefix-caching.Runs a synthetic benchmark with a fixed prompt repeated multiple times to directly measure cache hit efficiency. No dataset download required.
python3 benchmarks/benchmark_prefix_caching.py \
--model Qwen/Qwen3-8B \
--enable-prefix-caching \
--num-prompts 1 \
--repeat-count 100 \
--input-length-range 128:256To compare against the baseline without caching:
python3 benchmarks/benchmark_prefix_caching.py \
--model Qwen/Qwen3-8B \
--no-enable-prefix-caching \
--num-prompts 1 \
--repeat-count 100 \
--input-length-range 128:256Uses real-world conversational data from ShareGPT to evaluate prefix caching with naturally occurring prompt sharing.
First, download the dataset:
wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.jsonThen run the benchmark:
python3 benchmarks/benchmark_prefix_caching.py \
--model Qwen/Qwen3-8B \
--dataset-path ShareGPT_V3_unfiltered_cleaned_split.json \
--enable-prefix-caching \
--num-prompts 20 \
--repeat-count 5 \
--input-length-range 128:256Uses vllm bench serve with the synthetic prefix_repetition dataset to test caching via the serving API. This requires a running vLLM server.
First, start the server:
vllm serve Qwen/Qwen3-8BThen run the benchmark:
vllm bench serve \
--backend openai \
--model Qwen/Qwen3-8B \
--dataset-name prefix_repetition \
--num-prompts 100 \
--prefix-repetition-prefix-len 512 \
--prefix-repetition-suffix-len 128 \
--prefix-repetition-num-prefixes 5 \
--prefix-repetition-output-len 128Key parameters for prefix_repetition:
| Parameter | Description |
|---|---|
--prefix-repetition-prefix-len | Number of tokens in the shared prefix portion |
--prefix-repetition-suffix-len | Number of tokens in the unique suffix portion |
--prefix-repetition-num-prefixes | Number of distinct prefixes to cycle through |
--prefix-repetition-output-len | Number of output tokens to generate per request |
cd vllm).Qwen/Qwen3-8B) unless the user specifies a different one or the model is unavailable; change only --model.--repeat-count in Option 1 and 2 controls how many times each sampled prompt is replayed; higher values increase cache hit rate.--input-length-range accepts a min:max token range, e.g. 128:256.--tensor-parallel-size <N>.--prefix-caching-hash-algo xxhash (requires pip install xxhash).benchmark_prefix_caching.py| Argument | Required | Description |
|---|---|---|
--model | Yes | Model name or path (HuggingFace ID or local path) |
--num-prompts | Yes | Number of prompts to process |
--input-length-range | Yes | Token length range for inputs, e.g. 128:256 |
--repeat-count | No | Number of times each prompt is repeated (default: 1) |
--dataset-path | No | Path to a dataset file (e.g. ShareGPT JSON). Omit for synthetic fixed-prompt mode |
--prefix-len | No | Fixed prefix token length to prepend to every prompt |
--output-len | No | Number of output tokens to generate per request |
--sort | No | Sort prompts by length before benchmarking |
--enable-prefix-caching / --no-enable-prefix-caching | No | Toggle APC (recommended: enable to test caching) |
--prefix-caching-hash-algo | No | Hash algorithm: sha256, sha256_cbor, xxhash, xxhash_cbor |
--tensor-parallel-size | No | Number of GPUs for tensor parallelism |
--disable-detokenize | No | Skip detokenization to reduce overhead |
python3 benchmarks/*.py reports file not found, locate your local vLLM repository first and run the command from that repo root.git clone https://github.com/vllm-project/vllm
cd vllmexport HF_TOKEN=<your_token> or pass --hf-token <your_token>.xxhash or cbor2 is not installed and you use those hash algorithms, install them first: pip install xxhash cbor2.© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/vllm-skills/skills/vllm-prefix-cache-bench of vllm-project/vllm-skills.
Open the folder on GitHubat commit c996234
Vllm Prefix Cache Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm Prefix Cache Bench this skillvllm-project/vllm-skills | 103 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Prefix Cache Replaybenchflow-ai/skillsbench | 1.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| CI Fails Buildkiteguqiong96/Lvllm | 465 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
benchflow-ai/skillsbench
Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
guqiong96/Lvllm
Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.
ModelCloud/GPTQModel
Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.
vllm-project/vllm-skills
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.
vllm-project/vllm-skills
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
vllm-project/vllm-skills
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Works with
Categories
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Vllm Prefix Cache Bench is an agent skill from vllm-project/vllm-skills. This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.
Vllm Prefix Cache Bench fits situations like: the user asks to benchmark prefix caching hit rate; caching efficiency; repeated-prompt performance in vLLM.
Run `npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-prefix-cache-bench in vllm-project/vllm-skills) into .claude/skills/vllm-prefix-cache-bench in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-prefix-cache-bench in vllm-project/vllm-skills) into .agents/skills/vllm-prefix-cache-bench in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-prefix-cache-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-prefix-cache-bench, .gemini/skills/vllm-prefix-cache-bench, .github/skills/vllm-prefix-cache-bench and .opencode/skills/vllm-prefix-cache-bench in your project.
Going by SKILL.md and its folder, Vllm Prefix Cache Bench needs the command-line tools its instructions call (python3, pip, wget and git) and credentials named HF_TOKEN. Our summary lists: Python 3.
SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vllm Prefix Cache Bench is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vllm Prefix Cache Bench: SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Prefix Cache Replay (benchflow-ai/skillsbench, 1.8k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 103 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.
Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.