Agent skill

Vllm Bench Serve

by vllm-project in vllm-project/vllm-skills

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Bench Serve

skills CLI
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-skills vllm-bench-serve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-serve .claude/skills/vllm-bench-serve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-bench-serve
GitHub stars
102
Token cost
~1.6k tokens
SKILL.md length
397 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

  • Benchmarking LLM serving performance
  • SKILL.md covers Prerequisites, Quick Start, Core Arguments and Datasets, plus 6 more sections
  • Calls docker; needs API_KEY
  • Measuring TTFT/TPOT

What it does

Vllm Bench Serve is an agent skill from vllm-project/vllm-skills. Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result saving. Use when benchmarking LLM serving performance, measuring TTFT/TPOT, or load testing inference APIs.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and OpenAI. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.

When your agent uses it

  • Benchmarking LLM serving performance
  • Measuring TTFT/TPOT
  • Load testing inference APIs

Example prompts

  • “/vllm-bench-serve”

Requirements

  • Python 3
  • Docker
  • A credential in API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.vllm.ai
    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Bench Serve loads about 1.6k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 397 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 397 words, ~1,643 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-bench-serve/SKILL.md (or your agent's skills folder).
name
vllm-bench-serve
description
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result saving. Use when benchmarking LLM serving performance, measuring TTFT/TPOT, or load testing inference APIs.

vLLM Bench Serve

Benchmark vLLM or any OpenAI-compatible serving endpoint using the vllm bench serve CLI. Measures throughput, latency (TTFT, TPOT), and goodput against configurable request load.

Reference: vLLM Bench Serve Documentation

Prerequisites

  • vLLM installed (or any OpenAI-compatible server running)
  • A vLLM server or API endpoint already serving a model
  • Python environment with vLLM for the benchmark client

Quick Start

Basic benchmark against local vLLM server (default random dataset, 1000 prompts):

bash
vllm bench serve \
  --backend openai-chat \
  --host 127.0.0.1 \
  --port 8000 \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions

Save results to JSON:

bash
vllm bench serve \
  --backend openai-chat \
  --host 127.0.0.1 \
  --port 8000 \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions \
  --save-result \
  --result-dir ./bench-results \
  --metadata "version=0.6.0" "tp=1"

Note: When using --backend openai-chat, you must specify --endpoint /v1/chat/completions (default is /v1/completions).

Core Arguments

ArgumentDefaultDescription
--backendopenaiBackend type: openai, openai-chat, openai-embeddings, vllm, vllm-pooling, vllm-rerank, etc.
--host127.0.0.1Server host
--port8000Server port
--base-url-Alternative: full base URL instead of host:port
--endpoint/v1/completionsAPI endpoint; use /v1/chat/completions for openai-chat
--model(from /v1/models)Model name
--num-prompts1000Number of prompts to process
--request-rateinfRequests per second; inf = burst all at once
--max-concurrency-Max concurrent requests (caps parallelism)
--num-warmups0Warmup requests before measuring

Datasets

--dataset-nameUse Case
randomSynthetic random prompts (default)
sharegptShareGPT conversation format; requires --dataset-path
sonnetSonnet-style prompts
hfHuggingFace dataset; requires --dataset-path (dataset ID)
custom / custom_mmCustom dataset; requires --dataset-path
prefix_repetitionPrefix repetition benchmark
random-mmRandom multimodal (images/videos)
spec_benchSpec bench dataset

Dataset-specific options (examples):

bash
# Random: control input/output length
--dataset-name random --random-input-len 1024 --random-output-len 128

# Sonnet defaults: input 550, output 150, prefix 200
--dataset-name sonnet --sonnet-input-len 550 --sonnet-output-len 150

# HuggingFace dataset
--dataset-name hf --dataset-path "lmarena-ai/VisionArena-Chat" --hf-split test

# General overrides (map to dataset-specific args)
--input-len 512 --output-len 256

Load Control

bash
# Fixed request rate (Poisson process)
--request-rate 10

# More bursty arrivals (gamma distribution, burstiness < 1)
--request-rate 10 --burstiness 0.5

# Ramp-up from low to high RPS
--ramp-up-strategy linear --ramp-up-start-rps 1 --ramp-up-end-rps 50

# Limit concurrency (useful for rate-limited APIs)
--max-concurrency 32
Show full SKILL.md (187 more words)Show less

Results and Metrics

ArgumentDescription
--save-resultSave benchmark results to JSON
--save-detailedInclude per-request TTFT, TPOT, errors in JSON
--append-resultAppend to existing result file
--result-dirDirectory for result files
--result-filenameCustom filename (default: {label}-{request_rate}qps-{model}-{timestamp}.json)
--percentile-metricsMetrics for percentiles: ttft, tpot, itl, e2el (default: ttft,tpot,itl)
--metric-percentilesPercentile values, e.g. 25,50,99 (default: 99)
--goodputSLO for goodput: ttft:500 tpot:50 (ms)

Sampling Parameters (OpenAI-compatible backends)

bash
--temperature 0.7 --top-p 0.95 --top-k 50
--frequency-penalty 0 --presence-penalty 0 --repetition-penalty 1.0

Common Workflows

1. Throughput test with random dataset (burst):

bash
vllm bench serve --backend openai-chat --host 127.0.0.1 --port 8000 \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions \
  --dataset-name random \
  --num-prompts 500 --random-input-len 512 --random-output-len 128

2. Latency test with fixed QPS:

bash
vllm bench serve --backend openai-chat --host 127.0.0.1 --port 8000 \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions \
  --request-rate 5 --num-prompts 200 \
  --save-result --percentile-metrics ttft,tpot --metric-percentiles 50,99

3. Benchmark against remote API (base-url):

bash
vllm bench serve --backend openai-chat \
  --base-url "https://api.example.com/v1" \
  --model my-model \
  --header "Authorization=Bearer $API_KEY"

4. Run inside Docker (when vLLM client not on host):

bash
docker exec <container-name> vllm bench serve \
  --backend openai-chat --host 127.0.0.1 --port 8000 \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions \
  --dataset-name random --num-prompts 100

Troubleshooting

  • Connection refused: Ensure the server is running and --host/--port or --base-url are correct.
  • Model not found: Pass --model explicitly or ensure /v1/models returns the model.
  • URL must end with chat/completions: Use --endpoint /v1/chat/completions when --backend openai-chat.
  • Rate limit / 429: Reduce --request-rate or --max-concurrency.
  • Ready check: Use --ready-check-timeout-sec 60 to wait for the endpoint before benchmarking.
  • SSL: Use --insecure for self-signed certificates.

Notes

  • For embeddings/rerank benchmarks, use --backend openai-embeddings, vllm-pooling, or vllm-rerank.
  • --profile requires --profiler-config on the server for vLLM profiling.
  • Goodput SLOs are useful for SLA-style analysis; see DistServe paper for details.

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/vllm-skills/skills/vllm-bench-serve of vllm-project/vllm-skills.

Open the folder on GitHubat commit c996234

Compare with similar skills

Vllm Bench Serve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Bench Serve compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Bench Serve this skillvllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT
Vllm Serversickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT
VllmPrism-Shadow/penguin-harness2.5k—~1kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Server

    sickn33/agentic-awesome-skills

    Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm

    Prism-Shadow/penguin-harness

    Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.

    2.5k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Serving LLMs Vllm

    Luciole-Studio/Misaka-Agent

    vLLM: high-throughput LLM serving, OpenAI API, quantization.

    171 GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-skills

  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    102 GitHub stars~1.5k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    102 GitHub stars~2k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    102 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Prefix Cache Bench

    vllm-project/vllm-skills

    This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

    102 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    Auto-check: notes

Works with

Questions about Vllm Bench Serve

What does Vllm Bench Serve do?

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Vllm Bench Serve is an agent skill from vllm-project/vllm-skills. Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

When should I use Vllm Bench Serve?

Vllm Bench Serve fits situations like: benchmarking LLM serving performance; measuring TTFT/TPOT; load testing inference APIs.

How do I install Vllm Bench Serve in Claude Code?

Run `npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-bench-serve in vllm-project/vllm-skills) into .claude/skills/vllm-bench-serve in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Bench Serve in Codex?

Run `npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-bench-serve in vllm-project/vllm-skills) into .agents/skills/vllm-bench-serve in your project. Codex loads it when a task matches its description.

Can I use Vllm Bench Serve in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-bench-serve, .gemini/skills/vllm-bench-serve, .github/skills/vllm-bench-serve and .opencode/skills/vllm-bench-serve in your project.

What does Vllm Bench Serve need to run?

Going by SKILL.md and its folder, Vllm Bench Serve needs the command-line tools its instructions call (docker) and credentials named API_KEY. Our summary lists: Python 3; Docker; A credential in API_KEY.

Does Vllm Bench Serve access the network?

SKILL.md names 2 domains. As links in the text: docs.vllm.ai and arxiv.org. This is read from the text; nothing was executed.

Is Vllm Bench Serve safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Bench Serve use?

Vllm Bench Serve is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Bench Serve use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vllm Bench Serve?

Skills that share tags, products or a category with Vllm Bench Serve: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Server (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Bench Serve?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 102 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.

Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.