Agent skill

Vllm Bench Random Synthetic

by vllm-project in vllm-project/vllm-skills

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Bench Random Synthetic

skills CLI
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-random-synthetic -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-skills vllm-bench-random-synthetic --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-random-synthetic .claude/skills/vllm-bench-random-synthetic && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-bench-random-synthetic
GitHub stars
102
Token cost
~1.5k tokens
SKILL.md length
478 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

  • Works in 6 steps: Check if vLLM is installed: Run vllm… → Check if server is already running: Run… → Start vLLM server if needed: Run vllm… → …
  • The user wants to quickly test vLLM serving performance without downloading external datasets
  • SKILL.md covers When to use, Prerequisites, Quick Start and Parameters, plus 6 more sections
  • Calls curl and pip; needs HF_TOKEN

What it does

Vllm Bench Random Synthetic is an agent skill from vllm-project/vllm-skills. Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading external datasets.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and OpenAI. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.

When your agent uses it

  • The user wants to quickly test vLLM serving performance without downloading external datasets
  • Tasks that involve LLM inference and serving

Example prompts

  • “/vllm-bench-random-synthetic”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Check if vLLM is installed: Run vllm --version to verify
  2. Check if server is already running: Run curl http://localhost:8000/health to check
  3. Start vLLM server if needed: Run vllm serve (wait for "Application startup complete")
  4. Run benchmark: Execute vllm bench serve with appropriate parameters
  5. Review results: Check throughput and latency metrics
  6. Clean up: If the agent skill started the vLLM server (not a pre-existing one), stop it after benchmark completion using kill

What it can do on your machine

Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Bench Random Synthetic loads about 1.5k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 478 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 478 words, ~1,526 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-bench-random-synthetic/SKILL.md (or your agent's skills folder).
name
vllm-bench-random-synthetic
description
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading external datasets.

vLLM Benchmark with Random Synthetic Data

Run a quick performance benchmark on a vLLM server using synthetic random data. This skill measures core serving metrics including request throughput, token throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and inter-token latency.

When to use

  • User wants to quickly benchmark vLLM serving performance
  • User wants to measure throughput and latency metrics without downloading datasets
  • User wants to test a vLLM deployment with synthetic workload
  • User wants baseline performance numbers for a specific model

Prerequisites

  • vLLM must be installed (pip install vllm)
  • A vLLM server must be running (or can be started as part of the benchmark)
  • For GPU models, NVIDIA GPU with appropriate drivers must be available

Quick Start

The simplest way to run the benchmark:

bash
# Start vLLM server (in background or separate terminal)
vllm serve Qwen/Qwen2.5-1.5B-Instruct

# Run benchmark with random synthetic data
vllm bench serve \
  --backend openai-chat \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions \
  --dataset-name random \
  --num-prompts 10

Note:

  • Use --backend openai-chat with endpoint /v1/chat/completions for online benchmarks.

Parameters

ParameterDescriptionDefault
--backendBackend type: vllm, openai, openai-chatvllm
--modelModel name (must match the server)Required
--endpointAPI endpoint path/v1/completions or /v1/chat/completions
--dataset-nameDataset to userandom (synthetic)
--num-promptsNumber of requests to send10
--portServer port8000
--max-concurrencyMaximum concurrent requestsAuto
--save-resultSave results to fileOff
--result-dirDirectory to save results./

Expected Output

When successful, you will see output like:

============ Serving Benchmark Result ============
Successful requests:                     10
Benchmark duration (s):                  5.78
Total input tokens:                      1369
Total generated tokens:                  2212
Request throughput (req/s):              1.73
Output token throughput (tok/s):         382.89
Total token throughput (tok/s):          619.85
---------------Time to First Token----------------
Mean TTFT (ms):                          71.54
Median TTFT (ms):                        73.88
P99 TTFT (ms):                           79.49
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          7.91
Median TPOT (ms):                        7.96
P99 TPOT (ms):                           8.03
---------------Inter-token Latency----------------
Mean ITL (ms):                           7.74
Median ITL (ms):                         7.70
P99 ITL (ms):                            8.39
==================================================

Advanced Usage

With more prompts for better statistics
bash
vllm bench serve \
  --backend openai-chat \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions \
  --dataset-name random \
  --num-prompts 100
Save results to file
bash
vllm bench serve \
  --backend openai-chat \
  --model Qwen/Qwen2.5-1.5B-Instruct \
  --endpoint /v1/chat/completions \
  --dataset-name random \
  --num-prompts 50 \
  --save-result \
  --result-dir ./benchmark-results/
Custom port and concurrency
bash
vllm bench serve \
  --backend openai-chat \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --endpoint /v1/chat/completions \
  --dataset-name random \
  --num-prompts 100 \
  --port 8001 \
  --max-concurrency 4

Model Recommendations

For quick testing (small models, fast):

  • Qwen/Qwen2.5-1.5B-Instruct (recommended for quick tests)
  • facebook/opt-125m
  • facebook/opt-350m

For realistic benchmarks (medium models):

  • Qwen/Qwen2.5-7B-Instruct
  • meta-llama/Llama-3.1-8B-Instruct
  • mistralai/Mistral-7B-Instruct-v0.3

Workflow

  1. Check if vLLM is installed: Run vllm --version to verify
  2. Check if server is already running: Run curl http://localhost:8000/health to check
  3. Start vLLM server if needed: Run vllm serve <model-name> (wait for "Application startup complete")
  4. Run benchmark: Execute vllm bench serve with appropriate parameters
  5. Review results: Check throughput and latency metrics
  6. Clean up: If the agent skill started the vLLM server (not a pre-existing one), stop it after benchmark completion using kill <PID>
Show full SKILL.md (154 more words)Show less

Troubleshooting

Server not responding:

  • Check if server is running: curl http://localhost:8000/health
  • Verify port matches: Use --port flag if server is on different port

Model not found:

  • Ensure model name matches exactly between server and benchmark
  • Check HuggingFace access: export HF_TOKEN=<your_token> if needed

Out of memory:

  • Use a smaller model (e.g., Qwen2.5-1.5B-Instruct)
  • Reduce --num-prompts or --max-concurrency

Connection refused:

  • Server may still be starting (wait for "Application startup complete")
  • Check firewall or network settings

Notes

  • The random dataset generates synthetic prompts automatically
  • Benchmark duration scales with --num-prompts
  • For production benchmarking, use at least 100 prompts for stable statistics
  • Results may vary based on hardware, model size, and system load
  • First run may be slower due to model loading and compilation
  • Important: If the agent skill starts a vLLM server for benchmarking, it must stop the server after the benchmark completes to free up resources. Do not stop pre-existing servers that were already running before the benchmark.

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/vllm-skills/skills/vllm-bench-random-synthetic of vllm-project/vllm-skills.

Open the folder on GitHubat commit c996234

Compare with similar skills

Vllm Bench Random Synthetic next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Bench Random Synthetic compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Bench Random Synthetic this skillvllm-project/vllm-skills102—~1.5kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT
Vllm Serversickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT
VllmPrism-Shadow/penguin-harness2.5k—~1kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Server

    sickn33/agentic-awesome-skills

    Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm

    Prism-Shadow/penguin-harness

    Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.

    2.5k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Serving LLMs Vllm

    Luciole-Studio/Misaka-Agent

    vLLM: high-throughput LLM serving, OpenAI API, quantization.

    171 GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-skills

  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    102 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    102 GitHub stars~2k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    102 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Prefix Cache Bench

    vllm-project/vllm-skills

    This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

    102 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    Auto-check: notes

Works with

Questions about Vllm Bench Random Synthetic

What does Vllm Bench Random Synthetic do?

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Vllm Bench Random Synthetic is an agent skill from vllm-project/vllm-skills. Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

When should I use Vllm Bench Random Synthetic?

Vllm Bench Random Synthetic fits situations like: the user wants to quickly test vLLM serving performance without downloading external datasets; tasks that involve LLM inference and serving.

How do I install Vllm Bench Random Synthetic in Claude Code?

Run `npx skills add vllm-project/vllm-skills --skill vllm-bench-random-synthetic -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-bench-random-synthetic in vllm-project/vllm-skills) into .claude/skills/vllm-bench-random-synthetic in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Bench Random Synthetic in Codex?

Run `npx skills add vllm-project/vllm-skills --skill vllm-bench-random-synthetic -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-bench-random-synthetic in vllm-project/vllm-skills) into .agents/skills/vllm-bench-random-synthetic in your project. Codex loads it when a task matches its description.

Can I use Vllm Bench Random Synthetic in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-bench-random-synthetic -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-bench-random-synthetic, .gemini/skills/vllm-bench-random-synthetic, .github/skills/vllm-bench-random-synthetic and .opencode/skills/vllm-bench-random-synthetic in your project.

What does Vllm Bench Random Synthetic need to run?

Going by SKILL.md and its folder, Vllm Bench Random Synthetic needs the command-line tools its instructions call (curl and pip) and credentials named HF_TOKEN. Our summary lists: Python 3.

Does Vllm Bench Random Synthetic access the network?

SKILL.md contains no URLs. Its commands use curl and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vllm Bench Random Synthetic safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Bench Random Synthetic use?

Vllm Bench Random Synthetic is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Bench Random Synthetic use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vllm Bench Random Synthetic?

Skills that share tags, products or a category with Vllm Bench Random Synthetic: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Server (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Bench Random Synthetic?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 102 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.

Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.