Aider Delegate
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-skills vllm-bench-serve --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-serve .claude/skills/vllm-bench-serve && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-bench-serve" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serve into .claude/skills/vllm-bench-serve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-bench-serve", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serveType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-skills vllm-bench-serve --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-serve .agents/skills/vllm-bench-serve && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-bench-serve" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serve into .agents/skills/vllm-bench-serve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-bench-serve", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-skills vllm-bench-serve --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-serve .cursor/skills/vllm-bench-serve && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-bench-serve" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serve into .cursor/skills/vllm-bench-serve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-bench-serve", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-skills.git --path plugins/vllm-skills/skills/vllm-bench-serve--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-skills vllm-bench-serve --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-serve .gemini/skills/vllm-bench-serve && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-bench-serve" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serve into .gemini/skills/vllm-bench-serve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-bench-serve", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-skills vllm-bench-serveInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-serve .github/skills/vllm-bench-serve && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-bench-serve" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serve into .github/skills/vllm-bench-serve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-bench-serve", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-skills vllm-bench-serve --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-bench-serve .opencode/skills/vllm-bench-serve && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-bench-serve" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serve into .opencode/skills/vllm-bench-serve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-bench-serve", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-bench-serveBenchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
Vllm Bench Serve is an agent skill from vllm-project/vllm-skills. Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result saving. Use when benchmarking LLM serving performance, measuring TTFT/TPOT, or load testing inference APIs.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and OpenAI. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockerFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.vllm.aiarxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm Bench Serve loads about 1.6k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 397 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 397 words, ~1,643 tokens.
.claude/skills/vllm-bench-serve/SKILL.md (or your agent's skills folder).Benchmark vLLM or any OpenAI-compatible serving endpoint using the vllm bench serve CLI. Measures throughput, latency (TTFT, TPOT), and goodput against configurable request load.
Reference: vLLM Bench Serve Documentation
Basic benchmark against local vLLM server (default random dataset, 1000 prompts):
vllm bench serve \
--backend openai-chat \
--host 127.0.0.1 \
--port 8000 \
--model Qwen/Qwen2.5-1.5B-Instruct \
--endpoint /v1/chat/completionsSave results to JSON:
vllm bench serve \
--backend openai-chat \
--host 127.0.0.1 \
--port 8000 \
--model Qwen/Qwen2.5-1.5B-Instruct \
--endpoint /v1/chat/completions \
--save-result \
--result-dir ./bench-results \
--metadata "version=0.6.0" "tp=1"Note: When using
--backend openai-chat, you must specify--endpoint /v1/chat/completions(default is/v1/completions).
| Argument | Default | Description |
|---|---|---|
--backend | openai | Backend type: openai, openai-chat, openai-embeddings, vllm, vllm-pooling, vllm-rerank, etc. |
--host | 127.0.0.1 | Server host |
--port | 8000 | Server port |
--base-url | - | Alternative: full base URL instead of host:port |
--endpoint | /v1/completions | API endpoint; use /v1/chat/completions for openai-chat |
--model | (from /v1/models) | Model name |
--num-prompts | 1000 | Number of prompts to process |
--request-rate | inf | Requests per second; inf = burst all at once |
--max-concurrency | - | Max concurrent requests (caps parallelism) |
--num-warmups | 0 | Warmup requests before measuring |
--dataset-name | Use Case |
|---|---|
random | Synthetic random prompts (default) |
sharegpt | ShareGPT conversation format; requires --dataset-path |
sonnet | Sonnet-style prompts |
hf | HuggingFace dataset; requires --dataset-path (dataset ID) |
custom / custom_mm | Custom dataset; requires --dataset-path |
prefix_repetition | Prefix repetition benchmark |
random-mm | Random multimodal (images/videos) |
spec_bench | Spec bench dataset |
Dataset-specific options (examples):
# Random: control input/output length
--dataset-name random --random-input-len 1024 --random-output-len 128
# Sonnet defaults: input 550, output 150, prefix 200
--dataset-name sonnet --sonnet-input-len 550 --sonnet-output-len 150
# HuggingFace dataset
--dataset-name hf --dataset-path "lmarena-ai/VisionArena-Chat" --hf-split test
# General overrides (map to dataset-specific args)
--input-len 512 --output-len 256# Fixed request rate (Poisson process)
--request-rate 10
# More bursty arrivals (gamma distribution, burstiness < 1)
--request-rate 10 --burstiness 0.5
# Ramp-up from low to high RPS
--ramp-up-strategy linear --ramp-up-start-rps 1 --ramp-up-end-rps 50
# Limit concurrency (useful for rate-limited APIs)
--max-concurrency 32| Argument | Description |
|---|---|
--save-result | Save benchmark results to JSON |
--save-detailed | Include per-request TTFT, TPOT, errors in JSON |
--append-result | Append to existing result file |
--result-dir | Directory for result files |
--result-filename | Custom filename (default: {label}-{request_rate}qps-{model}-{timestamp}.json) |
--percentile-metrics | Metrics for percentiles: ttft, tpot, itl, e2el (default: ttft,tpot,itl) |
--metric-percentiles | Percentile values, e.g. 25,50,99 (default: 99) |
--goodput | SLO for goodput: ttft:500 tpot:50 (ms) |
--temperature 0.7 --top-p 0.95 --top-k 50
--frequency-penalty 0 --presence-penalty 0 --repetition-penalty 1.01. Throughput test with random dataset (burst):
vllm bench serve --backend openai-chat --host 127.0.0.1 --port 8000 \
--model Qwen/Qwen2.5-1.5B-Instruct \
--endpoint /v1/chat/completions \
--dataset-name random \
--num-prompts 500 --random-input-len 512 --random-output-len 1282. Latency test with fixed QPS:
vllm bench serve --backend openai-chat --host 127.0.0.1 --port 8000 \
--model Qwen/Qwen2.5-1.5B-Instruct \
--endpoint /v1/chat/completions \
--request-rate 5 --num-prompts 200 \
--save-result --percentile-metrics ttft,tpot --metric-percentiles 50,993. Benchmark against remote API (base-url):
vllm bench serve --backend openai-chat \
--base-url "https://api.example.com/v1" \
--model my-model \
--header "Authorization=Bearer $API_KEY"4. Run inside Docker (when vLLM client not on host):
docker exec <container-name> vllm bench serve \
--backend openai-chat --host 127.0.0.1 --port 8000 \
--model Qwen/Qwen2.5-1.5B-Instruct \
--endpoint /v1/chat/completions \
--dataset-name random --num-prompts 100--host/--port or --base-url are correct.--model explicitly or ensure /v1/models returns the model.--endpoint /v1/chat/completions when --backend openai-chat.--request-rate or --max-concurrency.--ready-check-timeout-sec 60 to wait for the endpoint before benchmarking.--insecure for self-signed certificates.--backend openai-embeddings, vllm-pooling, or vllm-rerank.--profile requires --profiler-config on the server for vLLM profiling.© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/vllm-skills/skills/vllm-bench-serve of vllm-project/vllm-skills.
Open the folder on GitHubat commit c996234
Vllm Bench Serve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm Bench Serve this skillvllm-project/vllm-skills | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Aider DelegateamElnagdy/delegate-skills | 2.3k | 2 repos | ~3k | Automated safety check: Pass | MIT | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Vllm Serversickn33/agentic-awesome-skills | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | |
| VllmPrism-Shadow/penguin-harness | 2.5k | — | ~1k | Automated safety check: Pass | Apache-2.0 |
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
sickn33/agentic-awesome-skills
Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.
Prism-Shadow/penguin-harness
Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.
Luciole-Studio/Misaka-Agent
vLLM: high-throughput LLM serving, OpenAI API, quantization.
vllm-project/vllm-skills
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
vllm-project/vllm-skills
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
vllm-project/vllm-skills
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Categories
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Vllm Bench Serve is an agent skill from vllm-project/vllm-skills. Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
Vllm Bench Serve fits situations like: benchmarking LLM serving performance; measuring TTFT/TPOT; load testing inference APIs.
Run `npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-bench-serve in vllm-project/vllm-skills) into .claude/skills/vllm-bench-serve in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-bench-serve in vllm-project/vllm-skills) into .agents/skills/vllm-bench-serve in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-bench-serve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-bench-serve, .gemini/skills/vllm-bench-serve, .github/skills/vllm-bench-serve and .opencode/skills/vllm-bench-serve in your project.
Going by SKILL.md and its folder, Vllm Bench Serve needs the command-line tools its instructions call (docker) and credentials named API_KEY. Our summary lists: Python 3; Docker; A credential in API_KEY.
SKILL.md names 2 domains. As links in the text: docs.vllm.ai and arxiv.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vllm Bench Serve is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vllm Bench Serve: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Server (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 102 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.
Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.