Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve.

No licenceAuto-check passedAI & LLM Engineering

Install Vllm Bench Serve

skills CLI
$ npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-bench-serve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ascend-ai-coding/awesome-ascend-skills vllm-bench-serve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ascend-ai-coding/awesome-ascend-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/inference/vllm-bench-serve .claude/skills/vllm-bench-serve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-bench-serve
GitHub stars
174
Token cost
~5.6k tokens
SKILL.md length
2,112 words
Files
18 (incl. scripts, references, assets)
Skills in repo
70
Repo updated
First seen
Licence
None found

At a glance

Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve.

  • Works in 12 steps: Scope & Boundaries → Trigger Examples → Workflow Overview → …
  • The user wants to benchmark
  • SKILL.md covers 1. Scope & Boundaries, 2. Trigger Examples, 3. Workflow Overview and 4. Phase 0 — Execution…, plus 7 more sections
  • Runs Python and Shell scripts from its folder; calls bash, python3 and curl

What it does

Vllm Bench Serve is an agent skill from ascend-ai-coding/awesome-ascend-skills. Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. Supports single benchmarks, multi-case batch execution with result aggregation, and auto-optimization to find optimal concurrency/throughput under latency SLO constraints (TTFT, TPOT, P99, success rate). Use this skill whenever the user wants to benchmark, stress test, or measure performance of a running LLM/multimodal/embedding inference service — even if they don't say "vllm bench serve" explicitly. Do NOT use for…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including scripts, reference files and assets (for example `assets/report_template.md`, `references/backend-mapping.md` and `references/dataset-guide.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving, Site reliability engineering and Load testing. It works with vLLM. The repository describes itself as: A comprehensive knowledge base for Huawei Ascend NPU development, structured as distributed Agent Skills. https://ascend-ai-coding.github.io/awesome-ascend-skills/.

When your agent uses it

  • The user wants to benchmark
  • Measure performance of a running LLM/multimodal/embedding inference service — even if they dont say vllm bench serve explicitly
  • Offline inference throughput
  • Service deployment/startup

Example prompts

  • “t say”
  • “/vllm-bench-serve”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Scope & Boundaries
  2. Trigger Examples
  3. Workflow Overview
  4. Phase 0 — Execution Environment
  5. Phase 1 — Service & Model
  6. Phase 2 — Mode Selection
  7. Phase 3 — Backend Selection
  8. Phase 4 — Dataset Selection
  9. Phase 5 — Benchmark Parameters
  10. Phase 6 — Parameter Validation
  11. Phase 7 — Command Generation & Execution
  12. Phase 8 — Auto-Optimize

What it can do on your machine

Read from SKILL.md and the folder at commit 62a4ecb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • python3
    • curl
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Bench Serve loads about 5.6k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 174 tokens; SKILL.md has 2,112 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~174
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 2,112 words (~5,586 tokens).

name
vllm-bench-serve
keywords
vllm bench serve, benchmark, online benchmark, 在线性能测试, 推理服务性能, throughput, latency, TTFT, TPOT, P99, concurrency, 并发测试, auto-optimize, 寻优, SLO, request rate…

Read the full SKILL.md on GitHub

Files

SKILL.md and 17 other files (scripts, references, assets) in skills/inference/vllm-bench-serve of ascend-ai-coding/awesome-ascend-skills.

  • SKILL.md
  • assets/report_template.md
  • references/backend-mapping.md
  • references/dataset-guide.md
  • references/environment-checks.md
  • references/optimization-strategy.md
  • references/param-reference.md
  • references/scenario-cookbook.md
  • references/troubleshooting.md
  • scripts/aggregate_results.py
  • scripts/auto_optimize.py
  • scripts/check_bench_env.sh
  • scripts/common.py
  • scripts/generate_bench_cmd.py
  • scripts/probe_service.sh
  • scripts/run_batch.sh
  • scripts/run_bench.sh
  • scripts/validate_params.py

Open the folder on GitHubat commit 62a4ecb

Compare with similar skills

Vllm Bench Serve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Bench Serve compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Bench Serve this skillascend-ai-coding/awesome-ascend-skills174—~5.6kAutomated safety check: PassNone
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Add Diffusion Modelvllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Metax Model Upgrade

    MetaX-MACA/vLLM-metax

    Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

    179 GitHub stars~2.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from ascend-ai-coding/awesome-ascend-skills

All 70 skills in this repo
  • Ascend Dmi

    ascend-ai-coding/awesome-ascend-skills

    当用户需要对华为昇腾 NPU 进行硬件层面的管理、测试或诊断时使用此 skill。典型场景: - 查看 NPU 卡的状态、温度、利用率 - 测试内存带宽(h2d/d2h/d2d/p2p) - 跑算力/功耗基准测试(TFLOPS、TOPS) - 诊断 NPU 硬件故障或做健康检查 - 对 NPU 卡做压力测试(aicore、内存) - 复位/恢复卡住或异常的 NPU 卡 典型用户问题(即使不提…

    174 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Ascendc

    ascend-ai-coding/awesome-ascend-skills

    End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.

    174 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Atc Model Converter

    ascend-ai-coding/awesome-ascend-skills

    Complete toolkit for Huawei Ascend NPU model conversion and end-to-end inference adaptation.

    174 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • External Cannbot Ops Pypto Op Develop

    ascend-ai-coding/awesome-ascend-skills

    当需要编写 PyPTO 算子实现时使用此 skill。基于需求规格、设计方案和参考实现,生成完整可运行的 PyPTO 算子实现与配套测试、文档。Triggers: 实现算子、写 kernel、编写实现、写 impl、算子编码、开始编码、code the op、写 test、生成测试、写实现代码、op develop、kernel 实现。

    174 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Change Analyzer

    ascend-ai-coding/awesome-ascend-skills

    Analyze official Megatron-LM commits, PRs, and branch change sets to identify feature evolution, candidate breaking changes, and migration-relevant events.

    174 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Commit Tracker

    ascend-ai-coding/awesome-ascend-skills

    Track and normalize change requests against the official Megatron-LM repository by branch, PR, commit, commit range, or time window.

    174 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Works with

Questions about Vllm Bench Serve

What does Vllm Bench Serve do?

Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. Vllm Bench Serve is an agent skill from ascend-ai-coding/awesome-ascend-skills. Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve.

When should I use Vllm Bench Serve?

Vllm Bench Serve fits situations like: the user wants to benchmark; measure performance of a running LLM/multimodal/embedding inference service — even if they dont say vllm bench serve explicitly; offline inference throughput; service deployment/startup.

How do I install Vllm Bench Serve in Claude Code?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-bench-serve -a claude-code`. Or copy the skill folder (skills/inference/vllm-bench-serve in ascend-ai-coding/awesome-ascend-skills) into .claude/skills/vllm-bench-serve in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Bench Serve in Codex?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-bench-serve -a codex`. Or copy the skill folder (skills/inference/vllm-bench-serve in ascend-ai-coding/awesome-ascend-skills) into .agents/skills/vllm-bench-serve in your project. Codex loads it when a task matches its description.

Can I use Vllm Bench Serve in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-bench-serve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-bench-serve, .gemini/skills/vllm-bench-serve, .github/skills/vllm-bench-serve and .opencode/skills/vllm-bench-serve in your project.

What does Vllm Bench Serve need to run?

Going by SKILL.md and its folder, Vllm Bench Serve needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (bash, python3, curl and docker). Our summary lists: Python 3; A Bash shell.

Does Vllm Bench Serve access the network?

SKILL.md contains no URLs. Its commands use curl and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vllm Bench Serve safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Vllm Bench Serve use?

No licence was found for Vllm Bench Serve or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Vllm Bench Serve use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Vllm Bench Serve?

Skills that share tags, products or a category with Vllm Bench Serve: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Bench Serve?

ascend-ai-coding (a GitHub organization) maintains it in ascend-ai-coding/awesome-ascend-skills, which has 174 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 9, 2026.

Source: ascend-ai-coding/awesome-ascend-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.