Agent skill

LLM Serving Framework Benchmark

by BBuf in BBuf/AI-Infra-Auto-Driven-SKILLS

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

No licenceAuto-check passedAI & LLM Engineering

Install LLM Serving Framework Benchmark

skills CLI
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-auto-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-auto-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-serving-auto-benchmark .claude/skills/llm-serving-auto-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-serving-auto-benchmark
GitHub stars
911
Token cost
~7.5k tokens
SKILL.md length
3,422 words
Files
62 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

  • Works in 8 steps: Preflight → Normalize The Workload → Pick A Search Tier → …
  • Picking the fastest serving framework for a model under a latency target
  • SKILL.md covers Overview, Source and experiment contracts, Validation Environment and Skill Scope, plus 7 more sections
  • Calls python, curl and ssh; needs HF_TOKEN and HUGGINGFACE_HUB_TOKEN

What it does

This skill compares LLM serving frameworks, including SGLang, vLLM, TensorRT-LLM and TokenSpeed, for one model under the same workload, GPU budget and latency SLA, in order to find the best deployment command. It is config-driven: fixed capacity choices go in each framework's `base_server_flags`, tunable options go in `search_space`, and every framework runs the same dataset scenarios. A bounded candidate list is generated from the search space with the baseline first, failed candidates stay in the results file, and the best candidate that meets the SLA is chosen after normalizing results.

For model-specific starting points it ships framework-neutral cookbook configs in `configs/cookbook-llm/`, which translate each model entry into native flags for each framework. A script, `validate_cookbook_configs.py`, loads them, checks flag names and renders candidate commands without launching any server. Native tooling is preferred, such as `sglang serve`, `vllm bench sweep serve`, `trtllm-serve` and `tokenspeed serve`. One hard scope rule: the TensorRT-LLM server is PyTorch-only here, so any candidate asking for `trt` or an engine backend is rejected and the reason recorded.

When your agent uses it

  • Picking the fastest serving framework for a model under a latency target
  • Sweeping server flags to find the best deployment command for a GPU budget
  • Validating cookbook configs before a benchmark run

Example prompts

  • “Compare SGLang and vLLM for gpt-oss-120b on my GPUs under a latency SLA.”
  • “Validate the cookbook YAML configs for the DeepSeek models before we run anything.”
  • “Find the best SLA-passing launch flags for this model across the frameworks and keep the failed candidates in the results.”

Requirements

  • GPUs with the serving frameworks you want to compare installed
  • Python, for the config validation script

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Preflight
  2. Normalize The Workload
  3. Pick A Search Tier
  4. Tune SGLang
  5. Tune vLLM
  6. Tune TensorRT-LLM
  7. Tune TokenSpeed
  8. Normalize Results

What it can do on your machine

Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • curl
    • ssh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • HUGGINGFACE_HUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Serving Framework Benchmark loads about 7.5k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 3,422 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~7.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 3,422 words (~7,532 tokens).

“Use this skill to compare LLM serving frameworks such as SGLang, vLLM, TensorRT-LLM, and TokenSpeed for the same model and workload.”

— opening of SKILL.md by BBuf
name
llm-serving-auto-benchmark

Read the full SKILL.md on GitHub

Files

SKILL.md and 61 other files (scripts, references) in skills/llm-serving-auto-benchmark of BBuf/AI-Infra-Auto-Driven-SKILLS.

  • SKILL.md
  • configs/cookbook-llm/README.md
  • configs/cookbook-llm/deepseek-math-v2.yaml
  • configs/cookbook-llm/deepseek-r1-0528.yaml
  • configs/cookbook-llm/deepseek-v3.1.yaml
  • configs/cookbook-llm/deepseek-v3.2.yaml
  • configs/cookbook-llm/deepseek-v3.yaml
  • configs/cookbook-llm/devstral-small-2-24b-instruct-2512.yaml
  • configs/cookbook-llm/ernie-4.5-21b-a3b-pt.yaml
  • configs/cookbook-llm/gigachat3.5.yaml
  • configs/cookbook-llm/glm-4.5.yaml
  • configs/cookbook-llm/glm-4.6.yaml
  • configs/cookbook-llm/glm-4.7-flash.yaml
  • configs/cookbook-llm/glm-4.7.yaml
  • configs/cookbook-llm/glm-5-fp8.yaml
  • configs/cookbook-llm/glm-5.3-bf16.yaml
  • configs/cookbook-llm/glm-5.3-flash.yaml
  • configs/cookbook-llm/glyph.yaml
  • configs/cookbook-llm/gpt-oss-120b.yaml
  • … and 43 more

Open the folder on GitHubat commit 6dc9c66

Compare with similar skills

LLM Serving Framework Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Serving Framework Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Serving Framework Benchmark this skillBBuf/AI-Infra-Auto-Driven-SKILLS911—~7.5kAutomated safety check: PassNone
Magpie Kernel Evaluatoramd/skills398—~2.3kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
One EvalOpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.0
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    398 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • One Eval

    OpenDCAI/One-Eval

    驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

    165 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 5 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from BBuf/AI-Infra-Auto-Driven-SKILLS

All 10 skills in this repo
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    911 GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    911 GitHub stars~2.5k tokensUpdated 3 days ago
    Auto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    Auto-check passed
  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    911 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Model Compute Simulator

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

    911 GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • SGLang Maintainer-Style Review

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.

    911 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed

Questions about LLM Serving Framework Benchmark

What does LLM Serving Framework Benchmark do?

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. This skill compares LLM serving frameworks, including SGLang, vLLM, TensorRT-LLM and TokenSpeed, for one model under the same workload, GPU budget and latency SLA, in order to find the best deployment command. It is config-driven: fixed capacity choices go in each framework's `base_server_flags`, tunable options go in `search_space`, and every framework runs the same dataset scenarios.

When should I use LLM Serving Framework Benchmark?

LLM Serving Framework Benchmark fits situations like: picking the fastest serving framework for a model under a latency target; sweeping server flags to find the best deployment command for a GPU budget; validating cookbook configs before a benchmark run.

How do I install LLM Serving Framework Benchmark in Claude Code?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-auto-benchmark -a claude-code`. Or copy the skill folder (skills/llm-serving-auto-benchmark in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/llm-serving-auto-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install LLM Serving Framework Benchmark in Codex?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-auto-benchmark -a codex`. Or copy the skill folder (skills/llm-serving-auto-benchmark in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/llm-serving-auto-benchmark in your project. Codex loads it when a task matches its description.

Can I use LLM Serving Framework Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-auto-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-serving-auto-benchmark, .gemini/skills/llm-serving-auto-benchmark, .github/skills/llm-serving-auto-benchmark and .opencode/skills/llm-serving-auto-benchmark in your project.

What does LLM Serving Framework Benchmark need to run?

Going by SKILL.md and its folder, LLM Serving Framework Benchmark needs the command-line tools its instructions call (python, curl and ssh) and credentials named HF_TOKEN and HUGGINGFACE_HUB_TOKEN. Our summary lists: GPUs with the serving frameworks you want to compare installed; Python, for the config validation script.

Does LLM Serving Framework Benchmark access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is LLM Serving Framework Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does LLM Serving Framework Benchmark use?

No licence was found for LLM Serving Framework Benchmark or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does LLM Serving Framework Benchmark use?

About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to LLM Serving Framework Benchmark?

Skills that share tags, products or a category with LLM Serving Framework Benchmark: Magpie Kernel Evaluator (amd/skills, 398 stars), Graphsignal (graphsignal/graphsignal, 257 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and One Eval (OpenDCAI/One-Eval, 165 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Serving Framework Benchmark?

BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 911 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.

Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.