Agent skill

LLM Serving Capacity Planner

by BBuf in BBuf/AI-Infra-Auto-Driven-SKILLS

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

No licenceAuto-check passedAI & LLM Engineering

Install LLM Serving Capacity Planner

skills CLI
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-capacity-planner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-serving-capacity-planner .claude/skills/llm-serving-capacity-planner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-serving-capacity-planner
GitHub stars
925
Token cost
~2.5k tokens
SKILL.md length
1,238 words
Files
4 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

  • Works in 4 steps: Collect the serving log → Optionally capture nvidia-smi data → Run the analyzer → …
  • Explaining where GPU memory went after an SGLang or vLLM server starts
  • SKILL.md covers Overview, Inputs, Workflow and When To Use It, plus 6 more sections
  • Runs Python scripts from its folder; calls python3, docker and mamba

What it does

This skill analyzes the startup log of an SGLang or vLLM server to explain how GPU memory is divided. A script, scripts/capacity_analyzer.py, pulls out the weight load, KV cache pool, CUDA graph, framework overhead and token-capacity lines, then estimates concurrent requests for common request lengths. If you give none, it uses 4096, 6144 and 8192 tokens.

The log file is the only required input. The GPU type can be detected from the log, nvidia-smi output adds per-rank memory for cross-checking, and the model's config.json lets the skill calculate KV cache bytes in theory. Reference files hold GPU specs and log patterns. For DeepSeek-V4.1 it warns that per-token byte figures are payload sizes, not allocation after page padding, and points to a separate KV layout reference to read first.

When your agent uses it

  • Explaining where GPU memory went after an SGLang or vLLM server starts
  • Triaging an out-of-memory failure from a serving log
  • Comparing mem-fraction-static settings and their effect on the KV cache budget
  • Estimating maximum concurrent requests at a given token length

Example prompts

  • “Here is my SGLang startup log; explain the memory breakdown on each GPU.”
  • “Compare these two vLLM logs and tell me how many 8192-token requests fit in each.”
  • “My server ran out of memory on startup; use the log and the nvidia-smi output to find out why.”

Requirements

  • A SGLang or vLLM serving startup log
  • Optional nvidia-smi output and the model's config.json

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Collect the serving log
  2. Optionally capture nvidia-smi data
  3. Run the analyzer
  4. Review and interpret results

What it can do on your machine

Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • docker
    • mamba

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Serving Capacity Planner loads about 2.5k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 1,238 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,238 words (~2,518 tokens).

“Use this when a serving log has enough memory lines to explain where GPU HBM went. The analyzer reads SGLang/vLLM startup logs, extracts weight load, KV pool, CUDA graph, framework overhead, and token-capacity lines, then estimates concurrent requests for common…”

— opening of SKILL.md by BBuf
name
llm-serving-capacity-planner

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (scripts, references) in skills/llm-serving-capacity-planner of BBuf/AI-Infra-Auto-Driven-SKILLS.

  • SKILL.md
  • references/gpu-specs.json
  • references/log-patterns.md
  • scripts/capacity_analyzer.py

Open the folder on GitHubat commit 6dc9c66

Compare with similar skills

LLM Serving Capacity Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Serving Capacity Planner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Serving Capacity Planner this skillBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.5kAutomated safety check: PassNone
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills406—~2.3kAutomated safety check: PassMIT
Jetson Inference Mem TuneNVIDIA/skills3.5k1 repos~2.9kAutomated safety check: PassApache-2.0
Jetson LLM ServeNVIDIA/skills3.5k1 repos~3kAutomated safety check: NotesApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    406 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

    3.5k GitHub starsUsed in 1 repo~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Jetson LLM Serve

    NVIDIA/skills

    Official

    Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

    3.5k GitHub starsUsed in 1 repo~3k tokens
    AI & LLM EngineeringAuto-check: notes
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from BBuf/AI-Infra-Auto-Driven-SKILLS

All 10 skills in this repo
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    925 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check passed
  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    925 GitHub stars~1.2k tokensUpdated 4 days ago
    Auto-check passed
  • Model Compute Simulator

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

    925 GitHub stars~4.5k tokensUpdated 4 days ago
    Auto-check passed
  • SGLang Maintainer-Style Review

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.

    925 GitHub stars~4.6k tokensUpdated 4 days ago
    Auto-check passed
  • Torch Profiler Layer Track

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

    925 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about LLM Serving Capacity Planner

What does LLM Serving Capacity Planner do?

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. This skill analyzes the startup log of an SGLang or vLLM server to explain how GPU memory is divided.py, pulls out the weight load, KV cache pool, CUDA graph, framework overhead and token-capacity lines, then estimates concurrent requests for common request lengths.

When should I use LLM Serving Capacity Planner?

LLM Serving Capacity Planner fits situations like: explaining where GPU memory went after an SGLang or vLLM server starts; triaging an out-of-memory failure from a serving log; comparing mem-fraction-static settings and their effect on the KV cache budget; estimating maximum concurrent requests at a given token length.

How do I install LLM Serving Capacity Planner in Claude Code?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a claude-code`. Or copy the skill folder (skills/llm-serving-capacity-planner in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/llm-serving-capacity-planner in your project. Claude Code loads it when a task matches its description.

How do I install LLM Serving Capacity Planner in Codex?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a codex`. Or copy the skill folder (skills/llm-serving-capacity-planner in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/llm-serving-capacity-planner in your project. Codex loads it when a task matches its description.

Can I use LLM Serving Capacity Planner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-serving-capacity-planner, .gemini/skills/llm-serving-capacity-planner, .github/skills/llm-serving-capacity-planner and .opencode/skills/llm-serving-capacity-planner in your project.

What does LLM Serving Capacity Planner need to run?

Going by SKILL.md and its folder, LLM Serving Capacity Planner needs Python for the scripts in its folder and the command-line tools its instructions call (python3, docker and mamba). Our summary lists: A SGLang or vLLM serving startup log; Optional nvidia-smi output and the model's config.json.

Does LLM Serving Capacity Planner access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is LLM Serving Capacity Planner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does LLM Serving Capacity Planner use?

No licence was found for LLM Serving Capacity Planner or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does LLM Serving Capacity Planner use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to LLM Serving Capacity Planner?

Skills that share tags, products or a category with LLM Serving Capacity Planner: Graphsignal (graphsignal/graphsignal, 257 stars), Magpie Kernel Evaluator (amd/skills, 406 stars), Jetson Inference Mem Tune (NVIDIA/skills, 3.5k stars) and Jetson LLM Serve (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Serving Capacity Planner?

BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 925 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.

Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.