Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-capacity-planner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-serving-capacity-planner .claude/skills/llm-serving-capacity-planner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "llm-serving-capacity-planner" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/llm-serving-capacity-planner into .claude/skills/llm-serving-capacity-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-serving-capacity-planner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/llm-serving-capacity-plannerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-capacity-planner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/llm-serving-capacity-planner .agents/skills/llm-serving-capacity-planner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "llm-serving-capacity-planner" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/llm-serving-capacity-planner into .agents/skills/llm-serving-capacity-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-serving-capacity-planner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-capacity-planner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/llm-serving-capacity-planner .cursor/skills/llm-serving-capacity-planner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "llm-serving-capacity-planner" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/llm-serving-capacity-planner into .cursor/skills/llm-serving-capacity-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-serving-capacity-planner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git --path skills/llm-serving-capacity-planner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-capacity-planner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/llm-serving-capacity-planner .gemini/skills/llm-serving-capacity-planner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "llm-serving-capacity-planner" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/llm-serving-capacity-planner into .gemini/skills/llm-serving-capacity-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-serving-capacity-planner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-capacity-plannerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/llm-serving-capacity-planner .github/skills/llm-serving-capacity-planner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "llm-serving-capacity-planner" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/llm-serving-capacity-planner into .github/skills/llm-serving-capacity-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-serving-capacity-planner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-serving-capacity-planner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/llm-serving-capacity-planner .opencode/skills/llm-serving-capacity-planner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "llm-serving-capacity-planner" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/llm-serving-capacity-planner into .opencode/skills/llm-serving-capacity-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-serving-capacity-planner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
llm-serving-capacity-plannerReads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.
This skill analyzes the startup log of an SGLang or vLLM server to explain how GPU memory is divided. A script, scripts/capacity_analyzer.py, pulls out the weight load, KV cache pool, CUDA graph, framework overhead and token-capacity lines, then estimates concurrent requests for common request lengths. If you give none, it uses 4096, 6144 and 8192 tokens.
The log file is the only required input. The GPU type can be detected from the log, nvidia-smi output adds per-rank memory for cross-checking, and the model's config.json lets the skill calculate KV cache bytes in theory. Reference files hold GPU specs and log patterns. For DeepSeek-V4.1 it warns that per-token byte figures are payload sizes, not allocation after page padding, and points to a separate KV layout reference to read first.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3dockermambaFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
LLM Serving Capacity Planner loads about 2.5k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 1,238 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Without a licence we can't republish the file, so here is its outline and opening line. It has 1,238 words (~2,518 tokens).
“Use this when a serving log has enough memory lines to explain where GPU HBM went. The analyzer reads SGLang/vLLM startup logs, extracts weight load, KV pool, CUDA graph, framework overhead, and token-capacity lines, then estimates concurrent requests for common…”
SKILL.md and 3 other files (scripts, references) in skills/llm-serving-capacity-planner of BBuf/AI-Infra-Auto-Driven-SKILLS.
Open the folder on GitHubat commit 6dc9c66
LLM Serving Capacity Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLM Serving Capacity Planner this skillBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~2.5k | Automated safety check: Pass | None | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Magpie Kernel Evaluatoramd/skills | 406 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Jetson Inference Mem TuneNVIDIA/skills | 3.5k | 1 repos | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Jetson LLM ServeNVIDIA/skills | 3.5k | 1 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
NVIDIA/skills
Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.
NVIDIA/skills
Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
BBuf/AI-Infra-Auto-Driven-SKILLS
Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
BBuf/AI-Infra-Auto-Driven-SKILLS
Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.
BBuf/AI-Infra-Auto-Driven-SKILLS
Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.
BBuf/AI-Infra-Auto-Driven-SKILLS
Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.
BBuf/AI-Infra-Auto-Driven-SKILLS
Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.
Categories
Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. This skill analyzes the startup log of an SGLang or vLLM server to explain how GPU memory is divided.py, pulls out the weight load, KV cache pool, CUDA graph, framework overhead and token-capacity lines, then estimates concurrent requests for common request lengths.
LLM Serving Capacity Planner fits situations like: explaining where GPU memory went after an SGLang or vLLM server starts; triaging an out-of-memory failure from a serving log; comparing mem-fraction-static settings and their effect on the KV cache budget; estimating maximum concurrent requests at a given token length.
Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a claude-code`. Or copy the skill folder (skills/llm-serving-capacity-planner in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/llm-serving-capacity-planner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a codex`. Or copy the skill folder (skills/llm-serving-capacity-planner in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/llm-serving-capacity-planner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-serving-capacity-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-serving-capacity-planner, .gemini/skills/llm-serving-capacity-planner, .github/skills/llm-serving-capacity-planner and .opencode/skills/llm-serving-capacity-planner in your project.
Going by SKILL.md and its folder, LLM Serving Capacity Planner needs Python for the scripts in its folder and the command-line tools its instructions call (python3, docker and mamba). Our summary lists: A SGLang or vLLM serving startup log; Optional nvidia-smi output and the model's config.json.
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
No licence was found for LLM Serving Capacity Planner or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with LLM Serving Capacity Planner: Graphsignal (graphsignal/graphsignal, 257 stars), Magpie Kernel Evaluator (amd/skills, 406 stars), Jetson Inference Mem Tune (NVIDIA/skills, 3.5k stars) and Jetson LLM Serve (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 925 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.
Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.