SGLang Structured Serving
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.
$ npx skills add uw-syfi/vibesys --skill serving-systems -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install uw-syfi/vibesys serving-systems --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .claude/skills && cp -r skills-src/resources/skills/serving-systems .claude/skills/serving-systems && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "serving-systems" agent skill from https://github.com/uw-syfi/vibesys/tree/main/resources/skills/serving-systems into .claude/skills/serving-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-systems", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/uw-syfi/vibesys/tree/main/resources/skills/serving-systemsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add uw-syfi/vibesys --skill serving-systems -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install uw-syfi/vibesys serving-systems --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .agents/skills && cp -r skills-src/resources/skills/serving-systems .agents/skills/serving-systems && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "serving-systems" agent skill from https://github.com/uw-syfi/vibesys/tree/main/resources/skills/serving-systems into .agents/skills/serving-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-systems", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add uw-syfi/vibesys --skill serving-systems -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install uw-syfi/vibesys serving-systems --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/resources/skills/serving-systems .cursor/skills/serving-systems && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "serving-systems" agent skill from https://github.com/uw-syfi/vibesys/tree/main/resources/skills/serving-systems into .cursor/skills/serving-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-systems", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/uw-syfi/vibesys.git --path resources/skills/serving-systems--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add uw-syfi/vibesys --skill serving-systems -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install uw-syfi/vibesys serving-systems --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/resources/skills/serving-systems .gemini/skills/serving-systems && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "serving-systems" agent skill from https://github.com/uw-syfi/vibesys/tree/main/resources/skills/serving-systems into .gemini/skills/serving-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-systems", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install uw-syfi/vibesys serving-systemsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add uw-syfi/vibesys --skill serving-systems -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .github/skills && cp -r skills-src/resources/skills/serving-systems .github/skills/serving-systems && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "serving-systems" agent skill from https://github.com/uw-syfi/vibesys/tree/main/resources/skills/serving-systems into .github/skills/serving-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-systems", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add uw-syfi/vibesys --skill serving-systems -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install uw-syfi/vibesys serving-systems --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/resources/skills/serving-systems .opencode/skills/serving-systems && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "serving-systems" agent skill from https://github.com/uw-syfi/vibesys/tree/main/resources/skills/serving-systems into .opencode/skills/serving-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-systems", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
serving-systemsLLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.
Serving Systems is an agent skill from uw-syfi/vibesys. LLM and multimodal serving systems. Activate on inference servers, latency / throughput / TTFT / TPOT, KV-cache, batching, attention kernels, graph capture, speculative decoding, structured output, quantization, MoE, prefix caching, vision/speech/image/video serving, porting a model to vLLM / SGLang / TensorRT-LLM, or serving on NVIDIA, AMD ROCm, Apple Silicon (MLX), or Trainium (Neuron, NKI).
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 96 other files, including reference files (for example `CLAUDE.md`, `OVERVIEW.md` and `README.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving and Structured output and tool calling. It works with NVIDIA AI Platform, SGLang and vLLM. The repository describes itself as: Can AI Agents Build Bespoke Systems? The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c7784eb. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Serving Systems loads about 2.9k tokens when it runs, and up to ~168k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 991 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from uw-syfi/vibesys at commit c7784eb, republished under its MIT licence (© uw-syfi). 991 words, ~2,884 tokens.
.claude/skills/serving-systems/SKILL.md (or your agent's skills folder). This skill also uses 94 other files; get the full folder from GitHub.This skill bundles the curated reference material for LLM and multimodal serving-system development as a topic library under references/. Open the specific reference whose topic matches the task; do not preload everything.
references/platforms/ first. Exactly one backend's directory is present — the one this run targets. Its floor.md is the optimization floor for your hardware.references/<tier>/<topic>.md directly with your file-read tool. Each is self-contained.The default-on optimizations are not the same across hardware, and applying one platform's floor to another produces wrong work — eliminating padding is correct on NVIDIA and inverted on Trainium; graph capture is required on NVIDIA and does not exist on Apple Silicon.
Open references/platforms/<backend>/floor.md for the backend present in this workspace. Only that platform's directory is materialized, so there is no ambiguity about which applies.
Read past floor.md before writing a conclusion. Each platform directory also holds the measurement discipline: turning a counter capture into a bound verdict, proving which kernel library actually ran, and the A/B protocol for before/after claims. Before quoting a "compute-bound"/"bandwidth-bound" verdict, a percent-of-peak or percent-of-bandwidth number, or a kernel-library-tuning recommendation, use that discipline instead of estimating from assumed model geometry (weight bytes, layer split, KV head/dim): an assumption is a hypothesis, not a measurement.
Topics split into two kinds, and the distinction is load-bearing:
algorithms/, models/, tooling/, frameworks/) state the problem, the invariants any implementation must satisfy, and the failure modes. These are the same on every backend.platforms/<backend>/) give the technique for specific hardware.Where a contract has a platform implementation, the contract links to it. Read the contract first — it tells you what must be true; the platform file tells you how to get there here.
Each entry is one file under references/. The bracketed phrase shows what triggers it.
One directory per compute backend, each with floor.md, hardware.md, and profiler.md plus its own kernel and framework notes. Only the selected backend's directory is present.
references/platforms/ — start at floor.md.references/algorithms/async-scheduling.md — Hide host scheduler work behind accelerator compute. Contract; mechanism is per-platform.
references/algorithms/batched-sampling.md — Per-request sampling parameters in one kernel pipeline, without per-request host sync.
references/algorithms/chunked-prefill.md — Split long prompts into chunks interleaved with decode, preventing a long prefill from stalling decode latency.
references/algorithms/continuous-batching.md — Requests join a running generation loop between steps. Contract; the KV strategy inverts between backends.
references/algorithms/cross-attention-kv-cache.md — Cross-attention KV cache for encoder-decoder decode (Whisper, mllama): compute encoder-context K/V once at prefill, read every step. Non-causal, no RoPE, separate pool.
references/algorithms/disaggregated-serving.md — Separate prefill and decode worker pools with KV transfer between them.
references/algorithms/heterogeneous-kv-cache.md — Memory management and prefix caching for hybrid models (full-attn + sliding-window, attention + SSM/Mamba, attention + linear).
references/algorithms/moe-routing-dispatch.md — MoE routing and dispatch — top-k gating, token-to-expert dispatch/combine, grouped-GEMM expert FFN, expert parallelism, expert load balancing.
references/algorithms/paged-attention.md — Block-based non-contiguous KV storage with a page table per request. Applies where the backend has a discrete memory pool.
references/algorithms/parallelism.md — TP, PP, EP, DP, SP and combinations. Multi-device backends only.
references/algorithms/quantization-schemes.md — Precision, granularity, calibration, checkpoint layout. Hardware support is generation-gated.
references/algorithms/radix-prefix-caching.md — Share KV cache across requests with common prefixes via a radix tree with LRU eviction.
references/algorithms/speculative-decoding.md — Draft proposals verified in one target pass. Contract; variable accept length is handled per-platform.
references/algorithms/structured-output.md — Grammar-guided decoding (XGrammar, Outlines, llguidance), JSON mode, regex constraints, tool calling, logits biasing.
references/models/attention-variants.md — Attention variants across three axes: head sharing (MHA / MQA / GQA / MLA), masking pattern, complexity class.
references/models/image-generation.md — Image generation serving — diffusion (U-Net, DiT) and flow-matching.
references/models/omni-multimodal.md — Omni-modal serving — multi-modality in AND out.
references/models/speech-generation.md — Speech generation serving — TTS and speech-to-speech.
references/models/speech-language.md — Speech-language serving — ASR, speech translation, audio-text chat.
references/models/ssm-hybrid.md — State-space and hybrid SSM+attention serving — Mamba/Mamba2, Jamba, Zamba, Nemotron-H, Falcon-Mamba.
references/models/text-dense.md — The foundational architecture most modern LLMs build on.
references/models/text-moe.md — Mixture-of-Experts text decoders — Mixtral, DeepSeek V2/V3/R1, Qwen3-MoE, Llama-4.
references/models/video-generation.md — Video generation serving — diffusion with 3D attention, large activations.
references/models/vision-language.md — Vision-language serving — LLaVA, Qwen-VL, InternVL, mllama, Molmo, DeepSeek-VL.
references/frameworks/pytorch.md — PyTorch idioms for serving — weight loading, torch.compile, state_dict remapping, custom ops, inference_mode.
references/frameworks/triton.md — Triton as a framework-level decision — when a custom Triton kernel pays off vs reusing an existing kernel library.
Platform-specific frameworks (MLX, torch-neuronx, NxD) live under that platform's directory.
Written against NVIDIA-first upstream trees; ROCm paths exist in vLLM and SGLang but are not the primary codepath.
references/engines/sglang.md — SGLang source-code lookup.references/engines/trtllm.md — TensorRT-LLM source-code lookup.references/engines/vllm.md — vLLM source-code lookup.references/engines/vllm-profiling.md — vLLM profiling capture quirks: multi-process default, offline single-process capture, torch.profiler interface, post-capture hang.references/tooling/accuracy-checker.md — Verify a custom generation implementation against HuggingFace model.generate().
references/tooling/fastapi-serving.md — Production-ready FastAPI inference server for HuggingFace models.
references/tooling/io-handling.md — Tokenization and chat templates, image/video/audio preprocessing, detokenization and UTF-8-safe streaming, tool-call parsing.
references/tooling/lora-serving.md — Multi-adapter LoRA serving — one base model dispatching different adapters per request.
references/tooling/openai-api.md — OpenAI-compatible HTTP per modality — text, image, TTS, STT, video, realtime audio.
references/tooling/performance-modeling.md — Analytical serving-performance modeling — roofline, Amdahl bounds, end-to-end time accounting, architecture ceilings, profiler calibration, and plateau-driven hypothesis selection.
references/tooling/profiler.md — Profiling discipline and altitudes. The contract is portable; the concrete toolchain is per-platform.
references/tooling/profiling-serving-engines.md — Index: how to fill the generic capture tools' lifecycle arguments (command, env, ready_command, stop_signal, grace_s) for a serving engine, linking to each engine's own profiling file under engines/.
references/tooling/serving-benchmark.md — Benchmark an LLM serving endpoint — TTFT, TPOT, ITL, end-to-end latency, throughput, p50/p95/p99 across concurrency and ISL/OSL sweeps.
Kernel implementation (writing CUDA / Triton / CUTLASS / HIP). For that, use the separate agent-gpu-skills collection.
Exception — NKI: writing NeuronCore kernels for AWS Trainium is in scope here, via the bundled neuron-nki-* skills (neuron-nki-writing, -docs, -debugging, -profiling, -profile-querying); there is no separate Trainium kernel collection.
The repos/ directory (excluded from materialization to agents) holds full source trees of vLLM, SGLang, and TensorRT-LLM as git submodules. Engine-source-map references cite paths like $SERVE_REPOS/<engine>/...; export SERVE_REPOS=$(git rev-parse --show-toplevel)/resources/skills/serving-systems/repos or substitute inline.
© uw-syfi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 94 other files (references) in resources/skills/serving-systems of uw-syfi/vibesys.
Open the folder on GitHubat commit c7784eb
Serving Systems next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Serving Systems this skilluw-syfi/vibesys | 103 | — | ~2.9k | Automated safety check: Pass | MIT | |
| SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~2.8k | Automated safety check: Pass | None | |
| Model PR History KnowledgeBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~1.5k | Automated safety check: Pass | None |
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
BBuf/AI-Infra-Auto-Driven-SKILLS
A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.
NVIDIA/skills
Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.
uw-syfi/vibesys
This skill guides using the cli to generate NKI kernel profiles (NEFF + NTFF pairs) to analyze performance on Neuron hardware.
uw-syfi/vibesys
This skill guides debugging NKI compilation errors on Neuron hardware.
uw-syfi/vibesys
Research NKI documentation for API lookups, tutorials, error codes, architecture, and optimization guides.
uw-syfi/vibesys
Query and analyze NKI kernel profile data from neuron-explorer parquet files.
uw-syfi/vibesys
Guide for writing and modifying NKI kernels. An agent skill from uw-syfi/vibesys.
uw-syfi/vibesys
Triage the open pull requests of the VibeSys repository. An agent skill from uw-syfi/vibesys.
Works with
Categories
LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. Serving Systems is an agent skill from uw-syfi/vibesys. LLM and multimodal serving systems.
Serving Systems fits situations like: tasks that involve LLM inference and serving; tasks that involve Structured output and tool calling.
Run `npx skills add uw-syfi/vibesys --skill serving-systems -a claude-code`. Or copy the skill folder (resources/skills/serving-systems in uw-syfi/vibesys) into .claude/skills/serving-systems in your project. Claude Code loads it when a task matches its description.
Run `npx skills add uw-syfi/vibesys --skill serving-systems -a codex`. Or copy the skill folder (resources/skills/serving-systems in uw-syfi/vibesys) into .agents/skills/serving-systems in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add uw-syfi/vibesys --skill serving-systems -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-systems, .gemini/skills/serving-systems, .github/skills/serving-systems and .opencode/skills/serving-systems in your project.
Going by SKILL.md and its folder, Serving Systems needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Serving Systems is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 165k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Serving Systems: SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
uw-syfi (a GitHub organization) maintains it in uw-syfi/vibesys, which has 103 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.
Source: uw-syfi/vibesys on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.