Agent skill

Serving Systems

by uw-syfi in uw-syfi/vibesys

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.

MITAuto-check passedAI & LLM Engineering

Install Serving Systems

skills CLI
$ npx skills add uw-syfi/vibesys --skill serving-systems -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install uw-syfi/vibesys serving-systems --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .claude/skills && cp -r skills-src/resources/skills/serving-systems .claude/skills/serving-systems && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
serving-systems
GitHub stars
103
Token cost
~2.9k tokens
SKILL.md length
991 words
Files
95 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.

  • Works in 4 steps: Read this file once to learn what's… → Open references/platforms/ first.… → For the active task, identify the one or… → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers How to use this skill, Start here: your platform's…, Portable contracts vs platform… and Reference index, plus 2 more sections
  • Calls git

What it does

Serving Systems is an agent skill from uw-syfi/vibesys. LLM and multimodal serving systems. Activate on inference servers, latency / throughput / TTFT / TPOT, KV-cache, batching, attention kernels, graph capture, speculative decoding, structured output, quantization, MoE, prefix caching, vision/speech/image/video serving, porting a model to vLLM / SGLang / TensorRT-LLM, or serving on NVIDIA, AMD ROCm, Apple Silicon (MLX), or Trainium (Neuron, NKI).

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 96 other files, including reference files (for example `CLAUDE.md`, `OVERVIEW.md` and `README.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Structured output and tool calling. It works with NVIDIA AI Platform, SGLang and vLLM. The repository describes itself as: Can AI Agents Build Bespoke Systems? The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Structured output and tool calling

Example prompts

  • “/serving-systems”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read this file once to learn what's covered.
  2. Open references/platforms/ first. Exactly one backend's directory is present — the one this run targets. Its floor.md is the optimization…
  3. For the active task, identify the one or two topics that match it (use the index below).
  4. Open references//.md directly with your file-read tool. Each is self-contained.

What it can do on your machine

Read from SKILL.md and the folder at commit c7784eb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Serving Systems loads about 2.9k tokens when it runs, and up to ~168k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 991 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~168k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from uw-syfi/vibesys at commit c7784eb, republished under its MIT licence (© uw-syfi). 991 words, ~2,884 tokens.

Download SKILL.mdSave it as .claude/skills/serving-systems/SKILL.md (or your agent's skills folder). This skill also uses 94 other files; get the full folder from GitHub.
name
serving-systems
description
LLM and multimodal serving systems. Activate on inference servers, latency / throughput / TTFT / TPOT, KV-cache, batching, attention kernels, graph capture, speculative decoding, structured output, quantization, MoE, prefix caching, vision/speech/image/video serving, porting a model to vLLM / SGLang / TensorRT-LLM, or serving on NVIDIA, AMD ROCm, Apple Silicon (MLX), or Trainium (Neuron, NKI).

serving-systems

This skill bundles the curated reference material for LLM and multimodal serving-system development as a topic library under references/. Open the specific reference whose topic matches the task; do not preload everything.

How to use this skill

  1. Read this file once to learn what's covered.
  2. Open references/platforms/ first. Exactly one backend's directory is present — the one this run targets. Its floor.md is the optimization floor for your hardware.
  3. For the active task, identify the one or two topics that match it (use the index below).
  4. Open references/<tier>/<topic>.md directly with your file-read tool. Each is self-contained.

Start here: your platform's floor

The default-on optimizations are not the same across hardware, and applying one platform's floor to another produces wrong work — eliminating padding is correct on NVIDIA and inverted on Trainium; graph capture is required on NVIDIA and does not exist on Apple Silicon.

Open references/platforms/<backend>/floor.md for the backend present in this workspace. Only that platform's directory is materialized, so there is no ambiguity about which applies.

Read past floor.md before writing a conclusion. Each platform directory also holds the measurement discipline: turning a counter capture into a bound verdict, proving which kernel library actually ran, and the A/B protocol for before/after claims. Before quoting a "compute-bound"/"bandwidth-bound" verdict, a percent-of-peak or percent-of-bandwidth number, or a kernel-library-tuning recommendation, use that discipline instead of estimating from assumed model geometry (weight bytes, layer split, KV head/dim): an assumption is a hypothesis, not a measurement.

Portable contracts vs platform implementations

Topics split into two kinds, and the distinction is load-bearing:

  • Contracts (algorithms/, models/, tooling/, frameworks/) state the problem, the invariants any implementation must satisfy, and the failure modes. These are the same on every backend.
  • Implementations (platforms/<backend>/) give the technique for specific hardware.

Where a contract has a platform implementation, the contract links to it. Read the contract first — it tells you what must be true; the platform file tells you how to get there here.

Reference index

Each entry is one file under references/. The bracketed phrase shows what triggers it.

Platforms

One directory per compute backend, each with floor.md, hardware.md, and profiler.md plus its own kernel and framework notes. Only the selected backend's directory is present.

Serving algorithms (portable contracts)
Show full SKILL.md (406 more words)Show less
Model architectures
Frameworks (cross-platform)

Platform-specific frameworks (MLX, torch-neuronx, NxD) live under that platform's directory.

Engine source maps

Written against NVIDIA-first upstream trees; ROCm paths exist in vLLM and SGLang but are not the primary codepath.

API / benchmark / profiler tooling

Out of scope

Kernel implementation (writing CUDA / Triton / CUTLASS / HIP). For that, use the separate agent-gpu-skills collection.

Exception — NKI: writing NeuronCore kernels for AWS Trainium is in scope here, via the bundled neuron-nki-* skills (neuron-nki-writing, -docs, -debugging, -profiling, -profile-querying); there is no separate Trainium kernel collection.

Reference repos

The repos/ directory (excluded from materialization to agents) holds full source trees of vLLM, SGLang, and TensorRT-LLM as git submodules. Engine-source-map references cite paths like $SERVE_REPOS/<engine>/...; export SERVE_REPOS=$(git rev-parse --show-toplevel)/resources/skills/serving-systems/repos or substitute inline.

© uw-syfi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 94 other files (references) in resources/skills/serving-systems of uw-syfi/vibesys.

  • SKILL.md
  • CLAUDE.md
  • OVERVIEW.md
  • README.md
  • references/algorithms/async-scheduling.md
  • references/algorithms/batched-sampling.md
  • references/algorithms/chunked-prefill.md
  • references/algorithms/continuous-batching.md
  • references/algorithms/cross-attention-kv-cache.md
  • references/algorithms/disaggregated-serving.md
  • references/algorithms/heterogeneous-kv-cache.md
  • references/algorithms/moe-routing-dispatch.md
  • references/algorithms/paged-attention.md
  • references/algorithms/parallelism.md
  • references/algorithms/quantization-schemes.md
  • references/algorithms/radix-prefix-caching.md
  • references/algorithms/speculative-decoding.md
  • references/algorithms/structured-output.md
  • references/engines
  • … and 76 more

Open the folder on GitHubat commit c7784eb

Compare with similar skills

Serving Systems next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Serving Systems compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Serving Systems this skilluw-syfi/vibesys103—~2.9kAutomated safety check: PassMIT
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone
Model PR History KnowledgeBBuf/AI-Infra-Auto-Driven-SKILLS925—~1.5kAutomated safety check: PassNone

Similar skills

  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Model PR History Knowledge

    BBuf/AI-Infra-Auto-Driven-SKILLS

    A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

    925 GitHub stars~1.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

    3.5k GitHub starsUsed in 1 repo~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from uw-syfi/vibesys

All 15 skills in this repo
  • Neuron Nki Profiling

    uw-syfi/vibesys

    This skill guides using the cli to generate NKI kernel profiles (NEFF + NTFF pairs) to analyze performance on Neuron hardware.

    103 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Neuron Nki Debugging

    uw-syfi/vibesys

    This skill guides debugging NKI compilation errors on Neuron hardware.

    103 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Neuron Nki Docs

    uw-syfi/vibesys

    Research NKI documentation for API lookups, tutorials, error codes, architecture, and optimization guides.

    103 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Query and analyze NKI kernel profile data from neuron-explorer parquet files.

    103 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Neuron Nki Writing

    uw-syfi/vibesys

    Guide for writing and modifying NKI kernels. An agent skill from uw-syfi/vibesys.

    103 GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Triage PRs

    uw-syfi/vibesys

    Triage the open pull requests of the VibeSys repository. An agent skill from uw-syfi/vibesys.

    103 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about Serving Systems

What does Serving Systems do?

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. Serving Systems is an agent skill from uw-syfi/vibesys. LLM and multimodal serving systems.

When should I use Serving Systems?

Serving Systems fits situations like: tasks that involve LLM inference and serving; tasks that involve Structured output and tool calling.

How do I install Serving Systems in Claude Code?

Run `npx skills add uw-syfi/vibesys --skill serving-systems -a claude-code`. Or copy the skill folder (resources/skills/serving-systems in uw-syfi/vibesys) into .claude/skills/serving-systems in your project. Claude Code loads it when a task matches its description.

How do I install Serving Systems in Codex?

Run `npx skills add uw-syfi/vibesys --skill serving-systems -a codex`. Or copy the skill folder (resources/skills/serving-systems in uw-syfi/vibesys) into .agents/skills/serving-systems in your project. Codex loads it when a task matches its description.

Can I use Serving Systems in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add uw-syfi/vibesys --skill serving-systems -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-systems, .gemini/skills/serving-systems, .github/skills/serving-systems and .opencode/skills/serving-systems in your project.

What does Serving Systems need to run?

Going by SKILL.md and its folder, Serving Systems needs the command-line tools its instructions call (git).

Does Serving Systems access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Serving Systems safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Serving Systems use?

Serving Systems is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Serving Systems use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 165k tokens, read only when the agent opens those files.

What are the alternatives to Serving Systems?

Skills that share tags, products or a category with Serving Systems: SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Serving Systems?

uw-syfi (a GitHub organization) maintains it in uw-syfi/vibesys, which has 103 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.

Source: uw-syfi/vibesys on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.