Agent skill

Sglang Prod Incident Triage

by sgl-project in sgl-project/sglang

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Sglang Prod Incident Triage

skills CLI
$ npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang sglang-prod-incident-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/sglang-prod-incident-triage .claude/skills/sglang-prod-incident-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sglang-prod-incident-triage
GitHub stars
37k
Used in
3 other repos
Token cost
~2.1k tokens
SKILL.md length
887 words
Files
7 (incl. scripts, references)
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

  • Works in 5 steps: Collect a baseline bundle → Save the failing request → Replay on a clean target → …
  • Recent server shows health-check failures
  • SKILL.md covers Overview, Output Contract, When To Use It and Workflow, plus 2 more sections
  • Runs Python scripts from its folder; calls python3, curl and git; needs SGLANG_BEARER_TOKEN

What it does

Sglang Prod Incident Triage is an agent skill from sgl-project/sglang. Replay-first debug flow for SGLang serving problems. Use when a live or recent server shows health-check failures, latency or throughput regressions, queue growth, timeouts, distributed stalls, crash dumps, wrong outputs after deploys, or PD/EP/HiCache issues, and the job is to turn the problem into a replay plus the right next debug tool.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/case-studies.md`, `references/decision-tree.md` and `references/endpoints-and-signals.md`).

It sits in AI & LLM Engineering. It works with SGLang and CUDA. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Recent server shows health-check failures
  • Throughput regressions
  • Distributed stalls
  • Wrong outputs after deploys

Example prompts

  • “/sglang-prod-incident-triage”

Requirements

  • Python 3
  • A credential in SGLANG_BEARER_TOKEN

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Collect a baseline bundle
  2. Save the failing request
  3. Replay on a clean target
  4. Only go deeper after replay
  5. Switch tools when the boundary is clear

What it can do on your machine

Read from SKILL.md and the folder at commit dab108b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • curl
    • git
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SGLANG_BEARER_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sglang Prod Incident Triage loads about 2.1k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 92 tokens; SKILL.md has 887 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit dab108b, republished under its Apache-2.0 licence (© sgl-project). 887 words, ~2,134 tokens.

Download SKILL.mdSave it as .claude/skills/sglang-prod-incident-triage/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
sglang-prod-incident-triage
description
Replay-first debug flow for SGLang serving problems. Use when a live or recent server shows health-check failures, latency or throughput regressions, queue growth, timeouts, distributed stalls, crash dumps, wrong outputs after deploys, or PD/EP/HiCache issues, and the job is to turn the problem into a replay plus the right next debug tool.

SGLang Serving Debug

Overview

Use this skill to turn a live serving problem into a debug path you can replay.

Use one loop:

  • collect a baseline bundle
  • save the failing request or crash dump
  • replay on a clean target
  • only then switch tools

Do not start with profiling.

This skill should work with more focused skills instead of re-implementing them:

  • debug-cuda-crash when replay plus coredump points to a CUDA crash path
  • debug-distributed-hang when the problem is clearly a TP/PP/DP/EP hang
  • llm-torch-profiler-analysis when the issue is already narrowed to a compute-side path

Three examples are included:

  • TTFT spike with low queue time
  • replay-first CUDA crash flow
  • request-shaped distributed hang flow

Output Contract

Return:

  • problem class
  • what was checked
  • strongest signal so far
  • current best guess
  • what was ruled out
  • next step
  • production risk

When To Use It

  • /health or /health_generate is unhealthy
  • latency or throughput regressed under serving load
  • queue size grows while health still looks green
  • one request class times out or hangs
  • the server crashes only after some requests
  • outputs changed after a deploy, topology change, or weight switch
  • one older commit is known-good and a newer commit is known-bad

Workflow

1. Collect a baseline bundle

If a live server is reachable, collect a read-only bundle before anything more intrusive:

bash
python3 scripts/incident_artifact_tool.py collect-bundle \
  --base-url http://127.0.0.1:30000 \
  --outdir /tmp/incident_bundle

python3 scripts/incident_artifact_tool.py summarize-bundle \
  /tmp/incident_bundle

If the server is protected:

bash
python3 scripts/incident_artifact_tool.py collect-bundle \
  --base-url http://127.0.0.1:30000 \
  --token "$SGLANG_BEARER_TOKEN" \
  --outdir /tmp/incident_bundle

The bundle script collects:

  • /health
  • /health_generate
  • /model_info
  • /server_info
  • /v1/loads?include=all
  • /v1/loads?include=core,queues,disagg,spec
  • /metrics
  • /hicache/storage-backend on a best-effort basis

Use the summary for a quick read on:

  • health vs. active health state
  • topology and runtime flags
  • point-in-time queue and token usage
  • TTFT / E2E / queue-time heuristics from Prometheus metrics

If the summary says the bundle was captured while the server was idle, recollect it during traffic or move quickly to dump plus replay.

If no live server is reachable, start from the best dump or log already available:

  • crash dump
  • request dump
  • logs
  • CUDA coredump
  • OTel trace
  • torch profile
2. Save the failing request

Read references/decision-tree.md only if the problem class is still unclear:

  • server down or unhealthy
  • latency or throughput regression
  • wrong output or behavior regression
  • intermittent timeout or hang

Then preserve the request payload that actually triggers the problem:

  • crash path: use --crash-dump-folder
  • non-crash path: enable request dump or save the exact trigger request

Do not jump straight from a live symptom to low-level debugging without first saving something you can replay.

3. Replay on a clean target

Read references/endpoints-and-signals.md when you need help reading the baseline bundle or the replay target.

Read references/replay-trace-profile.md when you need the replay, trace, profile, or bisect paths.

Standard order:

  1. collect baseline bundle
  2. capture request dump or crash dump
  3. restart a clean debug target if needed
  4. replay the same issue
  5. collect replay-time logs and dumps
4. Only go deeper after replay
Replay

Use replay when:

  • a crash dump exists
  • a request dump exists
  • the problem depends on request shape or workload mix

If a crash dump exists, summarize it first:

bash
python3 scripts/incident_artifact_tool.py summarize-dump \
  --input-file /path/to/crash_dump.pkl

Then replay:

bash
python3 /path/to/sglang/scripts/playground/replay_request_dump.py \
  --input-file /path/to/crash_dump.pkl \
  --host 127.0.0.1 \
  --port 30000 \
  --parallel 128

If safe_pickle_load blocks a locally captured trusted dump, use:

bash
python3 scripts/replay_trusted_request_dump.py \
  --input-file /path/to/request_dump.pkl \
  --host 127.0.0.1 \
  --port 30000 \
  --parallel 1

If replay indicates a CUDA crash path, restart the same build with coredumps enabled before reproducing again:

bash
SGLANG_CUDA_COREDUMP=1 \
SGLANG_CUDA_COREDUMP_DIR=/tmp/sglang_cuda_coredumps \
python -m sglang.launch_server \
  --model-path ... \
  --crash-dump-folder /tmp/sglang_crash_dump \
  ...

Then inspect the generated coredump:

bash
cuda-gdb "$(which python3)" \
  -ex "target cudacore /tmp/sglang_cuda_coredumps/cuda_coredump_<host>.<pid>.<ts>"

For a replay-first crash example, read references/case-studies.md.

Show full SKILL.md (362 more words)Show less
OTel trace

Use tracing when:

  • request-stage timing is unclear
  • router vs. worker attribution is unclear
  • PD prefill/decode transfer may be implicated

If tracing was enabled at startup, you can change the level without restart:

bash
curl "http://127.0.0.1:30000/set_trace_level?level=1"
curl "http://127.0.0.1:30000/set_trace_level?level=2"
Torch profile

Use profiling when:

  • the issue is already narrowed to compute-side ownership
  • replay already reproduces the problem
  • metrics and loads do not explain the regression

At that point, switch to llm-torch-profiler-analysis. Do not duplicate its profiling workflow here.

For a low-noise latency example, read references/case-studies.md.

Distributed hang

If this looks like a collective stall, save the failing request, replay it on a clean target, collect the replay-time bundle and stacks, then switch to debug-distributed-hang.

For an example of that flow, read references/case-studies.md.

Regression between two commits

If one commit is known-good and another is known-bad, build a deterministic harness before doing deeper manual debugging:

  1. choose a stable reproducer: request replay, benchmark command, or correctness check
  2. make the harness return 0 on good behavior and non-zero on bad behavior
  3. run git bisect start <bad> <good>
  4. run git bisect run <harness>
  5. return here only after a candidate commit is isolated

Prefer replay-backed bisect when the regression depends on request shape or long-running serving state.

6. Switch tools when the boundary is clear

Switch tools once the fault class is clear:

  • llm-torch-profiler-analysis for kernel and overlap attribution
  • debug-distributed-hang for collective or rank-divergence hangs
  • debug-cuda-crash for CUDA crash reproduction and kernel API logging

Do not switch tools before collecting the first bundle unless the user already has decisive logs or dumps.

References

Load only what the current step needs:

Scripts

If a live bundle was collected, include its path.

If replay, trace, or profiling was chosen, say why bundle plus dump were not enough.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in .agents/skills/sglang-prod-incident-triage of sgl-project/sglang.

  • SKILL.md
  • references/case-studies.md
  • references/decision-tree.md
  • references/endpoints-and-signals.md
  • references/replay-trace-profile.md
  • scripts/incident_artifact_tool.py
  • scripts/replay_trusted_request_dump.py

Open the folder on GitHubat commit dab108b

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sglang Prod Incident Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sglang Prod Incident Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sglang Prod Incident Triage this skillsgl-project/sglang37k3 repos~2.1kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone
Add Jit Kernelguqiong96/Lsglang1441 repos~10kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills408—~2.3kAutomated safety check: PassMIT
Jetson Memory AuditNVIDIA/skills3.6k1 repos~2.3kAutomated safety check: NotesApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Add Jit Kernel

    guqiong96/Lsglang

    Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

    144 GitHub starsUsed in 1 repo~10k tokens
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    408 GitHub stars~2.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Jetson Memory Audit

    NVIDIA/skills

    Official

    Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

    3.6k GitHub starsUsed in 1 repo~2.3k tokens
    AI & LLM EngineeringAuto-check: notes
  • Install Miles Diffusion

    radixark/miles_diffusion

    Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them.

    110 GitHub stars~1.6k tokensUpdated today
    DevOps & CloudAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Kl Consistency Test

    sgl-project/sglang

    Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…

    37k GitHub starsUsed in 2 repos~3.7k tokens
    Auto-check passed

Works with

Questions about Sglang Prod Incident Triage

What does Sglang Prod Incident Triage do?

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang. Sglang Prod Incident Triage is an agent skill from sgl-project/sglang. Replay-first debug flow for SGLang serving problems.

When should I use Sglang Prod Incident Triage?

Sglang Prod Incident Triage fits situations like: recent server shows health-check failures; throughput regressions; distributed stalls; wrong outputs after deploys.

How do I install Sglang Prod Incident Triage in Claude Code?

Run `npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a claude-code`. Or copy the skill folder (.agents/skills/sglang-prod-incident-triage in sgl-project/sglang) into .claude/skills/sglang-prod-incident-triage in your project. Claude Code loads it when a task matches its description.

How do I install Sglang Prod Incident Triage in Codex?

Run `npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a codex`. Or copy the skill folder (.agents/skills/sglang-prod-incident-triage in sgl-project/sglang) into .agents/skills/sglang-prod-incident-triage in your project. Codex loads it when a task matches its description.

Can I use Sglang Prod Incident Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-prod-incident-triage, .gemini/skills/sglang-prod-incident-triage, .github/skills/sglang-prod-incident-triage and .opencode/skills/sglang-prod-incident-triage in your project.

What does Sglang Prod Incident Triage need to run?

Going by SKILL.md and its folder, Sglang Prod Incident Triage needs Python for the scripts in its folder, the command-line tools its instructions call (python3, curl, git and python) and credentials named SGLANG_BEARER_TOKEN. Our summary lists: Python 3; A credential in SGLANG_BEARER_TOKEN.

Does Sglang Prod Incident Triage access the network?

SKILL.md contains no URLs. Its commands use curl and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sglang Prod Incident Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sglang Prod Incident Triage use?

Sglang Prod Incident Triage is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sglang Prod Incident Triage use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.

What are the alternatives to Sglang Prod Incident Triage?

Skills that share tags, products or a category with Sglang Prod Incident Triage: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Add Jit Kernel (guqiong96/Lsglang, 144 stars) and Magpie Kernel Evaluator (amd/skills, 408 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sglang Prod Incident Triage?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,973 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 11, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.