Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
$ npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang sglang-prod-incident-triage --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/sglang-prod-incident-triage .claude/skills/sglang-prod-incident-triage && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "sglang-prod-incident-triage" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/sglang-prod-incident-triage into .claude/skills/sglang-prod-incident-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-prod-incident-triage", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/.agents/skills/sglang-prod-incident-triageType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang sglang-prod-incident-triage --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/sglang-prod-incident-triage .agents/skills/sglang-prod-incident-triage && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "sglang-prod-incident-triage" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/sglang-prod-incident-triage into .agents/skills/sglang-prod-incident-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-prod-incident-triage", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang sglang-prod-incident-triage --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/sglang-prod-incident-triage .cursor/skills/sglang-prod-incident-triage && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "sglang-prod-incident-triage" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/sglang-prod-incident-triage into .cursor/skills/sglang-prod-incident-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-prod-incident-triage", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path .agents/skills/sglang-prod-incident-triage--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang sglang-prod-incident-triage --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/sglang-prod-incident-triage .gemini/skills/sglang-prod-incident-triage && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "sglang-prod-incident-triage" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/sglang-prod-incident-triage into .gemini/skills/sglang-prod-incident-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-prod-incident-triage", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang sglang-prod-incident-triageInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/sglang-prod-incident-triage .github/skills/sglang-prod-incident-triage && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "sglang-prod-incident-triage" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/sglang-prod-incident-triage into .github/skills/sglang-prod-incident-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-prod-incident-triage", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang sglang-prod-incident-triage --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/sglang-prod-incident-triage .opencode/skills/sglang-prod-incident-triage && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "sglang-prod-incident-triage" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/sglang-prod-incident-triage into .opencode/skills/sglang-prod-incident-triage/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-prod-incident-triage", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
sglang-prod-incident-triageReplay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
Sglang Prod Incident Triage is an agent skill from sgl-project/sglang. Replay-first debug flow for SGLang serving problems. Use when a live or recent server shows health-check failures, latency or throughput regressions, queue growth, timeouts, distributed stalls, crash dumps, wrong outputs after deploys, or PD/EP/HiCache issues, and the job is to turn the problem into a replay plus the right next debug tool.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/case-studies.md`, `references/decision-tree.md` and `references/endpoints-and-signals.md`).
It sits in AI & LLM Engineering. It works with SGLang and CUDA. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit dab108b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3curlgitpythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
SGLANG_BEARER_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Sglang Prod Incident Triage loads about 2.1k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 92 tokens; SKILL.md has 887 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from sgl-project/sglang at commit dab108b, republished under its Apache-2.0 licence (© sgl-project). 887 words, ~2,134 tokens.
.claude/skills/sglang-prod-incident-triage/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.Use this skill to turn a live serving problem into a debug path you can replay.
Use one loop:
Do not start with profiling.
This skill should work with more focused skills instead of re-implementing them:
debug-cuda-crash when replay plus coredump points to a CUDA crash pathdebug-distributed-hang when the problem is clearly a TP/PP/DP/EP hangllm-torch-profiler-analysis when the issue is already narrowed to a
compute-side pathThree examples are included:
Return:
/health or /health_generate is unhealthyIf a live server is reachable, collect a read-only bundle before anything more intrusive:
python3 scripts/incident_artifact_tool.py collect-bundle \
--base-url http://127.0.0.1:30000 \
--outdir /tmp/incident_bundle
python3 scripts/incident_artifact_tool.py summarize-bundle \
/tmp/incident_bundleIf the server is protected:
python3 scripts/incident_artifact_tool.py collect-bundle \
--base-url http://127.0.0.1:30000 \
--token "$SGLANG_BEARER_TOKEN" \
--outdir /tmp/incident_bundleThe bundle script collects:
/health/health_generate/model_info/server_info/v1/loads?include=all/v1/loads?include=core,queues,disagg,spec/metrics/hicache/storage-backend on a best-effort basisUse the summary for a quick read on:
If the summary says the bundle was captured while the server was idle, recollect it during traffic or move quickly to dump plus replay.
If no live server is reachable, start from the best dump or log already available:
Read references/decision-tree.md only if the problem class is still unclear:
Then preserve the request payload that actually triggers the problem:
--crash-dump-folderDo not jump straight from a live symptom to low-level debugging without first saving something you can replay.
Read references/endpoints-and-signals.md when you need help reading the baseline bundle or the replay target.
Read references/replay-trace-profile.md when you need the replay, trace, profile, or bisect paths.
Standard order:
Use replay when:
If a crash dump exists, summarize it first:
python3 scripts/incident_artifact_tool.py summarize-dump \
--input-file /path/to/crash_dump.pklThen replay:
python3 /path/to/sglang/scripts/playground/replay_request_dump.py \
--input-file /path/to/crash_dump.pkl \
--host 127.0.0.1 \
--port 30000 \
--parallel 128If safe_pickle_load blocks a locally captured trusted dump, use:
python3 scripts/replay_trusted_request_dump.py \
--input-file /path/to/request_dump.pkl \
--host 127.0.0.1 \
--port 30000 \
--parallel 1If replay indicates a CUDA crash path, restart the same build with coredumps enabled before reproducing again:
SGLANG_CUDA_COREDUMP=1 \
SGLANG_CUDA_COREDUMP_DIR=/tmp/sglang_cuda_coredumps \
python -m sglang.launch_server \
--model-path ... \
--crash-dump-folder /tmp/sglang_crash_dump \
...Then inspect the generated coredump:
cuda-gdb "$(which python3)" \
-ex "target cudacore /tmp/sglang_cuda_coredumps/cuda_coredump_<host>.<pid>.<ts>"For a replay-first crash example, read references/case-studies.md.
Use tracing when:
If tracing was enabled at startup, you can change the level without restart:
curl "http://127.0.0.1:30000/set_trace_level?level=1"
curl "http://127.0.0.1:30000/set_trace_level?level=2"Use profiling when:
At that point, switch to llm-torch-profiler-analysis. Do not duplicate
its profiling workflow here.
For a low-noise latency example, read references/case-studies.md.
If this looks like a collective stall, save the failing request, replay it on a
clean target, collect the replay-time bundle and stacks, then switch to
debug-distributed-hang.
For an example of that flow, read references/case-studies.md.
If one commit is known-good and another is known-bad, build a deterministic harness before doing deeper manual debugging:
0 on good behavior and non-zero on bad behaviorgit bisect start <bad> <good>git bisect run <harness>Prefer replay-backed bisect when the regression depends on request shape or long-running serving state.
Switch tools once the fault class is clear:
llm-torch-profiler-analysis for kernel and overlap attributiondebug-distributed-hang for collective or rank-divergence hangsdebug-cuda-crash for CUDA crash reproduction and kernel API loggingDo not switch tools before collecting the first bundle unless the user already has decisive logs or dumps.
Load only what the current step needs:
safe_pickle_load blocks stock replayIf a live bundle was collected, include its path.
If replay, trace, or profiling was chosen, say why bundle plus dump were not enough.
© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, references) in .agents/skills/sglang-prod-incident-triage of sgl-project/sglang.
Open the folder on GitHubat commit dab108b
We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
Sglang Prod Incident Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Sglang Prod Incident Triage this skillsgl-project/sglang | 37k | 3 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.8k | Automated safety check: Pass | None | |
| Add Jit Kernelguqiong96/Lsglang | 144 | 1 repos | ~10k | Automated safety check: Pass | Apache-2.0 | |
| Magpie Kernel Evaluatoramd/skills | 408 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Jetson Memory AuditNVIDIA/skills | 3.6k | 1 repos | ~2.3k | Automated safety check: Notes | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
guqiong96/Lsglang
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
NVIDIA/skills
Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.
radixark/miles_diffusion
Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
sgl-project/sglang
Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…
Categories
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang. Sglang Prod Incident Triage is an agent skill from sgl-project/sglang. Replay-first debug flow for SGLang serving problems.
Sglang Prod Incident Triage fits situations like: recent server shows health-check failures; throughput regressions; distributed stalls; wrong outputs after deploys.
Run `npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a claude-code`. Or copy the skill folder (.agents/skills/sglang-prod-incident-triage in sgl-project/sglang) into .claude/skills/sglang-prod-incident-triage in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a codex`. Or copy the skill folder (.agents/skills/sglang-prod-incident-triage in sgl-project/sglang) into .agents/skills/sglang-prod-incident-triage in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-prod-incident-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-prod-incident-triage, .gemini/skills/sglang-prod-incident-triage, .github/skills/sglang-prod-incident-triage and .opencode/skills/sglang-prod-incident-triage in your project.
Going by SKILL.md and its folder, Sglang Prod Incident Triage needs Python for the scripts in its folder, the command-line tools its instructions call (python3, curl, git and python) and credentials named SGLANG_BEARER_TOKEN. Our summary lists: Python 3; A credential in SGLANG_BEARER_TOKEN.
SKILL.md contains no URLs. Its commands use curl and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Sglang Prod Incident Triage is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Sglang Prod Incident Triage: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Add Jit Kernel (guqiong96/Lsglang, 144 stars) and Magpie Kernel Evaluator (amd/skills, 408 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,973 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 11, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.