Cuda Cpp Kernel
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging
$ npx skills add sgl-project/sglang --skill debug-cuda-crash -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang debug-cuda-crash --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/debug-cuda-crash .claude/skills/debug-cuda-crash && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "debug-cuda-crash" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-cuda-crash into .claude/skills/debug-cuda-crash/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-cuda-crash", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-cuda-crashType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill debug-cuda-crash -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang debug-cuda-crash --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/debug-cuda-crash .agents/skills/debug-cuda-crash && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "debug-cuda-crash" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-cuda-crash into .agents/skills/debug-cuda-crash/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-cuda-crash", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill debug-cuda-crash -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang debug-cuda-crash --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/debug-cuda-crash .cursor/skills/debug-cuda-crash && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "debug-cuda-crash" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-cuda-crash into .cursor/skills/debug-cuda-crash/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-cuda-crash", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path .agents/skills/debug-cuda-crash--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill debug-cuda-crash -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang debug-cuda-crash --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/debug-cuda-crash .gemini/skills/debug-cuda-crash && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "debug-cuda-crash" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-cuda-crash into .gemini/skills/debug-cuda-crash/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-cuda-crash", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang debug-cuda-crashInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill debug-cuda-crash -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/debug-cuda-crash .github/skills/debug-cuda-crash && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "debug-cuda-crash" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-cuda-crash into .github/skills/debug-cuda-crash/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-cuda-crash", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill debug-cuda-crash -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang debug-cuda-crash --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/debug-cuda-crash .opencode/skills/debug-cuda-crash && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "debug-cuda-crash" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/debug-cuda-crash into .opencode/skills/debug-cuda-crash/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debug-cuda-crash", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
debug-cuda-crashCall this skill when you need to debug CUDA crashes in SGLang using kernel API logging
Debug Cuda Crash is an agent skill from sgl-project/sglang. Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Deep learning and Debugging. It works with CUDA, SGLang and Qwen. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b7b2975. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Debug Cuda Crash loads about 4.9k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 1,348 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sgl-project/sglang at commit b7b2975, republished under its Apache-2.0 licence (© sgl-project). 1,348 words, ~4,926 tokens.
.claude/skills/debug-cuda-crash/SKILL.md (or your agent's skills folder).This tutorial shows you how to debug CUDA crashes and errors in SGLang using the @debug_kernel_api logging decorator.
When your code crashes with CUDA errors such as illegal memory access, device-side assert, out-of-bounds, or NaN/Inf, use kernel API logging to:
Problem: CUDA errors often crash the program before normal debugging output is flushed.
Solution: SGLang's @debug_kernel_api decorator logs inputs before execution, so you can still see what caused the crash even after the program aborts.
The current logging coverage focuses on the highest-value kernel boundaries in SGLang:
register_custom_op(...)register_custom_op_from_extern(...)torch.ops.sglang.* hotspots and model-specific bypassesThis means the logging is useful for both LLM and diffusion kernel debugging, but it does not automatically cover every pure PyTorch call in the repository.
export SGLANG_KERNEL_API_LOGLEVEL=1
export SGLANG_KERNEL_API_LOGDEST=stdout
python my_script.pyOutput:
================================================================================
[2026-03-19 00:47:06] SGLang Kernel API Call: RMSNorm.forward
================================================================================
[2026-03-19 00:47:06] SGLang Kernel API Call: sglang.quant_method.UnquantizedLinearMethod.apply
================================================================================
[2026-03-19 00:47:06] SGLang Kernel API Call: sglang.custom_op.fused_inplace_qknormThis is a real level-1 excerpt captured from Qwen/Qwen3-0.6B.
export SGLANG_KERNEL_API_LOGLEVEL=3
export SGLANG_KERNEL_API_LOGDEST=debug.log
python my_script.pyOutput in debug.log:
================================================================================
[2026-03-19 00:47:30] SGLang Kernel API Call: sglang.quant_method.UnquantizedLinearMethod.apply
Positional input arguments:
arg[0]=QKVParallelLinear(
repr=QKVParallelLinear(in_features=1024, output_features=4096, bias=False, tp_size=1, gather_output=False)
)
arg[1]=Tensor(
shape=(1, 1024)
dtype=torch.bfloat16
device=cuda:0
requires_grad=False
is_contiguous=True
)
arg[2]=None
Output:
return=Tensor(
shape=(1, 4096)
dtype=torch.bfloat16
device=cuda:0
requires_grad=False
is_contiguous=True
)This is a real level-3 excerpt captured from Qwen/Qwen3-0.6B.
export SGLANG_KERNEL_API_LOGLEVEL=5
export SGLANG_KERNEL_API_LOGDEST=debug.log
python my_script.pyAdditional output:
================================================================================
[2026-03-19 01:00:42] SGLang Kernel API Call: diffusion.quant_method.UnquantizedLinearMethod.apply
Positional input arguments:
arg[1]=Tensor(
shape=(1, 77, 768)
dtype=torch.bfloat16
device=cuda:0
requires_grad=False
is_contiguous=True
min=-27.250000
max=28.500000
mean=0.011723
nan_count=0
inf_count=0
)
Output:
return=Tensor(
shape=(1, 77, 2304)
dtype=torch.bfloat16
device=cuda:0
requires_grad=False
is_contiguous=True
min=-8.937500
max=9.375000
mean=0.009460
nan_count=0
inf_count=0
)This is a real level-5 excerpt captured from black-forest-labs/FLUX.1-dev.
export SGLANG_KERNEL_API_LOGLEVEL=10
export SGLANG_KERNEL_API_LOGDEST=debug.log
export SGLANG_KERNEL_API_DUMP_DIR=/tmp/sglang_kernel_api_dumps
python my_script.pyAt level 10, SGLang saves the inputs before execution. If the kernel crashes, the dump directory still contains the inputs and exception metadata.
If CUDA graph capture is active, tensor dumps are skipped automatically to avoid capture-time CUDA errors. In that case, you still get the kernel API call log, but not inputs.pt / outputs.pt.
Level-10 dumps are best understood as crash-safe call snapshots. They always preserve the observed call boundary. They do not guarantee one-click replay for every method, because some methods depend on module state that is not serialized into the dump.
Real level-10 dump layout from Qwen/Qwen3-0.6B:
/tmp/sglang_kernel_api_validation/qwen_qwen3_0_6b_level10_dumps
/tmp/sglang_kernel_api_validation/qwen_qwen3_0_6b_level10_dumps/20260319_004821_182_pid919286_RotaryEmbedding.forward_call0001
/tmp/sglang_kernel_api_validation/qwen_qwen3_0_6b_level10_dumps/20260319_004821_182_pid919286_RotaryEmbedding.forward_call0001/inputs.pt
/tmp/sglang_kernel_api_validation/qwen_qwen3_0_6b_level10_dumps/20260319_004821_182_pid919286_RotaryEmbedding.forward_call0001/metadata.json
/tmp/sglang_kernel_api_validation/qwen_qwen3_0_6b_level10_dumps/20260319_004821_182_pid919286_RotaryEmbedding.forward_call0001/outputs.ptReal metadata.json excerpt:
{
"function_name": "RotaryEmbedding.forward",
"timestamp": "20260319_004821_182",
"process_id": 919286,
"execution_status": "completed",
"input_tensor_keys": ["arg_0", "arg_1", "arg_2"],
"output_tensor_keys": ["result_0", "result_1"]
}Create a temporary reproducer:
python3 - <<'PY'
from pathlib import Path
Path("/tmp/sglang_llm_crash.py").write_text(
"import torch\\n"
"import torch.nn.functional as F\\n"
"from sglang.srt.utils.custom_op import register_custom_op\\n\\n"
"def _fake_embedding(indices, table):\\n"
" return torch.empty((*indices.shape, table.shape[-1]), device=table.device, dtype=table.dtype)\\n\\n"
"@register_custom_op(op_name='mock_llm_cuda_crash', fake_impl=_fake_embedding)\\n"
"def mock_llm_cuda_crash(indices, table):\\n"
" out = F.embedding(indices, table)\\n"
" torch.cuda.synchronize()\\n"
" return out\\n\\n"
"table = torch.randn(4, 8, device='cuda', dtype=torch.float16)\\n"
"indices = torch.tensor([0, 7], device='cuda', dtype=torch.long)\\n"
"mock_llm_cuda_crash(indices, table)\\n"
)
PY
SGLANG_KERNEL_API_LOGLEVEL=1 \
SGLANG_KERNEL_API_LOGDEST=/tmp/sglang_llm_level1.log \
python3 /tmp/sglang_llm_crash.pyWhat to expect:
device-side assertTry the same example at level 3:
SGLANG_KERNEL_API_LOGLEVEL=3 \
SGLANG_KERNEL_API_LOGDEST=/tmp/sglang_llm_level3.log \
python3 /tmp/sglang_llm_crash.pyNow the log shows tensor metadata before the crash.
Try level 10:
SGLANG_KERNEL_API_LOGLEVEL=10 \
SGLANG_KERNEL_API_LOGDEST=/tmp/sglang_llm_level10.log \
SGLANG_KERNEL_API_DUMP_DIR=/tmp/sglang_llm_level10_dumps \
python3 /tmp/sglang_llm_crash.pyNow you should see:
sglang.custom_op.mock_llm_cuda_crashinputs.ptmetadata.json showing execution_status: "exception"outputs.pt, because the kernel crashed before producing outputFor real-model success-path level-10 dumps, it is often easier to temporarily disable CUDA graph and piecewise CUDA graph for the debug run.
Create a temporary diffusion-side reproducer:
python3 - <<'PY'
from pathlib import Path
Path("/tmp/sglang_diffusion_crash.py").write_text(
"import torch\\n"
"import torch.nn.functional as F\\n"
"from sglang.multimodal_gen.runtime.layers.utils import register_custom_op\\n\\n"
"def _fake_embedding(positions, cache):\\n"
" return torch.empty((*positions.shape, cache.shape[-1]), device=cache.device, dtype=cache.dtype)\\n\\n"
"@register_custom_op(op_name='mock_diffusion_cuda_crash', fake_impl=_fake_embedding)\\n"
"def mock_diffusion_cuda_crash(positions, cache):\\n"
" out = F.embedding(positions, cache)\\n"
" torch.cuda.synchronize()\\n"
" return out\\n\\n"
"cache = torch.randn(4, 64, device='cuda', dtype=torch.float16)\\n"
"positions = torch.tensor([0, 9], device='cuda', dtype=torch.long)\\n"
"mock_diffusion_cuda_crash(positions, cache)\\n"
)
PY
SGLANG_KERNEL_API_LOGLEVEL=1 \
SGLANG_KERNEL_API_LOGDEST=/tmp/sglang_diffusion_level1.log \
python3 /tmp/sglang_diffusion_crash.pyTry level 3:
SGLANG_KERNEL_API_LOGLEVEL=3 \
SGLANG_KERNEL_API_LOGDEST=/tmp/sglang_diffusion_level3.log \
python3 /tmp/sglang_diffusion_crash.pyTry level 10:
SGLANG_KERNEL_API_LOGLEVEL=10 \
SGLANG_KERNEL_API_LOGDEST=/tmp/sglang_diffusion_level10.log \
SGLANG_KERNEL_API_DUMP_DIR=/tmp/sglang_diffusion_level10_dumps \
python3 /tmp/sglang_diffusion_crash.pyIf your local environment has unrelated FlashInfer import issues, resolve them in the shell before running the example. The example itself does not set any FLASHINFER_* environment variable.
When running with multiple GPUs or worker processes, use %i in the log path:
export SGLANG_KERNEL_API_LOGLEVEL=3
export SGLANG_KERNEL_API_LOGDEST=debug_rank_%i.log
torchrun --nproc_per_node=4 my_script.pyThis creates separate logs such as:
debug_rank_12345.logdebug_rank_12346.logdebug_rank_12347.logdebug_rank_12348.logReal multi-process example from a 2-GPU Qwen/Qwen2.5-0.5B-Instruct run:
/tmp/sglang_kernel_api_validation_multi/qwen_qwen2_5_0_5b_instruct_level3_950201.log
/tmp/sglang_kernel_api_validation_multi/qwen_qwen2_5_0_5b_instruct_level3_950349.log
/tmp/sglang_kernel_api_validation_multi/qwen_qwen2_5_0_5b_instruct_level3_950350.log
/tmp/sglang_kernel_api_validation_multi/qwen_qwen2_5_0_5b_instruct_level3_950351.logYou should usually do the same for level-10 dump directories:
export SGLANG_KERNEL_API_LOGLEVEL=10
export SGLANG_KERNEL_API_LOGDEST=debug_rank_%i.log
export SGLANG_KERNEL_API_DUMP_DIR=/tmp/sglang_kernel_api_dumps_%iThis avoids multiple ranks writing into the same dump directory tree.
If level 10 is too noisy, restrict dumps to specific APIs:
export SGLANG_KERNEL_API_LOGLEVEL=10
export SGLANG_KERNEL_API_LOGDEST=debug.log
export SGLANG_KERNEL_API_DUMP_DIR=/tmp/sglang_kernel_api_dumps
export SGLANG_KERNEL_API_DUMP_INCLUDE='sglang.custom_op.*'
export SGLANG_KERNEL_API_DUMP_EXCLUDE='*.fake_impl'SGLANG_KERNEL_API_DUMP_INCLUDE and SGLANG_KERNEL_API_DUMP_EXCLUDE use shell-style wildcard matching.
Typical errors:
RuntimeError: CUDA error: an illegal memory access was encountered
torch.AcceleratorError: CUDA error: device-side assert triggeredUse:
export SGLANG_KERNEL_API_LOGLEVEL=3Check in the logs:
Typical shape-mismatch pattern:
SGLang Kernel API Call: ...
arg[0]=Tensor(shape=(..., 128), ...) # ✅ expected dimension
arg[1]=Tensor(shape=(..., 64), ...) # ❌ mismatchThis often points to head-dim, hidden-dim, or cache-layout mismatch rather than a random CUDA failure.
Use:
export SGLANG_KERNEL_API_LOGLEVEL=5Check:
minmaxmeannan_countinf_countTypical bad pattern:
Tensor(
...
min=-1234567.000000 # ❌ suspiciously large
max=9876543.000000 # ❌ suspiciously large
mean=nan # ❌ bad
nan_count=128 # ❌ found NaNs
inf_count=0 # ✅ no Infs here
)This usually means the bad values were already present before the crashing kernel.
Use:
export SGLANG_KERNEL_API_LOGLEVEL=3Check:
Also check whether a supposedly per-token or per-frame tensor accidentally became full-sequence or full-image sized.
Typical bad pattern:
Tensor(
shape=(1024, 8192, 128, 128) # ❌ way too large
...
)Suppose the failing API log looks like this:
[2026-03-19 00:47:30] SGLang Kernel API Call: RotaryEmbedding.forward
Positional input arguments:
arg[0]=Tensor(shape=(1, 8), dtype=torch.int64, ...)
arg[1]=Tensor(shape=(1, 8, 8, 256), dtype=torch.bfloat16, ...) # ✅ query
arg[2]=Tensor(shape=(1, 8, 4, 64), dtype=torch.bfloat16, ...) # ❌ key head_dim mismatchWhat this tells you:
That usually means the bug is in projection layout, head packing, or cache format rather than in the rotary kernel itself.
For harder bugs, combine kernel API logging with CUDA memory checking:
export SGLANG_KERNEL_API_LOGLEVEL=3
export SGLANG_KERNEL_API_LOGDEST=debug.log
compute-sanitizer --tool memcheck python3 /tmp/sglang_llm_crash.pyUse debug.log to see the exact inputs that reached the crashing API boundary.
Typical compute-sanitizer output:
========= COMPUTE-SANITIZER
========= Invalid __global__ write of size 4 bytes
========= at 0x1234 in SomeKernel
========= by thread (256,0,0) in block (10,0,0)
========= Address 0x... is out of boundsUse the sanitizer output to identify the failing kernel and use debug.log to identify the exact tensors that reached the API boundary right before it.
If you need more synchronous host-side error reporting, you can try CUDA_LAUNCH_BLOCKING=1 as a separate follow-up experiment. It is not part of the default workflow because it changes execution timing and can hide concurrency-related behavior.
For crashes that need a stack trace instead of only memory diagnostics:
export SGLANG_KERNEL_API_LOGLEVEL=3
export SGLANG_KERNEL_API_LOGDEST=debug.log
cuda-gdb --args python3 /tmp/sglang_llm_crash.pyInside cuda-gdb:
(cuda-gdb) run
(cuda-gdb) whereThen correlate the backtrace with debug.log.
When you own the CUDA kernel, printf() is still useful for narrowing down bad indices, bad launch geometry, or broken state propagation.
Basic pattern:
__global__ void MyKernel(const float* input, float* output, int n) {
int idx = blockIdx.x * blockDim.x + threadIdx.x;
if (threadIdx.x == 0 && blockIdx.x == 0) {
printf("n=%d input0=%f\n", n, input[0]);
}
if (idx < n) {
output[idx] = input[idx] * 2.0f;
}
}After launch, force the output to flush:
my_kernel(...)
torch.cuda.synchronize()For warp-specialized kernels, do not blindly print only on threadIdx.x == 0. Pick one representative thread per warp or per specialization group instead.
Problem:
threadIdx.x == 0 only prints from the first warp in the blockBetter pattern:
__global__ void WarpSpecializedKernel(...) {
// Example: first lane of each warp
if ((threadIdx.x % 32) == 0) {
printf("warp=%d\n", threadIdx.x / 32);
}
}Or, if the kernel is organized in larger specialization groups, print once per group instead of once per block.
Common mistake:
// Only warp 0 prints
if (threadIdx.x == 0) {
printf("warp=%d\n", threadIdx.x / 32);
}| Kernel Type | Print Condition | Notes |
|---|---|---|
| Simple kernel | threadIdx.x == 0 | One thread per block is usually enough |
| Warp-specialized kernel | one representative lane per warp | e.g. threadIdx.x % 32 == 0 |
| Group-specialized kernel | one representative lane per group | choose based on the kernel's scheduling layout |
assert(value >= 0.0f && "value must be non-negative");
static_assert(BLOCK_SIZE % 32 == 0, "BLOCK_SIZE must be warp aligned");| Variable | Values | Description |
|---|---|---|
SGLANG_KERNEL_API_LOGLEVEL | 0 | No logging (default) |
1 | Function names only | |
3 | Inputs and outputs with metadata | |
5 | Level 3 plus tensor statistics | |
10 | Level 5 plus crash-safe tensor dumps | |
SGLANG_KERNEL_API_LOGDEST | stdout | Log to stdout |
stderr | Log to stderr | |
<path> | Log to file | |
log_%i.txt | %i expands to process ID | |
SGLANG_KERNEL_API_DUMP_DIR | <path> | Directory for level-10 dumps |
SGLANG_KERNEL_API_DUMP_INCLUDE | wildcard list | Only dump matching API names |
SGLANG_KERNEL_API_DUMP_EXCLUDE | wildcard list | Skip matching API names |
export SGLANG_KERNEL_API_LOGLEVEL=3Level 3 is usually enough to catch wrong shapes, wrong dtypes, and wrong devices.
export SGLANG_KERNEL_API_LOGLEVEL=5Use it when you suspect NaN or Inf values.
export SGLANG_KERNEL_API_LOGLEVEL=10This is the most useful mode when the process crashes before you can inspect live tensors.
If you need successful input/output dumps from a real model run, temporarily disable CUDA graph for that debug session.
When level 10 is too noisy, pair it with SGLANG_KERNEL_API_DUMP_INCLUDE / SGLANG_KERNEL_API_DUMP_EXCLUDE instead of dumping every covered API.
export SGLANG_KERNEL_API_LOGDEST=crash.logFile logs are safer than stdout when the process aborts.
unset SGLANG_KERNEL_API_LOGLEVELWhen disabled, the decorator returns the original callable and adds no runtime logging overhead.
Check:
echo $SGLANG_KERNEL_API_LOGLEVELecho $SGLANG_KERNEL_API_LOGDESTReduce the level:
export SGLANG_KERNEL_API_LOGLEVEL=3If you see:
statistics=[skipped: CUDA graph capture in progress]That is expected. Level-5 statistics are intentionally skipped during CUDA graph capture to avoid synchronization side effects.
If you see:
Tensor dump skipped: CUDA graph capture in progressThat is also expected. Level-10 dumps require copying tensors to CPU, which is not allowed during CUDA graph capture.
© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/debug-cuda-crash of sgl-project/sglang.
Open the folder on GitHubat commit b7b2975
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
Debug Cuda Crash next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Debug Cuda Crash this skillsgl-project/sglang | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Cpp Kernelvipshop/cache-dit | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Aoti Debugpytorch/pytorch | 104k | 1 repos | ~1.7k | Automated safety check: Pass | Custom licence | |
| The Art of Debuggingstas00/the-art-of-debugging | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Veomni DebugByteDance-Seed/VeOmni | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 |
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
ByteDance-Seed/VeOmni
A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
Categories
Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging. Debug Cuda Crash is an agent skill from sgl-project/sglang.
Debug Cuda Crash fits situations like: tasks that involve Deep learning; tasks that involve Debugging.
Run `npx skills add sgl-project/sglang --skill debug-cuda-crash -a claude-code`. Or copy the skill folder (.agents/skills/debug-cuda-crash in sgl-project/sglang) into .claude/skills/debug-cuda-crash in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill debug-cuda-crash -a codex`. Or copy the skill folder (.agents/skills/debug-cuda-crash in sgl-project/sglang) into .agents/skills/debug-cuda-crash in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill debug-cuda-crash -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-cuda-crash, .gemini/skills/debug-cuda-crash, .github/skills/debug-cuda-crash and .opencode/skills/debug-cuda-crash in your project.
Going by SKILL.md and its folder, Debug Cuda Crash needs the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Debug Cuda Crash is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Debug Cuda Crash: Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars), Aoti Debug (pytorch/pytorch, 104k stars), The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,851 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 8, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.