Add Jit Kernel
guqiong96/Lsglang
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module
Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)
$ npx skills add sgl-project/sglang --skill add-sgl-kernel -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang add-sgl-kernel --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-sgl-kernel .claude/skills/add-sgl-kernel && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-sgl-kernel" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/add-sgl-kernel into .claude/skills/add-sgl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-sgl-kernel", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/.agents/skills/add-sgl-kernelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill add-sgl-kernel -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang add-sgl-kernel --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/add-sgl-kernel .agents/skills/add-sgl-kernel && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-sgl-kernel" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/add-sgl-kernel into .agents/skills/add-sgl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-sgl-kernel", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill add-sgl-kernel -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang add-sgl-kernel --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/add-sgl-kernel .cursor/skills/add-sgl-kernel && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-sgl-kernel" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/add-sgl-kernel into .cursor/skills/add-sgl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-sgl-kernel", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path .agents/skills/add-sgl-kernel--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill add-sgl-kernel -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang add-sgl-kernel --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/add-sgl-kernel .gemini/skills/add-sgl-kernel && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-sgl-kernel" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/add-sgl-kernel into .gemini/skills/add-sgl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-sgl-kernel", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang add-sgl-kernelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill add-sgl-kernel -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/add-sgl-kernel .github/skills/add-sgl-kernel && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-sgl-kernel" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/add-sgl-kernel into .github/skills/add-sgl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-sgl-kernel", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill add-sgl-kernel -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang add-sgl-kernel --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/add-sgl-kernel .opencode/skills/add-sgl-kernel && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-sgl-kernel" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/add-sgl-kernel into .opencode/skills/add-sgl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-sgl-kernel", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-sgl-kernelStep-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)
Add Sgl Kernel is an agent skill from sgl-project/sglang. Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. It works with SGLang, C++, CUDA and Python. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1c42ad3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
makepytestpythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add Sgl Kernel loads about 3.4k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 726 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sgl-project/sglang at commit 1c42ad3, republished under its Apache-2.0 licence (© sgl-project). 726 words, ~3,377 tokens.
.claude/skills/add-sgl-kernel/SKILL.md (or your agent's skills folder).sgl-kernel (AOT / Heavyweight)Apply kernel-organization for the public operator namespace, logical grouping, lazy registry metadata, and test placement. The implementation tutorial below does not replace that API contract.
This tutorial walks through adding a simple element-wise scale operation as an AOT kernel. We'll implement scale(x, factor) = x * factor to demonstrate the complete workflow.
Add a new operation that scales each element of a tensor by a scalar factor:
x (CUDA) and scalar factor (float)x * factor (element-wise, in-place or into pre-allocated out)torch.float16), BF16 (torch.bfloat16), FP32 (torch.float32)DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16 macro (defined in python/sglang/kernels/aot/include/utils.h)python/sglang/kernels/jit first when the kernel does not depend on CUTLASS or another large C++ project. This is the default path for lightweight kernels that benefit from rapid iteration.sgl-kernel when the kernel does depend on CUTLASS or another large C++ project, or when it should be part of the AOT wheel / torch op registration flow.flashinfer, or CUTLASS that is already provided through flashinfer, the kernel can still be implemented as jit_kernel.In addition, every new kernel must ship with:
marker.do_bench from sglang.kernels.jit.benchmark — it works for any callable, not only JIT kernels)You will typically touch these files/areas:
python/sglang/kernels/aot/csrc/elementwise/scale.cu (pick the right subdirectory)python/sglang/kernels/aot/include/sgl_kernel_ops.hpython/sglang/kernels/aot/csrc/common_extension.ccpython/sglang/kernels/aot/CMakeLists.txt (set(SOURCES ...))python/sglang/kernels/aot/python/sgl_kernel/ and python/sglang/kernels/aot/python/sgl_kernel/__init__.pypython/sglang/kernels/aot/tests/test_scale.pypython/sglang/kernels/aot/benchmark/bench_scale.pycsrc/Pick the right subdirectory:
csrc/elementwise/ — for element-wise ops (our example)csrc/gemm/, csrc/attention/, csrc/moe/ — for other categoriesCreate python/sglang/kernels/aot/csrc/elementwise/scale.cu:
#include <ATen/cuda/CUDAContext.h>
#include <c10/cuda/CUDAGuard.h>
#include <torch/all.h>
#include "utils.h" // DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16
// scale_kernel: out[i] = input[i] * factor
// Supports float, half (__half), __nv_bfloat16 via template T
template <typename T>
__global__ void scale_kernel(T* __restrict__ out,
const T* __restrict__ input,
float factor,
int64_t n) {
int64_t idx = static_cast<int64_t>(blockIdx.x) * blockDim.x + threadIdx.x;
if (idx < n) {
out[idx] = static_cast<T>(static_cast<float>(input[idx]) * factor);
}
}
void scale(at::Tensor& out, const at::Tensor& input, double factor) {
TORCH_CHECK(input.is_cuda(), "input must be a CUDA tensor");
TORCH_CHECK(input.is_contiguous(), "input must be contiguous");
TORCH_CHECK(out.is_cuda(), "out must be a CUDA tensor");
TORCH_CHECK(out.is_contiguous(), "out must be contiguous");
TORCH_CHECK(out.sizes() == input.sizes(), "out and input must have the same shape");
TORCH_CHECK(out.scalar_type() == input.scalar_type(),
"out and input must have the same dtype");
const int64_t n = input.numel();
const int threads = 256;
const int blocks = (n + threads - 1) / threads;
const cudaStream_t stream = at::cuda::getCurrentCUDAStream();
const at::cuda::OptionalCUDAGuard device_guard(device_of(input));
// Dispatches over float, float16, bfloat16
DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16(input.scalar_type(), c_type, [&] {
scale_kernel<c_type><<<blocks, threads, 0, stream>>>(
static_cast<c_type*>(out.data_ptr()),
static_cast<const c_type*>(input.data_ptr()),
static_cast<float>(factor),
n);
cudaError_t status = cudaGetLastError();
TORCH_CHECK(status == cudaSuccess,
"scale_kernel launch failed: ", cudaGetErrorString(status));
return true;
});
}Key points:
at::Tensor (PyTorch tensors), TORCH_CHECK for validation, at::cuda::getCurrentCUDAStream() for streamDISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16 covers float, half (FP16), __nv_bfloat16 (BF16)TORCH_CHECK and skip logic in testsinclude/sgl_kernel_ops.hEdit python/sglang/kernels/aot/include/sgl_kernel_ops.h, add to the elementwise section:
void scale(at::Tensor& out, const at::Tensor& input, double factor);csrc/common_extension.ccEdit python/sglang/kernels/aot/csrc/common_extension.cc, inside TORCH_LIBRARY_FRAGMENT(sgl_kernel, m):
// From csrc/elementwise
m.def("scale(Tensor! out, Tensor input, float factor) -> ()");
m.impl("scale", torch::kCUDA, &scale);Key points:
Tensor! means in-place / mutable output argumenttorch.compile and for consistent call signaturesfloat here), but note that the C++ launcher signature still needs double for scalar arguments accepted by torch::LibraryCMakeLists.txtEdit python/sglang/kernels/aot/CMakeLists.txt, add to set(SOURCES ...):
csrc/elementwise/scale.cuKey points:
python/sglang/kernels/aot/python/sgl_kernel/Prefer following the existing module organization first. For elementwise kernels, the usual pattern is:
python/sglang/kernels/aot/python/sgl_kernel/elementwise.pypython/sglang/kernels/aot/python/sgl_kernel/__init__.pyFor example, in python/sglang/kernels/aot/python/sgl_kernel/elementwise.py, add:
import torch
def scale(
input: torch.Tensor,
factor: float,
out: torch.Tensor | None = None,
) -> torch.Tensor:
"""
Element-wise scale: out = input * factor.
Supported dtypes: torch.float16, torch.bfloat16, torch.float32.
Parameters
----------
input : CUDA input tensor
factor : scale factor (float)
out : optional pre-allocated CUDA output tensor (same shape/dtype as input)
"""
if out is None:
out = torch.empty_like(input)
torch.ops.sgl_kernel.scale.default(out, input, factor)
return outThen re-export it from python/sglang/kernels/aot/python/sgl_kernel/__init__.py following the existing import style used by other kernels.
After exposing the AOT wheel symbol, add a lazy wrapper and a KernelSpec under
python/sglang/kernels/ops/<group>/. Set backend=KernelBackend.AOT, record
actual device/architecture support, and keep sgl_kernel imports inside the
implementation path. SGLang runtime and integration tests import that wrapper.
For this example, use sglang.kernels.ops.elementwise.scale.
The wheel-level tests below validate its standalone API/build. They do not
replace CI-registered SGLang correctness tests in
test/registered/kernels/ops/elementwise/test_scale.py and benchmarks in
test/registered/kernels/benchmark/elementwise/bench_scale.py. Follow
write-sglang-test for their registration and CI budget; do not register tests
under the shipped python/sglang/ package.
Create python/sglang/kernels/aot/tests/test_scale.py:
import pytest
import torch
import sgl_kernel
@pytest.mark.parametrize("dtype", [torch.float16, torch.bfloat16, torch.float32])
@pytest.mark.parametrize("size", [128, 1024, 4096, 65536])
@pytest.mark.parametrize("factor", [0.5, 1.0, 2.0])
def test_scale_correctness(dtype, size, factor):
input = torch.randn(size, dtype=dtype, device="cuda")
out = torch.empty_like(input)
result = sgl_kernel.scale(input, factor, out=out)
assert result is out
expected = input * factor
rtol, atol = (1e-5, 1e-6) if dtype == torch.float32 else (1e-2, 1e-2)
torch.testing.assert_close(out, expected, rtol=rtol, atol=atol)
def test_scale_shape_mismatch():
input = torch.randn(128, dtype=torch.float16, device="cuda")
out = torch.empty(256, dtype=torch.float16, device="cuda")
with pytest.raises(RuntimeError, match="same shape"):
sgl_kernel.scale(input, 2.0, out=out)
def test_scale_cpu_input():
input = torch.randn(128, dtype=torch.float16) # CPU
out = torch.empty_like(input)
with pytest.raises(RuntimeError, match="CUDA"):
sgl_kernel.scale(input, 2.0, out=out)
if __name__ == "__main__":
import sys
sys.exit(pytest.main([__file__, "-q"]))Every benchmark must account for L2 cache reuse — see rules/kernel-benchmark.md.
Create python/sglang/kernels/aot/benchmark/bench_scale.py:
import torch
import sgl_kernel
from sglang.kernels.jit.benchmark import marker
def sglang_scale(input: torch.Tensor, factor: float, out: torch.Tensor) -> None:
sgl_kernel.scale(input, factor, out=out)
def torch_scale(input: torch.Tensor, factor: float, out: torch.Tensor) -> None:
torch.mul(input, factor, out=out)
FN_MAP = {"sglang": sglang_scale, "torch": torch_scale}
@marker.parametrize("dtype", [torch.float16, torch.bfloat16, torch.float32], [torch.float16])
@marker.parametrize("size", [2**n for n in range(10, 20)], [4096]) # 1K .. 512K
@marker.benchmark("provider", ["sglang", "torch"])
def benchmark(dtype: torch.dtype, size: int, provider: str):
input = torch.randn(size, dtype=dtype, device="cuda")
out = torch.empty_like(input)
return marker.do_bench(
FN_MAP[provider],
# Pass every tensor through input_args (not a closure) so marker rotates
# them across CUDA-graph calls to defeat L2 reuse.
input_args=(input, 2.0, out),
# Bandwidth = bytes(input) + bytes(out), both already in input_args.
memory_output=None,
)
if __name__ == "__main__":
benchmark.run()Build:
cd python/sglang/kernels/aot
make build -j16If you need to limit host resource usage:
cd python/sglang/kernels/aot
make build -j1 MAX_JOBS=2 CMAKE_ARGS="-DSGL_KERNEL_COMPILE_THREADS=1"After building successfully, run the test and benchmark:
pytest python/sglang/kernels/aot/tests/test_scale.py -q
python python/sglang/kernels/aot/benchmark/bench_scale.pyPR CI also runs pr-test-sgl-kernel.yml, including the B200 job
sgl-kernel-b200-test when kernel changes are detected. Use that job as the
Blackwell coverage signal for AOT sgl-kernel changes.
CUDA_LAUNCH_BLOCKING=1compute-sanitizer --tool memcheck python ...MAX_JOBS and SGL_KERNEL_COMPILE_THREADSpython/sglang/kernels/aot/analyze_whl_kernel_sizes.py.cu file is missing from SOURCES, the symbol will be undefined at link timepython/sglang/kernels/aot/README.mdpython/sglang/kernels/aot/include/sgl_kernel_ops.hpython/sglang/kernels/aot/csrc/common_extension.ccpython/sglang/kernels/aot/CMakeLists.txtpython/sglang/kernels/aot/include/utils.h — DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16 macro and friendspython/sglang/kernels/aot/csrc/elementwise/activation.cu — reference for the FP16/BF16/FP32 dispatch patternpython/sglang/kernels/aot/csrc/elementwise/scale.cu # NEW: CUDA kernel + launcher
python/sglang/kernels/aot/include/sgl_kernel_ops.h # MODIFIED: C++ declaration
python/sglang/kernels/aot/csrc/common_extension.cc # MODIFIED: schema + dispatch registration
python/sglang/kernels/aot/CMakeLists.txt # MODIFIED: add source file (alphabetical)
python/sglang/kernels/aot/python/sgl_kernel/elementwise.py # MODIFIED: Python wrapper
python/sglang/kernels/aot/python/sgl_kernel/__init__.py # MODIFIED: re-export Python API
python/sglang/kernels/aot/tests/test_scale.py # NEW: tests
python/sglang/kernels/aot/benchmark/bench_scale.py # NEW: benchmark© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/add-sgl-kernel of sgl-project/sglang.
Open the folder on GitHubat commit 1c42ad3
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
Add Sgl Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add Sgl Kernel this skillsgl-project/sglang | 37k | 2 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| Add Jit Kernelguqiong96/Lsglang | 143 | 1 repos | ~10k | Automated safety check: Pass | Apache-2.0 | |
| Paddle BuildPaddlePaddle/Paddle | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Ako4allTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT | |
| Paddle Op DevPaddlePaddle/Paddle | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT |
guqiong96/Lsglang
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
vipshop/cache-dit
A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
Categories
Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks). Add Sgl Kernel is an agent skill from sgl-project/sglang.
Add Sgl Kernel fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add sgl-project/sglang --skill add-sgl-kernel -a claude-code`. Or copy the skill folder (.agents/skills/add-sgl-kernel in sgl-project/sglang) into .claude/skills/add-sgl-kernel in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill add-sgl-kernel -a codex`. Or copy the skill folder (.agents/skills/add-sgl-kernel in sgl-project/sglang) into .agents/skills/add-sgl-kernel in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill add-sgl-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-sgl-kernel, .gemini/skills/add-sgl-kernel, .github/skills/add-sgl-kernel and .opencode/skills/add-sgl-kernel in your project.
Going by SKILL.md and its folder, Add Sgl Kernel needs the command-line tools its instructions call (make, pytest and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Add Sgl Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add Sgl Kernel: Add Jit Kernel (guqiong96/Lsglang, 143 stars), Paddle Build (PaddlePaddle/Paddle, 24k stars), Ako4all (TongmingLAIC/AKO4ALL, 369 stars) and Paddle Op Dev (PaddlePaddle/Paddle, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,829 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.