Agent skill

Add Sgl Kernel

by sgl-project in sgl-project/sglang

Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add Sgl Kernel

skills CLI
$ npx skills add sgl-project/sglang --skill add-sgl-kernel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang add-sgl-kernel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-sgl-kernel .claude/skills/add-sgl-kernel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-sgl-kernel
GitHub stars
37k
Used in
2 other repos
Token cost
~3.4k tokens
SKILL.md length
726 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

  • Works in 9 steps: Implement the kernel in csrc/ → Add a C++ declaration in… → Register the op in… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Goal, Two rules of thumb (must follow), Repository integration map and Step 1: Implement the kernel…, plus 12 more sections
  • Calls make, pytest and python

What it does

Add Sgl Kernel is an agent skill from sgl-project/sglang. Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with SGLang, C++, CUDA and Python. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/add-sgl-kernel”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Implement the kernel in csrc/
  2. Add a C++ declaration in include/sgl_kernel_ops.h
  3. Register the op in csrc/common_extension.cc
  4. Add the new source file to CMakeLists.txt
  5. Expose a Python API under python/sglang/kernels/aot/python/sgl_kernel/
  6. Write tests (required)
  7. Add a benchmark (required)
  8. Build
  9. Validate

What it can do on your machine

Read from SKILL.md and the folder at commit 1c42ad3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make
    • pytest
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Sgl Kernel loads about 3.4k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 726 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit 1c42ad3, republished under its Apache-2.0 licence (© sgl-project). 726 words, ~3,377 tokens.

Download SKILL.mdSave it as .claude/skills/add-sgl-kernel/SKILL.md (or your agent's skills folder).
name
add-sgl-kernel
description
Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

Tutorial: Adding a New Kernel to sgl-kernel (AOT / Heavyweight)

Apply kernel-organization for the public operator namespace, logical grouping, lazy registry metadata, and test placement. The implementation tutorial below does not replace that API contract.

This tutorial walks through adding a simple element-wise scale operation as an AOT kernel. We'll implement scale(x, factor) = x * factor to demonstrate the complete workflow.

Goal

Add a new operation that scales each element of a tensor by a scalar factor:

  • Input: tensor x (CUDA) and scalar factor (float)
  • Output: x * factor (element-wise, in-place or into pre-allocated out)
  • Supported dtypes: FP16 (torch.float16), BF16 (torch.bfloat16), FP32 (torch.float32)
    • Dispatched via DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16 macro (defined in python/sglang/kernels/aot/include/utils.h)

Two rules of thumb (must follow)

  1. Prefer python/sglang/kernels/jit first when the kernel does not depend on CUTLASS or another large C++ project. This is the default path for lightweight kernels that benefit from rapid iteration.
  2. Prefer sgl-kernel when the kernel does depend on CUTLASS or another large C++ project, or when it should be part of the AOT wheel / torch op registration flow.
  3. Exception: if the dependency is flashinfer, or CUTLASS that is already provided through flashinfer, the kernel can still be implemented as jit_kernel.

In addition, every new kernel must ship with:

  • Tests (pytest)
  • A benchmark script (marker.do_bench from sglang.kernels.jit.benchmark — it works for any callable, not only JIT kernels)

Repository integration map

You will typically touch these files/areas:

  • Implementation: python/sglang/kernels/aot/csrc/elementwise/scale.cu (pick the right subdirectory)
  • Public declarations: python/sglang/kernels/aot/include/sgl_kernel_ops.h
  • Torch extension registration: python/sglang/kernels/aot/csrc/common_extension.cc
  • Build: python/sglang/kernels/aot/CMakeLists.txt (set(SOURCES ...))
  • Python API: python/sglang/kernels/aot/python/sgl_kernel/ and python/sglang/kernels/aot/python/sgl_kernel/__init__.py
  • Tests: python/sglang/kernels/aot/tests/test_scale.py
  • Benchmarks: python/sglang/kernels/aot/benchmark/bench_scale.py

Step 1: Implement the kernel in csrc/

Pick the right subdirectory:

  • csrc/elementwise/ — for element-wise ops (our example)
  • csrc/gemm/, csrc/attention/, csrc/moe/ — for other categories

Create python/sglang/kernels/aot/csrc/elementwise/scale.cu:

cpp
#include <ATen/cuda/CUDAContext.h>
#include <c10/cuda/CUDAGuard.h>
#include <torch/all.h>

#include "utils.h"  // DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16

// scale_kernel: out[i] = input[i] * factor
// Supports float, half (__half), __nv_bfloat16 via template T
template <typename T>
__global__ void scale_kernel(T* __restrict__ out,
                              const T* __restrict__ input,
                              float factor,
                              int64_t n) {
  int64_t idx = static_cast<int64_t>(blockIdx.x) * blockDim.x + threadIdx.x;
  if (idx < n) {
    out[idx] = static_cast<T>(static_cast<float>(input[idx]) * factor);
  }
}

void scale(at::Tensor& out, const at::Tensor& input, double factor) {
  TORCH_CHECK(input.is_cuda(),       "input must be a CUDA tensor");
  TORCH_CHECK(input.is_contiguous(), "input must be contiguous");
  TORCH_CHECK(out.is_cuda(),         "out must be a CUDA tensor");
  TORCH_CHECK(out.is_contiguous(),   "out must be contiguous");
  TORCH_CHECK(out.sizes() == input.sizes(),  "out and input must have the same shape");
  TORCH_CHECK(out.scalar_type() == input.scalar_type(),
              "out and input must have the same dtype");

  const int64_t n = input.numel();
  const int threads = 256;
  const int blocks  = (n + threads - 1) / threads;

  const cudaStream_t stream = at::cuda::getCurrentCUDAStream();
  const at::cuda::OptionalCUDAGuard device_guard(device_of(input));

  // Dispatches over float, float16, bfloat16
  DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16(input.scalar_type(), c_type, [&] {
    scale_kernel<c_type><<<blocks, threads, 0, stream>>>(
        static_cast<c_type*>(out.data_ptr()),
        static_cast<const c_type*>(input.data_ptr()),
        static_cast<float>(factor),
        n);
    cudaError_t status = cudaGetLastError();
    TORCH_CHECK(status == cudaSuccess,
                "scale_kernel launch failed: ", cudaGetErrorString(status));
    return true;
  });
}

Key points:

  • Use at::Tensor (PyTorch tensors), TORCH_CHECK for validation, at::cuda::getCurrentCUDAStream() for stream
  • Keep Python wrappers thin; do shape/dtype/device validation in C++ right around the launch path
  • DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16 covers float, half (FP16), __nv_bfloat16 (BF16)
  • Add device error checking after every kernel launch
  • If a kernel only works on certain architectures, enforce that with TORCH_CHECK and skip logic in tests

Step 2: Add a C++ declaration in include/sgl_kernel_ops.h

Edit python/sglang/kernels/aot/include/sgl_kernel_ops.h, add to the elementwise section:

cpp
void scale(at::Tensor& out, const at::Tensor& input, double factor);

Step 3: Register the op in csrc/common_extension.cc

Edit python/sglang/kernels/aot/csrc/common_extension.cc, inside TORCH_LIBRARY_FRAGMENT(sgl_kernel, m):

cpp
// From csrc/elementwise
m.def("scale(Tensor! out, Tensor input, float factor) -> ()");
m.impl("scale", torch::kCUDA, &scale);

Key points:

  • Tensor! means in-place / mutable output argument
  • The schema is important for torch.compile and for consistent call signatures
  • Keep the torch schema in PyTorch scalar types (float here), but note that the C++ launcher signature still needs double for scalar arguments accepted by torch::Library

Step 4: Add the new source file to CMakeLists.txt

Edit python/sglang/kernels/aot/CMakeLists.txt, add to set(SOURCES ...):

cmake
csrc/elementwise/scale.cu

Key points:

  • Keep the list alphabetically sorted (the file explicitly requires this)
  • If the kernel has arch constraints, reflect that in tests/benchmarks via skip logic

Show full SKILL.md (278 more words)Show less

Step 5: Expose a Python API under python/sglang/kernels/aot/python/sgl_kernel/

Prefer following the existing module organization first. For elementwise kernels, the usual pattern is:

  • implement the Python wrapper in python/sglang/kernels/aot/python/sgl_kernel/elementwise.py
  • then re-export it from python/sglang/kernels/aot/python/sgl_kernel/__init__.py

For example, in python/sglang/kernels/aot/python/sgl_kernel/elementwise.py, add:

python
import torch

def scale(
    input: torch.Tensor,
    factor: float,
    out: torch.Tensor | None = None,
) -> torch.Tensor:
    """
    Element-wise scale: out = input * factor.

    Supported dtypes: torch.float16, torch.bfloat16, torch.float32.

    Parameters
    ----------
    input  : CUDA input tensor
    factor : scale factor (float)
    out    : optional pre-allocated CUDA output tensor (same shape/dtype as input)
    """
    if out is None:
        out = torch.empty_like(input)
    torch.ops.sgl_kernel.scale.default(out, input, factor)
    return out

Then re-export it from python/sglang/kernels/aot/python/sgl_kernel/__init__.py following the existing import style used by other kernels.


SGLang integration (required for runtime use)

After exposing the AOT wheel symbol, add a lazy wrapper and a KernelSpec under python/sglang/kernels/ops/<group>/. Set backend=KernelBackend.AOT, record actual device/architecture support, and keep sgl_kernel imports inside the implementation path. SGLang runtime and integration tests import that wrapper. For this example, use sglang.kernels.ops.elementwise.scale.

The wheel-level tests below validate its standalone API/build. They do not replace CI-registered SGLang correctness tests in test/registered/kernels/ops/elementwise/test_scale.py and benchmarks in test/registered/kernels/benchmark/elementwise/bench_scale.py. Follow write-sglang-test for their registration and CI budget; do not register tests under the shipped python/sglang/ package.

Step 6: Write tests (required)

Create python/sglang/kernels/aot/tests/test_scale.py:

python
import pytest

import torch
import sgl_kernel

@pytest.mark.parametrize("dtype", [torch.float16, torch.bfloat16, torch.float32])
@pytest.mark.parametrize("size", [128, 1024, 4096, 65536])
@pytest.mark.parametrize("factor", [0.5, 1.0, 2.0])
def test_scale_correctness(dtype, size, factor):
    input = torch.randn(size, dtype=dtype, device="cuda")
    out   = torch.empty_like(input)

    result = sgl_kernel.scale(input, factor, out=out)
    assert result is out

    expected = input * factor
    rtol, atol = (1e-5, 1e-6) if dtype == torch.float32 else (1e-2, 1e-2)
    torch.testing.assert_close(out, expected, rtol=rtol, atol=atol)


def test_scale_shape_mismatch():
    input = torch.randn(128, dtype=torch.float16, device="cuda")
    out   = torch.empty(256, dtype=torch.float16, device="cuda")
    with pytest.raises(RuntimeError, match="same shape"):
        sgl_kernel.scale(input, 2.0, out=out)


def test_scale_cpu_input():
    input = torch.randn(128, dtype=torch.float16)  # CPU
    out   = torch.empty_like(input)
    with pytest.raises(RuntimeError, match="CUDA"):
        sgl_kernel.scale(input, 2.0, out=out)


if __name__ == "__main__":
    import sys
    sys.exit(pytest.main([__file__, "-q"]))

Step 7: Add a benchmark (required)

Every benchmark must account for L2 cache reuse — see rules/kernel-benchmark.md.

Create python/sglang/kernels/aot/benchmark/bench_scale.py:

python
import torch

import sgl_kernel
from sglang.kernels.jit.benchmark import marker


def sglang_scale(input: torch.Tensor, factor: float, out: torch.Tensor) -> None:
    sgl_kernel.scale(input, factor, out=out)


def torch_scale(input: torch.Tensor, factor: float, out: torch.Tensor) -> None:
    torch.mul(input, factor, out=out)


FN_MAP = {"sglang": sglang_scale, "torch": torch_scale}


@marker.parametrize("dtype", [torch.float16, torch.bfloat16, torch.float32], [torch.float16])
@marker.parametrize("size", [2**n for n in range(10, 20)], [4096])  # 1K .. 512K
@marker.benchmark("provider", ["sglang", "torch"])
def benchmark(dtype: torch.dtype, size: int, provider: str):
    input = torch.randn(size, dtype=dtype, device="cuda")
    out = torch.empty_like(input)
    return marker.do_bench(
        FN_MAP[provider],
        # Pass every tensor through input_args (not a closure) so marker rotates
        # them across CUDA-graph calls to defeat L2 reuse.
        input_args=(input, 2.0, out),
        # Bandwidth = bytes(input) + bytes(out), both already in input_args.
        memory_output=None,
    )


if __name__ == "__main__":
    benchmark.run()

Step 8: Build

Build:

bash
cd python/sglang/kernels/aot
make build -j16

If you need to limit host resource usage:

bash
cd python/sglang/kernels/aot
make build -j1 MAX_JOBS=2 CMAKE_ARGS="-DSGL_KERNEL_COMPILE_THREADS=1"

Step 9: Validate

After building successfully, run the test and benchmark:

bash
pytest python/sglang/kernels/aot/tests/test_scale.py -q
python python/sglang/kernels/aot/benchmark/bench_scale.py

PR CI also runs pr-test-sgl-kernel.yml, including the B200 job sgl-kernel-b200-test when kernel changes are detected. Use that job as the Blackwell coverage signal for AOT sgl-kernel changes.


Troubleshooting

  • Async CUDA errors: CUDA_LAUNCH_BLOCKING=1
  • Memory errors: compute-sanitizer --tool memcheck python ...
  • Build is too slow / OOM: reduce MAX_JOBS and SGL_KERNEL_COMPILE_THREADS
  • Binary bloat: use python/sglang/kernels/aot/analyze_whl_kernel_sizes.py
  • CMake sources list: if your .cu file is missing from SOURCES, the symbol will be undefined at link time

References

  • python/sglang/kernels/aot/README.md
  • python/sglang/kernels/aot/include/sgl_kernel_ops.h
  • python/sglang/kernels/aot/csrc/common_extension.cc
  • python/sglang/kernels/aot/CMakeLists.txt
  • python/sglang/kernels/aot/include/utils.h — DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FLOAT_FP16 macro and friends
  • python/sglang/kernels/aot/csrc/elementwise/activation.cu — reference for the FP16/BF16/FP32 dispatch pattern

Summary of Files Created/Modified

python/sglang/kernels/aot/csrc/elementwise/scale.cu          # NEW: CUDA kernel + launcher
python/sglang/kernels/aot/include/sgl_kernel_ops.h           # MODIFIED: C++ declaration
python/sglang/kernels/aot/csrc/common_extension.cc           # MODIFIED: schema + dispatch registration
python/sglang/kernels/aot/CMakeLists.txt                     # MODIFIED: add source file (alphabetical)
python/sglang/kernels/aot/python/sgl_kernel/elementwise.py   # MODIFIED: Python wrapper
python/sglang/kernels/aot/python/sgl_kernel/__init__.py      # MODIFIED: re-export Python API
python/sglang/kernels/aot/tests/test_scale.py                # NEW: tests
python/sglang/kernels/aot/benchmark/bench_scale.py           # NEW: benchmark

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/add-sgl-kernel of sgl-project/sglang.

Open the folder on GitHubat commit 1c42ad3

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Add Sgl Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Sgl Kernel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Sgl Kernel this skillsgl-project/sglang37k2 repos~3.4kAutomated safety check: PassApache-2.0
Add Jit Kernelguqiong96/Lsglang1431 repos~10kAutomated safety check: PassApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Paddle Op DevPaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT

Similar skills

  • Add Jit Kernel

    guqiong96/Lsglang

    Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

    143 GitHub starsUsed in 1 repo~10k tokens
    AI & LLM EngineeringAuto-check passed
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Cute Dsl Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

    1.3k GitHub stars~2.8k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 31 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Questions about Add Sgl Kernel

What does Add Sgl Kernel do?

Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks). Add Sgl Kernel is an agent skill from sgl-project/sglang.

When should I use Add Sgl Kernel?

Add Sgl Kernel fits situations like: AI & LLM Engineering work in your project.

How do I install Add Sgl Kernel in Claude Code?

Run `npx skills add sgl-project/sglang --skill add-sgl-kernel -a claude-code`. Or copy the skill folder (.agents/skills/add-sgl-kernel in sgl-project/sglang) into .claude/skills/add-sgl-kernel in your project. Claude Code loads it when a task matches its description.

How do I install Add Sgl Kernel in Codex?

Run `npx skills add sgl-project/sglang --skill add-sgl-kernel -a codex`. Or copy the skill folder (.agents/skills/add-sgl-kernel in sgl-project/sglang) into .agents/skills/add-sgl-kernel in your project. Codex loads it when a task matches its description.

Can I use Add Sgl Kernel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill add-sgl-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-sgl-kernel, .gemini/skills/add-sgl-kernel, .github/skills/add-sgl-kernel and .opencode/skills/add-sgl-kernel in your project.

What does Add Sgl Kernel need to run?

Going by SKILL.md and its folder, Add Sgl Kernel needs the command-line tools its instructions call (make, pytest and python). Our summary lists: Python 3.

Does Add Sgl Kernel access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add Sgl Kernel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Sgl Kernel use?

Add Sgl Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Sgl Kernel use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Sgl Kernel?

Skills that share tags, products or a category with Add Sgl Kernel: Add Jit Kernel (guqiong96/Lsglang, 143 stars), Paddle Build (PaddlePaddle/Paddle, 24k stars), Ako4all (TongmingLAIC/AKO4ALL, 369 stars) and Paddle Op Dev (PaddlePaddle/Paddle, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Sgl Kernel?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,829 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.