Agent skill

Large Class Style

by sgl-project in sgl-project/sglang

Code style for SGLang large classes Scheduler, TokenizerManager, and ModelRunner: frozen-code conventions and init orchestration style.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Large Class Style

skills CLI
$ npx skills add sgl-project/sglang --skill large-class-style -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang large-class-style --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/large-class-style .claude/skills/large-class-style && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
large-class-style
GitHub stars
37k
Used in
2 other repos
Token cost
~1.9k tokens
SKILL.md length
816 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Code style for SGLang large classes Scheduler, TokenizerManager, and ModelRunner: frozen-code conventions and init orchestration style.

  • Works in 2 steps: Frozen Code → init style
  • Modifying any of these three classes
  • SKILL.md covers 1. Frozen Code and 2. init style
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Large Class Style is an agent skill from sgl-project/sglang. Code style for SGLang large classes Scheduler, TokenizerManager, and ModelRunner: frozen-code conventions and init orchestration style. Use when modifying any of these three classes or reviewing changes to them.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with SGLang and Python. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Modifying any of these three classes
  • Reviewing changes to them

Example prompts

  • “/large-class-style”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Frozen Code
  2. init style

What it can do on your machine

Read from SKILL.md and the folder at commit b7b2975. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Large Class Style loads about 1.9k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 816 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit b7b2975, republished under its Apache-2.0 licence (© sgl-project). 816 words, ~1,935 tokens.

Download SKILL.mdSave it as .claude/skills/large-class-style/SKILL.md (or your agent's skills folder).
name
large-class-style
description
Code style for SGLang large classes `Scheduler`, `TokenizerManager`, and `ModelRunner`: frozen-code conventions and `__init__` orchestration style. Use when modifying any of these three classes or reviewing changes to them.

Code Style for Scheduler / TokenizerManager / ModelRunner

Conventions for SGLang's three large classes:

  • Scheduler — python/sglang/srt/managers/scheduler.py
  • TokenizerManager — python/sglang/srt/managers/tokenizer_manager.py
  • ModelRunner — python/sglang/srt/model_executor/model_runner.py

1. Frozen Code

  • Some core files are frozen: orchestration-only — a thin composition root that constructs collaborators, wires them, delegates to them, and coordinates the calls. They must stay that way.
  • Domain logic does not belong in a frozen file; it lives in a collaborator class in its own module.
1.1 Why
  • The file is a thin orchestrator over collaborator classes; freezing keeps it that way and stops it growing back into a god class.
  • Keeping domain logic in collaborators (their own files) is what makes per-file code ownership, single responsibility, and unit testing possible.
  • The orchestrator is the composition root: it may know about every collaborator, because wiring and sequencing them is its job. Coordination stays here — domain logic does not.
1.2 Frozen files
  • python/sglang/srt/model_executor/model_runner.py
1.3 Allowed: orchestration

Every statement refers to a collaborator and is one of:

  1. Construct — a short init_<thing> helper whose body is essentially a single construction (follows §2); use maybe_init_<thing> with a one-line gate when conditional.
  2. Wire — a short call that runs the helper from the orchestrator (e.g. in __init__).
  3. Delegate — calls to a collaborator's methods at the necessary call sites (self.foo.run(...)).
  4. Coordinate — the minimal control flow that selects or orders the above: an if choosing whether / which collaborator to wire or call, the order of calls, threading one call's result into the next.
  • Heuristic: a statement is allowed only if it constructs, wires, delegates, or selects/orders those — never if it computes or transforms a value beyond passing arguments and results through.
python
# model_runner.py — orchestration only.
def init_foo(self):                        # construct
    self.foo = FooManager(server_args=self.server_args, device=self.device)

self.init_foo()                            # wire (in __init__)

if self.server_args.enable_bar:            # coordinate: select
    self.bar.prepare(forward_batch)        # delegate
out = self.foo.run(forward_batch)          # delegate
self.baz.consume(out)                      # coordinate: thread result into next delegate
1.4 Not allowed: domain logic
  • Config building, data transformation, algorithm bodies, math, post-processing — any branch or loop that computes rather than coordinates.
  • It belongs in the collaborator.
python
# NOT allowed in a frozen file: domain logic inlined.
self.foo = None
if self.server_args.enable_foo:
    config = build_foo_config(self.model_config, self.device)   # config logic in frozen file
    self.foo = FooManager(config)                               # inline construction, not via (maybe_)init_foo
    out = [step(x) for x in batch]                              # computation, not coordination
  • Fix: move that body into FooManager (its __init__ or a factory) plus a (maybe_)init_foo helper.
1.5 Where coordination logic goes
  1. Default: extract. Pull cohesive coordination into a low-coupling collaborator (an initializer, a forward pipeline) and delegate to it.
  2. Residue stays. Coordination that can't be cohesively extracted may remain — but only the minimal Coordinate form above, kept pseudocode-readable. This is the explicit exception, not a fallback; note why it stays.
  • When the residue outgrows pseudocode, that is the signal to extract a dedicated coordinator — not to keep inlining.
1.6 Pass what the collaborator needs, not the god object
  • When you extract domain logic into a collaborator (a factory, an initializer, a pipeline), give it the specific values it needs — model_config, device, the sizes — not the whole frozen object (ModelRunner, Scheduler).

  • Passing the god object back re-creates the coupling the split was meant to remove: the module still reads dozens of attributes off it, can't be unit-tested without building the whole class, and every field rename ripples back in.

  • Default to narrow, keyword args. Reference shape: layer_setup.resolve_layer_indices(*, model, model_config, is_draft_worker, spec_algorithm).

  • Return a small frozen struct and let the orchestrator assign it onto its own fields. The collaborator should not reach back in and mutate the god object.

  • If a leaf genuinely needs the live object — its constructor contract already takes the runner, or it reads state that mutates after init — confine that dependency to the smallest leaf and pass narrow args everywhere above it. Note why it can't be narrowed.

Show full SKILL.md (270 more words)Show less
1.7 If you do pass the god object, keep it read-only
  • A callee that genuinely takes the live object should read fields off it and return results; it writes fields back only when there is genuinely no other way.
  • The orchestrator owns the assignment onto its own fields.
  • Why: a callee that mutates the god object scatters its writes across other modules — you can no longer see what ModelRunner owns by reading model_runner.py, the hidden writes race with the orchestrator's own ordering, and the callee silently depends on being invoked at exactly the right moment.
python
# Good — callee reads the runner and returns a small frozen struct; the orchestrator
# owns the writes.
# model_runner.py
class ModelRunner:
    def bar(self):
        self.foo_result = foo(self)

# another_file.py
def foo(model_runner) -> FooResult:
    return FooResult(a=xx, b=yy, c=zz)

# Avoid — callee reaches back in and writes the runner's fields.
# model_runner.py
class ModelRunner:
    def bar(self):
        foo(self)

# another_file.py
def foo(model_runner):
    model_runner.a = xx
    model_runner.b = yy
    model_runner.c = zz

2. __init__ style

Apply when modifying the __init__ of the three classes above.

2.1 Why
  • Downstream forks override one piece (tokenizer, KV cache, IPC, …).
  • Inline logic forces them to copy the whole __init__, which rots against upstream.
  • Splitting into init_* helpers lets them override exactly what they need.
  • Reference shape: TokenizerManager.__init__ in python/sglang/srt/managers/tokenizer_manager.py.
2.2 Rules
  • __init__ is an orchestrator. Sequence of self.init_*(...) calls + minimal glue. No non-trivial construction inlined.
  • One helper per overridable unit. Each init_* = one concern a subclass might swap. Don't lump.
  • Naming: init_<thing> (snake_case, names the component). Conditional construction → maybe_init_<thing>, gate inside the helper.
  • No silent state coupling. A helper only reads self.* set by earlier helpers. Ordering lives in __init__. Shared intermediates → pass as args, not via self.*.
  • New logic = new helper. Default to adding init_<thing>, not another inline block. One-line self.foo = server_args.foo is fine; structured logic is not.
  • Preserve override points. Prefer additive changes to existing init_* signatures. Breaking changes → call out in PR.
2.3 Scope
  • Only the three classes listed above.
  • Not other manager-style classes, not small dataclass/utility constructors.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/large-class-style of sgl-project/sglang.

Open the folder on GitHubat commit b7b2975

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Large Class Style next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Large Class Style compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Large Class Style this skillsgl-project/sglang37k2 repos~1.9kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
One EvalOpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.0
Add Jit Kernelguqiong96/Lsglang1431 repos~10kAutomated safety check: PassApache-2.0
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT
Hyperloom Workload Optimizeramd/skills398—~1.7kAutomated safety check: NotesMIT

Similar skills

  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • One Eval

    OpenDCAI/One-Eval

    驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

    165 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Jit Kernel

    guqiong96/Lsglang

    Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

    143 GitHub starsUsed in 1 repo~10k tokens
    AI & LLM EngineeringAuto-check passed
  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

    398 GitHub stars~1.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • LLM Serving Framework Benchmark

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

    911 GitHub stars~7.5k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 31 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Works with

Questions about Large Class Style

What does Large Class Style do?

Code style for SGLang large classes Scheduler, TokenizerManager, and ModelRunner: frozen-code conventions and init orchestration style. Large Class Style is an agent skill from sgl-project/sglang. Code style for SGLang large classes Scheduler, TokenizerManager, and ModelRunner: frozen-code conventions and init orchestration style.

When should I use Large Class Style?

Large Class Style fits situations like: modifying any of these three classes; reviewing changes to them.

How do I install Large Class Style in Claude Code?

Run `npx skills add sgl-project/sglang --skill large-class-style -a claude-code`. Or copy the skill folder (.agents/skills/large-class-style in sgl-project/sglang) into .claude/skills/large-class-style in your project. Claude Code loads it when a task matches its description.

How do I install Large Class Style in Codex?

Run `npx skills add sgl-project/sglang --skill large-class-style -a codex`. Or copy the skill folder (.agents/skills/large-class-style in sgl-project/sglang) into .agents/skills/large-class-style in your project. Codex loads it when a task matches its description.

Can I use Large Class Style in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill large-class-style -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/large-class-style, .gemini/skills/large-class-style, .github/skills/large-class-style and .opencode/skills/large-class-style in your project.

What does Large Class Style need to run?

SKILL.md names no scripts, command-line tools or credentials: Large Class Style is instructions for the agent only. Our summary lists: Python 3.

Does Large Class Style access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Large Class Style safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Large Class Style use?

Large Class Style is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Large Class Style use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Large Class Style?

Skills that share tags, products or a category with Large Class Style: Dstack Prototyping (dstackai/dstack, 2.3k stars), One Eval (OpenDCAI/One-Eval, 165 stars), Add Jit Kernel (guqiong96/Lsglang, 143 stars) and SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Large Class Style?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,851 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 8, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.