Agent skill

Sglang Diffusion Add Model

by sgl-project in sgl-project/sglang

A skill your agent uses when adding a new diffusion model or Diffusers pipeline to SGLang.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Sglang Diffusion Add Model

skills CLI
$ npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang sglang-diffusion-add-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model .claude/skills/sglang-diffusion-add-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sglang-diffusion-add-model
GitHub stars
37k
Used in
2 other repos
Token cost
~9.9k tokens
SKILL.md length
3,108 words
Files
2 (incl. references)
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when adding a new diffusion model or Diffusers pipeline to SGLang.

  • Works in 11 steps: Obtain and Study the Reference… → Evaluate Reuse of Existing Pipelines and… → Implement Model Components → …
  • Adding a new diffusion model
  • SKILL.md covers Three Pipeline Styles, Key Files and Directories, Step-by-Step Implementation and Reference Implementations, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Sglang Diffusion Add Model is an agent skill from sgl-project/sglang. Use when adding a new diffusion model or Diffusers pipeline to SGLang.

Its SKILL.md is about 9.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/testing-and-accuracy.md`).

It sits in AI & LLM Engineering, covering Diffusion and image models. It works with SGLang. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Adding a new diffusion model
  • Diffusers pipeline to SGLang

Example prompts

  • “/sglang-diffusion-add-model”

Requirements

  • Python 3

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Obtain and Study the Reference Implementation
  2. Evaluate Reuse of Existing Pipelines and Stages
  3. Implement Model Components
  4. Create Model Configs
  5. Create PipelineConfig
  6. Implement the BeforeDenoisingStage (Core Step)
  7. Define the Pipeline Class
  8. Register the Model
  9. Verify Output Quality
  10. Decide the ComfyUI Route (Optional)
  11. Opt In to BCG and Quality Fast Paths Only After Eager Parity

What it can do on your machine

Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sglang Diffusion Add Model loads about 9.9k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 24 tokens; SKILL.md has 3,108 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~24
When it runs · the whole SKILL.md, loaded when a task matches
~9.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 3,108 words, ~9,923 tokens.

Download SKILL.mdSave it as .claude/skills/sglang-diffusion-add-model/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
sglang-diffusion-add-model
description
Use when adding a new diffusion model or Diffusers pipeline to SGLang.

Add a Diffusion Model to SGLang

Use this skill when adding a new diffusion model or pipeline variant to sglang.multimodal_gen.

Three Pipeline Styles

The recommended default for most new models. Uses a three-stage structure:

BeforeDenoisingStage (model-specific)  -->  DenoisingStage (standard)  -->  DecodingStage (standard)
  • BeforeDenoisingStage: A single, model-specific stage that consolidates all pre-processing logic: input validation, text encoding, image encoding, latent preparation, timestep setup. This stage is unique per model.
  • DenoisingStage: Framework-standard stage for the denoising loop (DiT/UNet forward passes). Shared across models.
  • DecodingStage: Framework-standard stage for VAE decoding. Shared across models.

Why recommended? Modern diffusion models have highly heterogeneous pre-processing requirements (different text encoders, different latent formats, different conditioning mechanisms). The Hybrid approach keeps pre-processing isolated per model, avoids fragile shared stages with excessive conditional logic, and lets developers port Diffusers reference code quickly.

Style B: Modular Composition Style

Uses the framework's fine-grained standard stages (TextEncodingStage, LatentPreparationStage, TimestepPreparationStage, etc.) to build the pipeline by composition.

This style is appropriate when:

  • The new model's pre-processing can largely reuse existing stages — e.g., a model that uses standard CLIP/T5 text encoding + standard latent preparation with minimal customization. In this case, add_standard_t2i_stages() or add_standard_ti2i_stages() may be all you need.
  • A model-specific optimization needs to be extracted as a standalone stage — e.g., a specialized encoding or conditioning step that benefits from being a separate stage for profiling, parallelism control, or reuse across multiple pipeline variants.

See existing Modular examples: QwenImagePipeline (uses add_standard_t2i_stages), FluxPipeline, WanPipeline, SanaPipeline, StableDiffusion3Pipeline, and ZImagePipeline.

Style C: Native Task-Contract Pipeline

Use this only when one checkpoint exposes multiple tightly coupled modalities or request profiles that cannot be represented safely by generic image/video sampling fields. MiniMax-H3 is the reference: it selects FL2VA or Ref2VA weights from one root model ID, validates canonical task / conditions / target requests before queueing, packs text/video/audio tokens into one denoise sequence, and returns synchronized video plus audio.

This style still uses ComposedPipelineBase, but owns a model-specific chain under stages/model_specific_stages/<model>/. Keep request validation, media materialization, packed-sequence construction, per-modality encode/decode, and presentation as explicit stages. Do not force coupled state into the standard DenoisingStage / DecodingStage contract just to resemble a simpler model.

Choose this style only with source evidence that the public API, scheduler, or joint latent state needs it. Preserve one canonical request object from API admission through offline generation and server execution so the two entry points cannot silently diverge.

How to Choose
SituationRecommended Style
Model has unique/complex pre-processing (VLM captioning, AR token generation, custom latent packing, etc.)Hybrid — consolidate into a BeforeDenoisingStage
Model jointly denoises multiple modalities or exposes partitioned task contracts from one root checkpointNative task contract — use MiniMax-H3 as the reference and keep model-specific stages explicit
Model fits neatly into standard text-to-image or text+image-to-image patternModular — use add_standard_t2i_stages() / add_standard_ti2i_stages()
Porting a Diffusers pipeline with many custom stepsHybrid — copy the __call__ logic into a single stage
Adding a variant of an existing model that shares most logicModular — reuse existing stages, customize via PipelineConfig callbacks
A specific pre-processing step needs special parallelism or profiling isolationModular — extract that step as a dedicated stage

Key principle (standard-denoise styles): For Hybrid and Modular pipelines, the stage(s) before DenoisingStage must produce a Req batch object with all the standard tensor fields that DenoisingStage expects (latents, timesteps, prompt embeds, and model-specific conditioning). Native task-contract pipelines may own a different denoise/decode contract; keep that divergence explicit and covered by request-contract tests.


Key Files and Directories

PurposePath
Pipeline classespython/sglang/multimodal_gen/runtime/pipelines/
Model-specific stagespython/sglang/multimodal_gen/runtime/pipelines_core/stages/model_specific_stages/
PipelineStage base classpython/sglang/multimodal_gen/runtime/pipelines_core/stages/base.py
Pipeline base classpython/sglang/multimodal_gen/runtime/pipelines_core/composed_pipeline_base.py
Standard stages (Denoising, Decoding)python/sglang/multimodal_gen/runtime/pipelines_core/stages/
Pipeline configspython/sglang/multimodal_gen/configs/pipeline_configs/
Sampling paramspython/sglang/multimodal_gen/configs/sample/
DiT model implementationspython/sglang/multimodal_gen/runtime/models/dits/
VAE implementationspython/sglang/multimodal_gen/runtime/models/vaes/
Encoder implementationspython/sglang/multimodal_gen/runtime/models/encoders/
Scheduler implementationspython/sglang/multimodal_gen/runtime/models/schedulers/
Model/VAE/DiT configspython/sglang/multimodal_gen/configs/models/dits/, vaes/, encoders/
Central registrypython/sglang/multimodal_gen/registry.py
Model component registrypython/sglang/multimodal_gen/runtime/models/registry.py
Current support listdocs/docs/sglang-diffusion/compatibility_matrix.mdx

Step-by-Step Implementation

Step 1: Obtain and Study the Reference Implementation

Before writing any code, obtain the model's reference implementation or Diffusers pipeline code. You need the actual source code to work from — do not guess or assume the model's architecture. If the user already gave a HuggingFace model ID or repo, inspect that yourself first. Ask the user only when the reference implementation is private, ambiguous, or otherwise unavailable. Typical sources are:

  • The model's Diffusers pipeline source (e.g., the pipeline_*.py file from the diffusers library or HuggingFace repo)
  • Or the model's official reference implementation (e.g., from the model author's GitHub repo)
  • Or the HuggingFace model ID so you can look up model_index.json and the associated pipeline class

Once you have the reference code, study it thoroughly:

  1. Find the model's model_index.json to identify required modules (text_encoder, vae, transformer, scheduler, etc.)
  2. Read the Diffusers pipeline's __call__ method end-to-end. Identify:
    • How text prompts are encoded
    • How latents are prepared (shape, dtype, scaling)
    • How timesteps/sigmas are computed
    • What conditioning kwargs the DiT/UNet expects
    • How the denoising loop works (classifier-free guidance, etc.)
    • How VAE decoding is done (scaling factors, tiling, etc.)
Step 2: Evaluate Reuse of Existing Pipelines and Stages

Before creating any new files, check whether an existing pipeline or stage can be reused or extended. Only create new pipelines/stages when the existing ones would require extensive modifications or when no similar implementation exists.

Specifically:

  1. Compare the new model's architecture against existing pipelines before creating files. Current native families include MiniMax-H3, Krea-2, LTX-2/2.3/2.5, HunyuanVideo/FastHunyuan, Wan/FastWan/TurboWan/LingBot World/LingBot Video MoE, MOVA, FLUX/FLUX.2/Klein, LongCat-Image, Z-Image, Qwen-Image/edit/layered, GLM-Image, SD3, Hunyuan3D, Helios, Cosmos3 Nano/Super/Edge/distilled, SANA/SANA-Video/SANA-WM, FireRed, ERNIE-Image, JoyAI, and Ideogram4. If the new model shares most of its structure with an existing one (e.g., same text encoders, similar latent format, compatible denoising loop), prefer:
    • Adding a new config variant to the existing pipeline rather than creating a new pipeline class
    • Reusing the existing BeforeDenoisingStage with minor parameter differences
    • Using add_standard_t2i_stages() / add_standard_ti2i_stages() / add_standard_ti2v_stages() if the model fits standard patterns
  2. Check existing stages in runtime/pipelines_core/stages/ and stages/model_specific_stages/. If an existing stage handles 80%+ of what the new model needs, extend it rather than duplicating it.
  3. Check existing model components — many models share VAEs (e.g., AutoencoderKL), text encoders (CLIP, T5), and schedulers. Reuse these directly instead of re-implementing.

Rule of thumb: Only create a new file when the existing implementation would need substantial structural changes to accommodate the new model, or when no architecturally similar implementation exists.

Step 3: Implement Model Components

Adapt or implement the model's core components in the appropriate directories.

DiT/Transformer (runtime/models/dits/{model_name}.py):

python
# python/sglang/multimodal_gen/runtime/models/dits/my_model.py

import torch
import torch.nn as nn

from sglang.multimodal_gen.runtime.layers.layernorm import (
    LayerNormScaleShift,
    RMSNormScaleShift,
)
from sglang.multimodal_gen.runtime.layers.attention.selector import (
    get_attn_backend,
)


class MyModelTransformer2DModel(nn.Module):
    """DiT model for MyModel.

    Adapt from the Diffusers/reference implementation. Key points:
    - Use SGLang's fused LayerNorm/RMSNorm ops (see `existing-fast-paths.md` under the benchmark/profile skill)
    - Use SGLang's attention backend selector
    - Keep the same parameter naming as Diffusers for weight loading compatibility
    """

    def __init__(self, config):
        super().__init__()
        # ... model layers ...

    def forward(
        self,
        hidden_states: torch.Tensor,
        encoder_hidden_states: torch.Tensor,
        timestep: torch.Tensor,
        # ... model-specific kwargs ...
    ) -> torch.Tensor:
        # ... forward pass ...
        return output

Tensor Parallel (TP) and Sequence Parallel (SP): For multi-GPU deployment, it is recommended to add TP/SP support to the DiT model. This can be done incrementally after the single-GPU implementation is verified. Reference existing implementations and adapt to your model's architecture:

  • Wan model (runtime/models/dits/wanvideo.py) — Full TP + SP reference:
    • TP: Uses ColumnParallelLinear for Q/K/V projections, RowParallelLinear for output projections, attention heads divided by tp_size
    • SP: Sequence dimension sharding via get_sp_world_size(), padding for alignment, sequence_model_parallel_all_gather for aggregation
    • Cross-attention skips SP (skip_sequence_parallel=is_cross_attention)
  • Qwen-Image model (runtime/models/dits/qwen_image.py) — SP + USPAttention reference:
    • SP: Uses USPAttention (Ulysses + Ring Attention), configured via --ulysses-degree / --ring-degree
    • TP: Uses MergedColumnParallelLinear for QKV (with Nunchaku quantization), ReplicatedLinear otherwise

Important: These are references only — each model has its own architecture and parallelism requirements. Consider:

  • How attention heads can be divided across TP ranks
  • Whether the model's sequence dimension is naturally shardable for SP
  • Which linear layers benefit from column/row parallel sharding vs. replication
  • Whether cross-attention or other special modules need SP exclusion

Key imports for distributed support:

python
from sglang.multimodal_gen.runtime.distributed import (
    divide,
    get_sp_group,
    get_sp_world_size,
    get_tp_world_size,
    sequence_model_parallel_all_gather,
)
from sglang.multimodal_gen.runtime.layers.linear import (
    ColumnParallelLinear,
    RowParallelLinear,
    ReplicatedLinear,
)

VAE (runtime/models/vaes/{model_name}.py): Implement if the model uses a non-standard VAE. Many models reuse existing VAEs.

Encoders (runtime/models/encoders/{model_name}.py): Implement if the model uses custom text/image encoders.

Schedulers (runtime/models/schedulers/{scheduler_name}.py): Implement if the model requires a custom scheduler not available in Diffusers.

Step 4: Create Model Configs

DiT Config (configs/models/dits/{model_name}.py):

python
# python/sglang/multimodal_gen/configs/models/dits/mymodel.py

from dataclasses import dataclass, field

from sglang.multimodal_gen.configs.models.dits.base import DiTConfig


@dataclass
class MyModelDitConfig(DiTConfig):
    arch_config: dict = field(default_factory=lambda: {
        "in_channels": 16,
        "num_layers": 24,
        "patch_size": 2,
        # ... model-specific architecture params ...
    })

VAE Config (configs/models/vaes/{model_name}.py):

python
from dataclasses import dataclass, field

from sglang.multimodal_gen.configs.models.vaes.base import VAEConfig


@dataclass
class MyModelVAEConfig(VAEConfig):
    vae_scale_factor: int = 8
    # ... VAE-specific params ...

Sampling Params (configs/sample/{model_name}.py):

python
from dataclasses import dataclass

from sglang.multimodal_gen.configs.sample.base import SamplingParams


@dataclass
class MyModelSamplingParams(SamplingParams):
    num_inference_steps: int = 50
    guidance_scale: float = 7.5
    height: int = 1024
    width: int = 1024
    # ... model-specific defaults ...
Step 5: Create PipelineConfig

The PipelineConfig holds static model configuration and defines callback methods used by the standard DenoisingStage and DecodingStage.

python
# python/sglang/multimodal_gen/configs/pipeline_configs/my_model.py

from dataclasses import dataclass, field

import torch

from sglang.multimodal_gen.configs.models import DiTConfig, VAEConfig
from sglang.multimodal_gen.configs.pipeline_configs.base import (
    ImagePipelineConfig,
    ModelTaskType,
    # PipelineConfig,              # common base for many video pipelines
    # SpatialImagePipelineConfig,  # alternative base for spatial image models
)
from sglang.multimodal_gen.configs.models.dits.mymodel import MyModelDitConfig
from sglang.multimodal_gen.configs.models.vaes.mymodel import MyModelVAEConfig


@dataclass
class MyModelPipelineConfig(ImagePipelineConfig):
    """Pipeline config for MyModel.

    This config provides callbacks that the standard DenoisingStage and
    DecodingStage use during execution. The BeforeDenoisingStage handles
    all model-specific pre-processing independently.
    """

    task_type: ModelTaskType = ModelTaskType.T2I
    vae_precision: str = "bf16"
    should_use_guidance: bool = True
    vae_tiling: bool = False
    enable_autocast: bool = False

    dit_config: DiTConfig = field(default_factory=MyModelDitConfig)
    vae_config: VAEConfig = field(default_factory=MyModelVAEConfig)

    # --- Callbacks used by DenoisingStage ---

    def get_freqs_cis(self, batch, device, rotary_emb, dtype):
        """Prepare rotary position embeddings for the DiT."""
        # Model-specific RoPE computation
        ...
        return freqs_cis

    def prepare_pos_cond_kwargs(self, batch, latent_model_input, t, **kwargs):
        """Build positive conditioning kwargs for each denoising step."""
        return {
            "hidden_states": latent_model_input,
            "encoder_hidden_states": batch.prompt_embeds[0],
            "timestep": t,
            # ... model-specific kwargs ...
        }

    def prepare_neg_cond_kwargs(self, batch, latent_model_input, t, **kwargs):
        """Build negative conditioning kwargs for CFG."""
        return {
            "hidden_states": latent_model_input,
            "encoder_hidden_states": batch.negative_prompt_embeds[0],
            "timestep": t,
            # ... model-specific kwargs ...
        }

    # --- Callbacks used by DecodingStage ---

    def get_decode_scale_and_shift(self):
        """Return (scale, shift) for latent denormalization before VAE decode."""
        return self.vae_config.latents_std, self.vae_config.latents_mean

    def post_denoising_loop(self, latents, batch):
        """Optional post-processing after the denoising loop finishes."""
        return latents.to(torch.bfloat16)

    def post_decoding(self, frames, server_args):
        """Optional post-processing after VAE decoding."""
        return frames

There is no separate VideoPipelineConfig base class. For video models, choose ModelTaskType.T2V, ModelTaskType.I2V, or ModelTaskType.TI2V, and follow existing video configs such as Wan, LTX, Hunyuan, Helios, or MOVA when deciding whether to subclass PipelineConfig directly or use a model-specific base.

Important: The prepare_pos_cond_kwargs / prepare_neg_cond_kwargs methods define what the DiT receives at each denoising step. These must match the DiT's forward() signature.

Step 6: Implement the BeforeDenoisingStage (Core Step)

This is the heart of the Hybrid pattern. Create a single stage that handles ALL pre-processing.

python
# python/sglang/multimodal_gen/runtime/pipelines_core/stages/model_specific_stages/my_model.py

import torch
from typing import List, Optional, Union

from sglang.multimodal_gen.runtime.pipelines_core.schedule_batch import Req
from sglang.multimodal_gen.runtime.pipelines_core.stages.base import PipelineStage
from sglang.multimodal_gen.runtime.server_args import ServerArgs
from sglang.multimodal_gen.runtime.distributed import get_local_torch_device
from sglang.multimodal_gen.runtime.utils.logging_utils import init_logger

logger = init_logger(__name__)


class MyModelBeforeDenoisingStage(PipelineStage):
    """Monolithic pre-processing stage for MyModel.

    Consolidates all logic before the denoising loop:
    - Input validation
    - Text/image encoding
    - Latent preparation
    - Timestep/sigma computation

    This stage produces a Req batch with all fields required by
    the standard DenoisingStage.
    """

    def __init__(self, vae, text_encoder, tokenizer, transformer, scheduler):
        super().__init__()
        self.vae = vae
        self.text_encoder = text_encoder
        self.tokenizer = tokenizer
        self.transformer = transformer
        self.scheduler = scheduler
        # ... other initialization (image processors, scale factors, etc.) ...

    # --- Internal helper methods ---
    # Copy/adapt directly from the Diffusers reference pipeline.
    # These are private to this stage; no need to make them reusable.

    def _encode_prompt(self, prompt, device, dtype):
        """Encode text prompt into embeddings."""
        # ... model-specific text encoding logic ...
        return prompt_embeds, negative_prompt_embeds

    def _prepare_latents(self, batch_size, height, width, dtype, device, generator):
        """Create initial noisy latents."""
        # ... model-specific latent preparation ...
        return latents

    def _prepare_timesteps(self, num_inference_steps, device):
        """Compute the timestep/sigma schedule."""
        # ... model-specific timestep computation ...
        return timesteps, sigmas

    # --- Main forward method ---

    @torch.no_grad()
    def forward(self, batch: Req, server_args: ServerArgs) -> Req:
        """Execute all pre-processing and populate batch for DenoisingStage.

        This method mirrors the first half of a Diffusers pipeline __call__,
        up to (but not including) the denoising loop.
        """
        device = get_local_torch_device()
        dtype = torch.bfloat16
        generator = torch.Generator(device=device).manual_seed(batch.seed)

        # 1. Encode prompt
        prompt_embeds, negative_prompt_embeds = self._encode_prompt(
            batch.prompt, device, dtype
        )

        # 2. Prepare latents
        latents = self._prepare_latents(
            batch_size=1,
            height=batch.height,
            width=batch.width,
            dtype=dtype,
            device=device,
            generator=generator,
        )

        # 3. Prepare timesteps
        timesteps, sigmas = self._prepare_timesteps(
            batch.num_inference_steps, device
        )

        # 4. Populate batch with everything DenoisingStage needs
        batch.prompt_embeds = [prompt_embeds]
        batch.negative_prompt_embeds = [negative_prompt_embeds]
        batch.latents = latents
        batch.timesteps = timesteps
        batch.num_inference_steps = len(timesteps)
        batch.sigmas = sigmas
        batch.generator = generator
        batch.raw_latent_shape = latents.shape
        batch.height = batch.height
        batch.width = batch.width

        return batch

Key fields that DenoisingStage expects on the batch (set these in your forward):

FieldTypeDescription
batch.latentstorch.TensorInitial noisy latent tensor
batch.timestepstorch.TensorTimestep schedule
batch.num_inference_stepsintNumber of denoising steps
batch.sigmaslist[float]Sigma schedule (as a list, not numpy)
batch.prompt_embedslist[torch.Tensor]Positive prompt embeddings (wrapped in list)
batch.negative_prompt_embedslist[torch.Tensor]Negative prompt embeddings (wrapped in list)
batch.generatortorch.GeneratorRNG generator for reproducibility
batch.raw_latent_shapetupleOriginal latent shape before any packing
batch.height / batch.widthintOutput dimensions
Step 7: Define the Pipeline Class

The pipeline class is minimal -- it just wires the stages together.

python
# python/sglang/multimodal_gen/runtime/pipelines/my_model.py

from sglang.multimodal_gen.runtime.pipelines_core import LoRAPipeline
from sglang.multimodal_gen.runtime.pipelines_core.composed_pipeline_base import (
    ComposedPipelineBase,
)
from sglang.multimodal_gen.runtime.pipelines_core.stages import DenoisingStage
from sglang.multimodal_gen.runtime.pipelines_core.stages.model_specific_stages.my_model import (
    MyModelBeforeDenoisingStage,
)
from sglang.multimodal_gen.runtime.server_args import ServerArgs


class MyModelPipeline(LoRAPipeline, ComposedPipelineBase):
    pipeline_name = "MyModelPipeline"  # Must match model_index.json _class_name

    _required_config_modules = [
        "text_encoder",
        "tokenizer",
        "vae",
        "transformer",
        "scheduler",
        # ... list all modules from model_index.json ...
    ]

    def create_pipeline_stages(self, server_args: ServerArgs):
        # 1. Monolithic pre-processing (model-specific)
        self.add_stage(
            MyModelBeforeDenoisingStage(
                vae=self.get_module("vae"),
                text_encoder=self.get_module("text_encoder"),
                tokenizer=self.get_module("tokenizer"),
                transformer=self.get_module("transformer"),
                scheduler=self.get_module("scheduler"),
            ),
        )

        # 2. Standard denoising loop (framework-provided)
        self.add_stage(
            DenoisingStage(
                transformer=self.get_module("transformer"),
                scheduler=self.get_module("scheduler"),
            ),
        )

        # 3. Standard VAE decoding (framework-provided)
        self.add_standard_decoding_stage()


# REQUIRED: This is how the registry discovers the pipeline
EntryClass = [MyModelPipeline]
Step 8: Register the Model

In python/sglang/multimodal_gen/registry.py, register your configs:

python
register_configs(
    sampling_param_cls=MyModelSamplingParams,
    pipeline_config_cls=MyModelPipelineConfig,
    hf_model_paths=[
        "org/my-model-name",  # HuggingFace model ID(s)
    ],
    model_detectors=[
        lambda path: "my-model" in path.lower(),
    ],
)

register_configs() does not take a model_family argument. It registers the sampling and pipeline config classes, then resolves models by exact hf_model_paths or optional detector predicates. Prefer exact hf_model_paths for public checkpoints used in docs or tests; use detector predicates only for families where local mirrors, renamed repos, or generated paths are common.

The EntryClass in your pipeline file is automatically discovered by the registry's _discover_and_register_pipelines() function -- no additional registration needed for the pipeline class itself.

Step 9: Verify Output Quality

After implementation, you must verify that the generated output is not noise. A noisy or garbled output image/video is the most common sign of an incorrect implementation. Common causes include:

  • Incorrect latent scale/shift factors (get_decode_scale_and_shift returning wrong values)
  • Wrong timestep/sigma schedule (order, dtype, or value range)
  • Mismatched conditioning kwargs (fields not matching the DiT's forward() signature)
  • Incorrect VAE decoder configuration (wrong vae_scale_factor, missing denormalization)
  • Rotary embedding style mismatch (is_neox_style set incorrectly)
  • Wrong prompt embedding format (missing list wrapping, wrong encoder output selection)

If the output is noise, the implementation is incorrect — do not ship it. Debug by:

  1. Comparing intermediate tensor values (latents, prompt_embeds, timesteps) against the Diffusers reference pipeline
  2. Running the Diffusers pipeline and SGLang pipeline side-by-side with the same seed
  3. Checking each stage's output shape and value range independently
Step 10: Decide the ComfyUI Route (Optional)

A model is reachable from ComfyUI two ways. Pick one deliberately — the wrong choice costs several hundred lines of weight-mapping code that buys nothing.

Server route. ComfyUI sends an HTTP request and SGLang runs the whole pipeline. Choose this when the model needs conditioning ComfyUI cannot supply (audio, reference materials, task routing), produces more than one modality, or has its own request contract.

Cost: nothing, if the request fits the existing generate_image / generate_video fields. If the model has extra request fields, pass them through extra_fields — the request schemas accept unknown keys, so the client in apps/ComfyUI_SGLDiffusion/core/server_api.py does not need a per-model change. Add a node in nodes.py only when the inputs are worth surfacing as ComfyUI widgets. SGLDiffusionGenerateH3 is the worked example.

Executor route. ComfyUI's KSampler drives the denoise loop and SGLang replaces the DiT forward, using ComfyUI's own text encoders and VAE. Choose this only when the model denoises a single latent tensor that ComfyUI already knows how to build and decode.

Cost, per model: a runtime/pipelines/comfyui_<model>_pipeline.py that maps ComfyUI's single-file checkpoint layout onto the native module tree (350-690 lines in the existing three), an executor in apps/ComfyUI_SGLDiffusion/executors/ that adapts latent layout and conditioning to Req, and entries in both dicts in core/generator.py.

The deciding question is not model size or modality — it is whether ComfyUI's sampler can drive the model's loop unchanged. If reproducing the conditioning inside ComfyUI would duplicate stages the server already runs, take the server route.

Show full SKILL.md (1,176 more words)Show less
Step 11: Opt In to BCG and Quality Fast Paths Only After Eager Parity

Do not put a new model behind Breakable CUDA Graph (BCG) merely because one forward captures. Diffusion BCG support has three independent admission paths:

  1. register the exact model IDs and safe basename aliases in BREAKABLE_CUDA_GRAPH_SUPPORTED_MODEL_IDS
  2. register the resolved pipeline config class in BREAKABLE_CUDA_GRAPH_SUPPORTED_PIPELINE_CONFIGS
  3. implement or select the correct prompt padder under runtime/breakable_cuda_graph/model_padders/

The third item is model semantics, not a generic shape utility. Reuse pad_masked_prompt_kwargs only when the model already consumes a real mask and zero-padding every coupled text tensor leaves attention and RoPE unchanged. Existing special cases show the common contracts:

  • Qwen pads embeddings, masks, text RoPE caches, and sequence-length metadata together; it synthesizes a mask when the eager path did not need one.
  • Ideogram pads the combined text-image sequence and carries replay-local DynamicVarlenMaskMeta; stale capture-time varlen indices are incorrect.
  • Z-Image preserves native prompt length because extra tokens change its semantics even when the padding looks conventional.
  • MiniMax-H3 buckets only within compatible packed-sequence alignment groups.
  • LongCat-Image and default SANA-Video already produce fixed 512- and 300-token contracts, respectively, so their padders are pass-through.

Keep mask construction active for batch size one. A shortcut such as if batch > 1 can make eager B=1 appear valid while BCG B=1 attends padded tokens. Any object whose values depend on live lengths must be rebuilt from static replay buffers once per replay; do not bake Python lists, varlen indices, or weakly referenced tensors from warmup into the graph.

BCG validation must prove all of the following:

  • warmup logs [Diffusion BCG] captured
  • serving logs no support disable, capture failure, or serving signature MISSED
  • lossless Eager and BCG artifacts are byte-identical for the same prompt, seed, shape, steps, guidance, dtype, and topology
  • short/long prompts exercise every intended bucket and an over-limit prompt falls back deliberately
  • video frame count and conditioning shapes match the captured signature; --warmup-resolutions specifies only width and height
  • padder and support-gate unit tests cover aliases, pipeline config, fixed lengths, masks, RoPE/position tensors, and replay-local metadata

For a non-bit-exact optimization, integrate through the request-scoped site framework under sglang.kernels.ops.diffusion.sites. Mark sites during model construction and let QualityGatedFusion mount them for both quality="lossless" and quality="high"; quality="lossless" must keep the original code path. A high-only sparse, caching, or other approximate path must remain outside this fusion gate. Eligibility must be all-or-nothing for coupled sites and fail closed on dtype, shape, layout, backend, BCG, or compile incompatibility. Add clean site-level guard/parity tests and a model wiring test instead of embedding request-policy branches throughout the DiT.

Finally, use the benchmark/profile skill's --quality-bcg-matrix to run same-GPU ABBA pairs for Eager/BCG at exact/lossless/high. Report denoise and saved request e2e separately, require at least 1.5% repeated mean e2e improvement for an optimization PR, attach profile and generated-media A/B evidence, then delete the task-owned checkpoint cache and verify zero residual weight files in the cleanup ledger.

Reference Implementations

ModelPipelineBeforeDenoisingStagePipelineConfig
GLM-Imageruntime/pipelines/glm_image.pystages/model_specific_stages/glm_image.pyconfigs/pipeline_configs/glm_image.py
Qwen-Image-Layeredruntime/pipelines/qwen_image.py (QwenImageLayeredPipeline)stages/model_specific_stages/qwen_image_layered.pyconfigs/pipeline_configs/qwen_image.py (QwenImageLayeredPipelineConfig)
Cosmos3runtime/pipelines/cosmos3_pipeline.pystages/model_specific_stages/cosmos3.pyconfigs/pipeline_configs/cosmos3.py
LongCat-Imageruntime/pipelines/longcat_image.pystages/model_specific_stages/longcat_image.pyconfigs/pipeline_configs/longcat_image.py
ErnieImageruntime/pipelines/ernie_image.pystages/model_specific_stages/ernie_image_pe.pyconfigs/pipeline_configs/ernie_image.py
Hunyuan3Druntime/pipelines/hunyuan3d_pipeline.pystages/model_specific_stages/hunyuan3d/configs/pipeline_configs/hunyuan3d.py
SANA-WMruntime/pipelines/sana_wm_pipeline.py, sana_wm_realtime_pipeline.pystages/model_specific_stages/sana_wm/configs/pipeline_configs/sana_wm.py
LingBot World realtimeruntime/pipelines/lingbot_world_causal_dmd_pipeline.pystages/model_specific_stages/lingbot_world/configs/pipeline_configs/lingbot_world.py
Krea-2runtime/pipelines/krea2.pystages/model_specific_stages/krea2.pyconfigs/pipeline_configs/krea2.py
Modular Style (when standard stages fit well)
ModelPipelineNotes
Qwen-Image (T2I)runtime/pipelines/qwen_image.pyUses add_standard_t2i_stages() — standard text encoding + latent prep fits this model
Qwen-Image-Editruntime/pipelines/qwen_image.pyUses add_standard_ti2i_stages() — standard image-to-image flow
Fluxruntime/pipelines/flux.pyUses add_standard_t2i_stages() with custom prepare_mu
FLUX.2 / FLUX.2 Kleinruntime/pipelines/flux_2.py, flux_2_klein.pyReuses FLUX.2 stages; Klein differences live in config and sampling params
Z-Imageruntime/pipelines/zimage_pipeline.pyUses standard image pipeline stages plus Z-Image-specific config/model code
Ideogram4runtime/pipelines/ideogram.pyUses dedicated text encoding and denoising stages while keeping standard latent prep
SANAruntime/pipelines/sana.pySpatial image pipeline; reuse the spatial image config pattern
SANA-Videoruntime/pipelines/sana_video.pyNative 3D transformer with model-specific text encoding and otherwise standard T2V stages
Stable Diffusion 3/3.5runtime/pipelines/stable_diffusion_3.pySpatial image pipeline; compare scheduler, VAE scale, and conditioning layout
LTX-2 / LTX-2.3 / LTX-2.5runtime/pipelines/ltx_2_pipeline.pyVideo pipeline family with one-stage, two-stage, HQ, joint audio/video, and optional LTX-2.5 diffusion-decoder variants; prefer config/loader specialization over a new pipeline
Heliosruntime/pipelines/helios_pipeline.pyVideo pipeline family with custom denoising and decoding stages
FireRed/JoyAI image editruntime/pipelines/qwen_image.py, runtime/pipelines/joy_image.pyFireRed reuses Qwen edit-plus config; JoyAI has its own edit pipeline
Wanruntime/pipelines/wan_pipeline.pyUses add_standard_ti2v_stages()
LingBot Video MoE 30Bruntime/pipelines/lingbot_video_moe.pyUses a model-specific structured-JSON text-encoding stage, then standard latent/timestep preparation, denoising, and decoding
Native Task-Contract Style (coupled multimodal requests)
ModelPipelineRequest / stage references
MiniMax-H3runtime/pipelines/minimax_h3_pipeline.pyconfigs/sample/minimax_h3.py owns the canonical request fields; stages/model_specific_stages/minimax_h3/ owns admission, material I/O, packed video/audio/text denoising, separate video/audio VAE work, and synchronized presentation

Checklist

Before submitting, verify:

Common (all styles):

  • Pipeline file exists at runtime/pipelines/{model_name}.py with EntryClass
  • PipelineConfig at configs/pipeline_configs/{model_name}.py
  • SamplingParams at configs/sample/{model_name}.py
  • DiT model at runtime/models/dits/{model_name}.py
  • DiT config at configs/models/dits/{model_name}.py
  • VAE — reuse existing (e.g., AutoencoderKL) or create new at runtime/models/vaes/
  • VAE config — reuse existing or create new at configs/models/vaes/{model_name}.py
  • Registry entry in registry.py via register_configs()
  • pipeline_name matches Diffusers model_index.json _class_name
  • _required_config_modules lists all modules from model_index.json
  • PipelineConfig callbacks (prepare_pos_cond_kwargs, get_freqs_cis, etc.) match DiT's forward() signature
  • Latent scale/shift factors are correctly configured
  • Use fused kernels where possible (see existing-fast-paths.md under the benchmark/profile skill)
  • Weight names match Diffusers for automatic loading
  • TP/SP support considered for DiT model (recommended; reference wanvideo.py for TP+SP, qwen_image.py for USPAttention)
  • Output quality verified — generated images/videos are not noise; compared against Diffusers reference output
  • BCG admission is complete or intentionally absent — model ID, pipeline config, and model-specific padding contract agree
  • BCG replay is proven when enabled — capture marker present, no signature miss/fallback, and lossless artifact hash is exact
  • Quality fast paths are request-scoped — lossless remains untouched; high-quality sites fail closed and have guard/parity tests
  • Performance evidence is controlled — same-GPU repeated e2e, profile, generated-media comparison, and task-owned weight cleanup ledger

Hybrid style only:

  • BeforeDenoisingStage at stages/model_specific_stages/{model_name}.py
  • BeforeDenoisingStage.forward() populates all fields needed by DenoisingStage

Native task-contract style only:

  • Root checkpoint plus variant selection maps to the intended partition; do not require users to discover internal subdirectories
  • Offline generate and HTTP serving lower through the same validated request contract
  • Task, condition role/order, target canvas/time, and output container are rejected early when invalid
  • Joint-modality correctness covers every output stream; a valid video is insufficient when the model also generates audio or action data

Common Pitfalls

  1. batch.sigmas must be a Python list, not a numpy array. Use .tolist() to convert.
  2. batch.prompt_embeds is a list of tensors (one per encoder), not a single tensor. Wrap with [tensor].
  3. Don't forget batch.raw_latent_shape -- DecodingStage uses it to unpack latents.
  4. Rotary embedding style matters: is_neox_style=True = split-half rotation, is_neox_style=False = interleaved. Check the reference model carefully.
  5. VAE precision: Many VAEs need fp32 or bf16 for numerical stability. Set vae_precision in the PipelineConfig accordingly.
  6. Avoid forcing model-specific logic into shared stages: If your model's pre-processing doesn't naturally fit the existing standard stages, prefer the Hybrid pattern with a dedicated BeforeDenoisingStage rather than adding conditional branches to shared stages.

After Implementation: Tests and Performance Data

After the model produces non-noise output, read references/testing-and-accuracy.md before adding GPU cases, component-accuracy skips/hooks, suite entries, or benchmark claims. That reference tracks the current gpu_cases.py, DiffusionTestCase.run_component_accuracy_check, single_test_file/component_accuracy/, and run_suite.py split.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model of sgl-project/sglang.

  • SKILL.md
  • references/testing-and-accuracy.md

Open the folder on GitHubat commit f620d73

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sglang Diffusion Add Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sglang Diffusion Add Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sglang Diffusion Add Model this skillsgl-project/sglang37k2 repos~9.9kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Adapt New Diffusion Modelintel/auto-round1.6k—~2.8kAutomated safety check: PassApache-2.0
Add Pipelineverl-project/verl-omni1.2k—~1kAutomated safety check: PassApache-2.0
Clean Startup Logguqiong96/Lsglang1431 repos~4.5kAutomated safety check: PassApache-2.0
Comfyui AnimatoolShiroEirin/comfyui-good-anima481—~4.6kAutomated safety check: PassGPL-3.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

    1.6k GitHub stars~2.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Pipeline

    verl-project/verl-omni

    Router for adding a diffusion or omni pipeline to verl-omni.

    1.2k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Clean Startup Log

    guqiong96/Lsglang

    Clean up noisy startup warnings and spurious prints in SGLang server logs.

    143 GitHub starsUsed in 1 repo~4.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Comfyui Animatool

    ShiroEirin/comfyui-good-anima

    Route ALL Anima image generation: validate Danbooru hard anchors, form visual brief, assemble English prompts and args, then load comfyui-manager for workflow execution.

    481 GitHub stars~4.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Stage1 Add Vae

    End2End-Diffusion/diffusion-bench

    Add a new HuggingFace-supported VAE to the stage1 tokenizer pipeline.

    105 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Works with

Questions about Sglang Diffusion Add Model

What does Sglang Diffusion Add Model do?

A skill your agent uses when adding a new diffusion model or Diffusers pipeline to SGLang. Sglang Diffusion Add Model is an agent skill from sgl-project/sglang. Use when adding a new diffusion model or Diffusers pipeline to SGLang.

When should I use Sglang Diffusion Add Model?

Sglang Diffusion Add Model fits situations like: adding a new diffusion model; diffusers pipeline to SGLang.

How do I install Sglang Diffusion Add Model in Claude Code?

Run `npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a claude-code`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model in sgl-project/sglang) into .claude/skills/sglang-diffusion-add-model in your project. Claude Code loads it when a task matches its description.

How do I install Sglang Diffusion Add Model in Codex?

Run `npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a codex`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model in sgl-project/sglang) into .agents/skills/sglang-diffusion-add-model in your project. Codex loads it when a task matches its description.

Can I use Sglang Diffusion Add Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-diffusion-add-model, .gemini/skills/sglang-diffusion-add-model, .github/skills/sglang-diffusion-add-model and .opencode/skills/sglang-diffusion-add-model in your project.

What does Sglang Diffusion Add Model need to run?

SKILL.md names no scripts, command-line tools or credentials: Sglang Diffusion Add Model is instructions for the agent only. Our summary lists: Python 3.

Does Sglang Diffusion Add Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sglang Diffusion Add Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sglang Diffusion Add Model use?

Sglang Diffusion Add Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sglang Diffusion Add Model use?

About 9.9k tokens (SKILL.md is roughly 40k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Sglang Diffusion Add Model?

Skills that share tags, products or a category with Sglang Diffusion Add Model: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Adapt New Diffusion Model (intel/auto-round, 1.6k stars), Add Pipeline (verl-project/verl-omni, 1.2k stars) and Clean Startup Log (guqiong96/Lsglang, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sglang Diffusion Add Model?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.