SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
A skill your agent uses when adding a new diffusion model or Diffusers pipeline to SGLang.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-add-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model .claude/skills/sglang-diffusion-add-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "sglang-diffusion-add-model" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model into .claude/skills/sglang-diffusion-add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-add-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-add-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model .agents/skills/sglang-diffusion-add-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "sglang-diffusion-add-model" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model into .agents/skills/sglang-diffusion-add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-add-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-add-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model .cursor/skills/sglang-diffusion-add-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "sglang-diffusion-add-model" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model into .cursor/skills/sglang-diffusion-add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-add-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-add-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model .gemini/skills/sglang-diffusion-add-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "sglang-diffusion-add-model" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model into .gemini/skills/sglang-diffusion-add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-add-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang sglang-diffusion-add-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model .github/skills/sglang-diffusion-add-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "sglang-diffusion-add-model" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model into .github/skills/sglang-diffusion-add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-add-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-add-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model .opencode/skills/sglang-diffusion-add-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "sglang-diffusion-add-model" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model into .opencode/skills/sglang-diffusion-add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-add-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
sglang-diffusion-add-modelA skill your agent uses when adding a new diffusion model or Diffusers pipeline to SGLang.
Sglang Diffusion Add Model is an agent skill from sgl-project/sglang. Use when adding a new diffusion model or Diffusers pipeline to SGLang.
Its SKILL.md is about 9.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/testing-and-accuracy.md`).
It sits in AI & LLM Engineering, covering Diffusion and image models. It works with SGLang. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
11 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Sglang Diffusion Add Model loads about 9.9k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 24 tokens; SKILL.md has 3,108 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 3,108 words, ~9,923 tokens.
.claude/skills/sglang-diffusion-add-model/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use this skill when adding a new diffusion model or pipeline variant to sglang.multimodal_gen.
The recommended default for most new models. Uses a three-stage structure:
BeforeDenoisingStage (model-specific) --> DenoisingStage (standard) --> DecodingStage (standard)Why recommended? Modern diffusion models have highly heterogeneous pre-processing requirements (different text encoders, different latent formats, different conditioning mechanisms). The Hybrid approach keeps pre-processing isolated per model, avoids fragile shared stages with excessive conditional logic, and lets developers port Diffusers reference code quickly.
Uses the framework's fine-grained standard stages (TextEncodingStage, LatentPreparationStage, TimestepPreparationStage, etc.) to build the pipeline by composition.
This style is appropriate when:
add_standard_t2i_stages() or add_standard_ti2i_stages() may be all you need.See existing Modular examples: QwenImagePipeline (uses add_standard_t2i_stages), FluxPipeline, WanPipeline, SanaPipeline, StableDiffusion3Pipeline, and ZImagePipeline.
Use this only when one checkpoint exposes multiple tightly coupled modalities
or request profiles that cannot be represented safely by generic image/video
sampling fields. MiniMax-H3 is the reference: it selects FL2VA or Ref2VA
weights from one root model ID, validates canonical task / conditions /
target requests before queueing, packs text/video/audio tokens into one
denoise sequence, and returns synchronized video plus audio.
This style still uses ComposedPipelineBase, but owns a model-specific chain
under stages/model_specific_stages/<model>/. Keep request validation, media
materialization, packed-sequence construction, per-modality encode/decode, and
presentation as explicit stages. Do not force coupled state into the standard
DenoisingStage / DecodingStage contract just to resemble a simpler model.
Choose this style only with source evidence that the public API, scheduler, or joint latent state needs it. Preserve one canonical request object from API admission through offline generation and server execution so the two entry points cannot silently diverge.
| Situation | Recommended Style |
|---|---|
| Model has unique/complex pre-processing (VLM captioning, AR token generation, custom latent packing, etc.) | Hybrid — consolidate into a BeforeDenoisingStage |
| Model jointly denoises multiple modalities or exposes partitioned task contracts from one root checkpoint | Native task contract — use MiniMax-H3 as the reference and keep model-specific stages explicit |
| Model fits neatly into standard text-to-image or text+image-to-image pattern | Modular — use add_standard_t2i_stages() / add_standard_ti2i_stages() |
| Porting a Diffusers pipeline with many custom steps | Hybrid — copy the __call__ logic into a single stage |
| Adding a variant of an existing model that shares most logic | Modular — reuse existing stages, customize via PipelineConfig callbacks |
| A specific pre-processing step needs special parallelism or profiling isolation | Modular — extract that step as a dedicated stage |
Key principle (standard-denoise styles): For Hybrid and Modular pipelines,
the stage(s) before DenoisingStage must produce a Req batch object with all
the standard tensor fields that DenoisingStage expects (latents, timesteps,
prompt embeds, and model-specific conditioning). Native task-contract pipelines
may own a different denoise/decode contract; keep that divergence explicit and
covered by request-contract tests.
| Purpose | Path |
|---|---|
| Pipeline classes | python/sglang/multimodal_gen/runtime/pipelines/ |
| Model-specific stages | python/sglang/multimodal_gen/runtime/pipelines_core/stages/model_specific_stages/ |
| PipelineStage base class | python/sglang/multimodal_gen/runtime/pipelines_core/stages/base.py |
| Pipeline base class | python/sglang/multimodal_gen/runtime/pipelines_core/composed_pipeline_base.py |
| Standard stages (Denoising, Decoding) | python/sglang/multimodal_gen/runtime/pipelines_core/stages/ |
| Pipeline configs | python/sglang/multimodal_gen/configs/pipeline_configs/ |
| Sampling params | python/sglang/multimodal_gen/configs/sample/ |
| DiT model implementations | python/sglang/multimodal_gen/runtime/models/dits/ |
| VAE implementations | python/sglang/multimodal_gen/runtime/models/vaes/ |
| Encoder implementations | python/sglang/multimodal_gen/runtime/models/encoders/ |
| Scheduler implementations | python/sglang/multimodal_gen/runtime/models/schedulers/ |
| Model/VAE/DiT configs | python/sglang/multimodal_gen/configs/models/dits/, vaes/, encoders/ |
| Central registry | python/sglang/multimodal_gen/registry.py |
| Model component registry | python/sglang/multimodal_gen/runtime/models/registry.py |
| Current support list | docs/docs/sglang-diffusion/compatibility_matrix.mdx |
Before writing any code, obtain the model's reference implementation or Diffusers pipeline code. You need the actual source code to work from — do not guess or assume the model's architecture. If the user already gave a HuggingFace model ID or repo, inspect that yourself first. Ask the user only when the reference implementation is private, ambiguous, or otherwise unavailable. Typical sources are:
pipeline_*.py file from the diffusers library or HuggingFace repo)model_index.json and the associated pipeline classOnce you have the reference code, study it thoroughly:
model_index.json to identify required modules (text_encoder, vae, transformer, scheduler, etc.)__call__ method end-to-end. Identify:Before creating any new files, check whether an existing pipeline or stage can be reused or extended. Only create new pipelines/stages when the existing ones would require extensive modifications or when no similar implementation exists.
Specifically:
BeforeDenoisingStage with minor parameter differencesadd_standard_t2i_stages() / add_standard_ti2i_stages() / add_standard_ti2v_stages() if the model fits standard patternsruntime/pipelines_core/stages/ and stages/model_specific_stages/. If an existing stage handles 80%+ of what the new model needs, extend it rather than duplicating it.AutoencoderKL), text encoders (CLIP, T5), and schedulers. Reuse these directly instead of re-implementing.Rule of thumb: Only create a new file when the existing implementation would need substantial structural changes to accommodate the new model, or when no architecturally similar implementation exists.
Adapt or implement the model's core components in the appropriate directories.
DiT/Transformer (runtime/models/dits/{model_name}.py):
# python/sglang/multimodal_gen/runtime/models/dits/my_model.py
import torch
import torch.nn as nn
from sglang.multimodal_gen.runtime.layers.layernorm import (
LayerNormScaleShift,
RMSNormScaleShift,
)
from sglang.multimodal_gen.runtime.layers.attention.selector import (
get_attn_backend,
)
class MyModelTransformer2DModel(nn.Module):
"""DiT model for MyModel.
Adapt from the Diffusers/reference implementation. Key points:
- Use SGLang's fused LayerNorm/RMSNorm ops (see `existing-fast-paths.md` under the benchmark/profile skill)
- Use SGLang's attention backend selector
- Keep the same parameter naming as Diffusers for weight loading compatibility
"""
def __init__(self, config):
super().__init__()
# ... model layers ...
def forward(
self,
hidden_states: torch.Tensor,
encoder_hidden_states: torch.Tensor,
timestep: torch.Tensor,
# ... model-specific kwargs ...
) -> torch.Tensor:
# ... forward pass ...
return outputTensor Parallel (TP) and Sequence Parallel (SP): For multi-GPU deployment, it is recommended to add TP/SP support to the DiT model. This can be done incrementally after the single-GPU implementation is verified. Reference existing implementations and adapt to your model's architecture:
runtime/models/dits/wanvideo.py) — Full TP + SP reference:ColumnParallelLinear for Q/K/V projections, RowParallelLinear for output projections, attention heads divided by tp_sizeget_sp_world_size(), padding for alignment, sequence_model_parallel_all_gather for aggregationskip_sequence_parallel=is_cross_attention)runtime/models/dits/qwen_image.py) — SP + USPAttention reference:USPAttention (Ulysses + Ring Attention), configured via --ulysses-degree / --ring-degreeMergedColumnParallelLinear for QKV (with Nunchaku quantization), ReplicatedLinear otherwiseImportant: These are references only — each model has its own architecture and parallelism requirements. Consider:
Key imports for distributed support:
from sglang.multimodal_gen.runtime.distributed import (
divide,
get_sp_group,
get_sp_world_size,
get_tp_world_size,
sequence_model_parallel_all_gather,
)
from sglang.multimodal_gen.runtime.layers.linear import (
ColumnParallelLinear,
RowParallelLinear,
ReplicatedLinear,
)VAE (runtime/models/vaes/{model_name}.py): Implement if the model uses a non-standard VAE. Many models reuse existing VAEs.
Encoders (runtime/models/encoders/{model_name}.py): Implement if the model uses custom text/image encoders.
Schedulers (runtime/models/schedulers/{scheduler_name}.py): Implement if the model requires a custom scheduler not available in Diffusers.
DiT Config (configs/models/dits/{model_name}.py):
# python/sglang/multimodal_gen/configs/models/dits/mymodel.py
from dataclasses import dataclass, field
from sglang.multimodal_gen.configs.models.dits.base import DiTConfig
@dataclass
class MyModelDitConfig(DiTConfig):
arch_config: dict = field(default_factory=lambda: {
"in_channels": 16,
"num_layers": 24,
"patch_size": 2,
# ... model-specific architecture params ...
})VAE Config (configs/models/vaes/{model_name}.py):
from dataclasses import dataclass, field
from sglang.multimodal_gen.configs.models.vaes.base import VAEConfig
@dataclass
class MyModelVAEConfig(VAEConfig):
vae_scale_factor: int = 8
# ... VAE-specific params ...Sampling Params (configs/sample/{model_name}.py):
from dataclasses import dataclass
from sglang.multimodal_gen.configs.sample.base import SamplingParams
@dataclass
class MyModelSamplingParams(SamplingParams):
num_inference_steps: int = 50
guidance_scale: float = 7.5
height: int = 1024
width: int = 1024
# ... model-specific defaults ...The PipelineConfig holds static model configuration and defines callback methods used by the standard DenoisingStage and DecodingStage.
# python/sglang/multimodal_gen/configs/pipeline_configs/my_model.py
from dataclasses import dataclass, field
import torch
from sglang.multimodal_gen.configs.models import DiTConfig, VAEConfig
from sglang.multimodal_gen.configs.pipeline_configs.base import (
ImagePipelineConfig,
ModelTaskType,
# PipelineConfig, # common base for many video pipelines
# SpatialImagePipelineConfig, # alternative base for spatial image models
)
from sglang.multimodal_gen.configs.models.dits.mymodel import MyModelDitConfig
from sglang.multimodal_gen.configs.models.vaes.mymodel import MyModelVAEConfig
@dataclass
class MyModelPipelineConfig(ImagePipelineConfig):
"""Pipeline config for MyModel.
This config provides callbacks that the standard DenoisingStage and
DecodingStage use during execution. The BeforeDenoisingStage handles
all model-specific pre-processing independently.
"""
task_type: ModelTaskType = ModelTaskType.T2I
vae_precision: str = "bf16"
should_use_guidance: bool = True
vae_tiling: bool = False
enable_autocast: bool = False
dit_config: DiTConfig = field(default_factory=MyModelDitConfig)
vae_config: VAEConfig = field(default_factory=MyModelVAEConfig)
# --- Callbacks used by DenoisingStage ---
def get_freqs_cis(self, batch, device, rotary_emb, dtype):
"""Prepare rotary position embeddings for the DiT."""
# Model-specific RoPE computation
...
return freqs_cis
def prepare_pos_cond_kwargs(self, batch, latent_model_input, t, **kwargs):
"""Build positive conditioning kwargs for each denoising step."""
return {
"hidden_states": latent_model_input,
"encoder_hidden_states": batch.prompt_embeds[0],
"timestep": t,
# ... model-specific kwargs ...
}
def prepare_neg_cond_kwargs(self, batch, latent_model_input, t, **kwargs):
"""Build negative conditioning kwargs for CFG."""
return {
"hidden_states": latent_model_input,
"encoder_hidden_states": batch.negative_prompt_embeds[0],
"timestep": t,
# ... model-specific kwargs ...
}
# --- Callbacks used by DecodingStage ---
def get_decode_scale_and_shift(self):
"""Return (scale, shift) for latent denormalization before VAE decode."""
return self.vae_config.latents_std, self.vae_config.latents_mean
def post_denoising_loop(self, latents, batch):
"""Optional post-processing after the denoising loop finishes."""
return latents.to(torch.bfloat16)
def post_decoding(self, frames, server_args):
"""Optional post-processing after VAE decoding."""
return framesThere is no separate VideoPipelineConfig base class. For video models, choose
ModelTaskType.T2V, ModelTaskType.I2V, or ModelTaskType.TI2V, and follow
existing video configs such as Wan, LTX, Hunyuan, Helios, or MOVA when deciding
whether to subclass PipelineConfig directly or use a model-specific base.
Important: The prepare_pos_cond_kwargs / prepare_neg_cond_kwargs methods define what the DiT receives at each denoising step. These must match the DiT's forward() signature.
This is the heart of the Hybrid pattern. Create a single stage that handles ALL pre-processing.
# python/sglang/multimodal_gen/runtime/pipelines_core/stages/model_specific_stages/my_model.py
import torch
from typing import List, Optional, Union
from sglang.multimodal_gen.runtime.pipelines_core.schedule_batch import Req
from sglang.multimodal_gen.runtime.pipelines_core.stages.base import PipelineStage
from sglang.multimodal_gen.runtime.server_args import ServerArgs
from sglang.multimodal_gen.runtime.distributed import get_local_torch_device
from sglang.multimodal_gen.runtime.utils.logging_utils import init_logger
logger = init_logger(__name__)
class MyModelBeforeDenoisingStage(PipelineStage):
"""Monolithic pre-processing stage for MyModel.
Consolidates all logic before the denoising loop:
- Input validation
- Text/image encoding
- Latent preparation
- Timestep/sigma computation
This stage produces a Req batch with all fields required by
the standard DenoisingStage.
"""
def __init__(self, vae, text_encoder, tokenizer, transformer, scheduler):
super().__init__()
self.vae = vae
self.text_encoder = text_encoder
self.tokenizer = tokenizer
self.transformer = transformer
self.scheduler = scheduler
# ... other initialization (image processors, scale factors, etc.) ...
# --- Internal helper methods ---
# Copy/adapt directly from the Diffusers reference pipeline.
# These are private to this stage; no need to make them reusable.
def _encode_prompt(self, prompt, device, dtype):
"""Encode text prompt into embeddings."""
# ... model-specific text encoding logic ...
return prompt_embeds, negative_prompt_embeds
def _prepare_latents(self, batch_size, height, width, dtype, device, generator):
"""Create initial noisy latents."""
# ... model-specific latent preparation ...
return latents
def _prepare_timesteps(self, num_inference_steps, device):
"""Compute the timestep/sigma schedule."""
# ... model-specific timestep computation ...
return timesteps, sigmas
# --- Main forward method ---
@torch.no_grad()
def forward(self, batch: Req, server_args: ServerArgs) -> Req:
"""Execute all pre-processing and populate batch for DenoisingStage.
This method mirrors the first half of a Diffusers pipeline __call__,
up to (but not including) the denoising loop.
"""
device = get_local_torch_device()
dtype = torch.bfloat16
generator = torch.Generator(device=device).manual_seed(batch.seed)
# 1. Encode prompt
prompt_embeds, negative_prompt_embeds = self._encode_prompt(
batch.prompt, device, dtype
)
# 2. Prepare latents
latents = self._prepare_latents(
batch_size=1,
height=batch.height,
width=batch.width,
dtype=dtype,
device=device,
generator=generator,
)
# 3. Prepare timesteps
timesteps, sigmas = self._prepare_timesteps(
batch.num_inference_steps, device
)
# 4. Populate batch with everything DenoisingStage needs
batch.prompt_embeds = [prompt_embeds]
batch.negative_prompt_embeds = [negative_prompt_embeds]
batch.latents = latents
batch.timesteps = timesteps
batch.num_inference_steps = len(timesteps)
batch.sigmas = sigmas
batch.generator = generator
batch.raw_latent_shape = latents.shape
batch.height = batch.height
batch.width = batch.width
return batchKey fields that DenoisingStage expects on the batch (set these in your forward):
| Field | Type | Description |
|---|---|---|
batch.latents | torch.Tensor | Initial noisy latent tensor |
batch.timesteps | torch.Tensor | Timestep schedule |
batch.num_inference_steps | int | Number of denoising steps |
batch.sigmas | list[float] | Sigma schedule (as a list, not numpy) |
batch.prompt_embeds | list[torch.Tensor] | Positive prompt embeddings (wrapped in list) |
batch.negative_prompt_embeds | list[torch.Tensor] | Negative prompt embeddings (wrapped in list) |
batch.generator | torch.Generator | RNG generator for reproducibility |
batch.raw_latent_shape | tuple | Original latent shape before any packing |
batch.height / batch.width | int | Output dimensions |
The pipeline class is minimal -- it just wires the stages together.
# python/sglang/multimodal_gen/runtime/pipelines/my_model.py
from sglang.multimodal_gen.runtime.pipelines_core import LoRAPipeline
from sglang.multimodal_gen.runtime.pipelines_core.composed_pipeline_base import (
ComposedPipelineBase,
)
from sglang.multimodal_gen.runtime.pipelines_core.stages import DenoisingStage
from sglang.multimodal_gen.runtime.pipelines_core.stages.model_specific_stages.my_model import (
MyModelBeforeDenoisingStage,
)
from sglang.multimodal_gen.runtime.server_args import ServerArgs
class MyModelPipeline(LoRAPipeline, ComposedPipelineBase):
pipeline_name = "MyModelPipeline" # Must match model_index.json _class_name
_required_config_modules = [
"text_encoder",
"tokenizer",
"vae",
"transformer",
"scheduler",
# ... list all modules from model_index.json ...
]
def create_pipeline_stages(self, server_args: ServerArgs):
# 1. Monolithic pre-processing (model-specific)
self.add_stage(
MyModelBeforeDenoisingStage(
vae=self.get_module("vae"),
text_encoder=self.get_module("text_encoder"),
tokenizer=self.get_module("tokenizer"),
transformer=self.get_module("transformer"),
scheduler=self.get_module("scheduler"),
),
)
# 2. Standard denoising loop (framework-provided)
self.add_stage(
DenoisingStage(
transformer=self.get_module("transformer"),
scheduler=self.get_module("scheduler"),
),
)
# 3. Standard VAE decoding (framework-provided)
self.add_standard_decoding_stage()
# REQUIRED: This is how the registry discovers the pipeline
EntryClass = [MyModelPipeline]In python/sglang/multimodal_gen/registry.py, register your configs:
register_configs(
sampling_param_cls=MyModelSamplingParams,
pipeline_config_cls=MyModelPipelineConfig,
hf_model_paths=[
"org/my-model-name", # HuggingFace model ID(s)
],
model_detectors=[
lambda path: "my-model" in path.lower(),
],
)register_configs() does not take a model_family argument. It registers the
sampling and pipeline config classes, then resolves models by exact
hf_model_paths or optional detector predicates. Prefer exact hf_model_paths
for public checkpoints used in docs or tests; use detector predicates only for
families where local mirrors, renamed repos, or generated paths are common.
The EntryClass in your pipeline file is automatically discovered by the registry's _discover_and_register_pipelines() function -- no additional registration needed for the pipeline class itself.
After implementation, you must verify that the generated output is not noise. A noisy or garbled output image/video is the most common sign of an incorrect implementation. Common causes include:
get_decode_scale_and_shift returning wrong values)forward() signature)vae_scale_factor, missing denormalization)is_neox_style set incorrectly)If the output is noise, the implementation is incorrect — do not ship it. Debug by:
A model is reachable from ComfyUI two ways. Pick one deliberately — the wrong choice costs several hundred lines of weight-mapping code that buys nothing.
Server route. ComfyUI sends an HTTP request and SGLang runs the whole pipeline. Choose this when the model needs conditioning ComfyUI cannot supply (audio, reference materials, task routing), produces more than one modality, or has its own request contract.
Cost: nothing, if the request fits the existing generate_image /
generate_video fields. If the model has extra request fields, pass them
through extra_fields — the request schemas accept unknown keys, so the
client in apps/ComfyUI_SGLDiffusion/core/server_api.py does not need a
per-model change. Add a node in nodes.py only when the inputs are worth
surfacing as ComfyUI widgets. SGLDiffusionGenerateH3 is the worked example.
Executor route. ComfyUI's KSampler drives the denoise loop and SGLang replaces the DiT forward, using ComfyUI's own text encoders and VAE. Choose this only when the model denoises a single latent tensor that ComfyUI already knows how to build and decode.
Cost, per model: a runtime/pipelines/comfyui_<model>_pipeline.py that maps
ComfyUI's single-file checkpoint layout onto the native module tree (350-690
lines in the existing three), an executor in
apps/ComfyUI_SGLDiffusion/executors/ that adapts latent layout and
conditioning to Req, and entries in both dicts in core/generator.py.
The deciding question is not model size or modality — it is whether ComfyUI's sampler can drive the model's loop unchanged. If reproducing the conditioning inside ComfyUI would duplicate stages the server already runs, take the server route.
Do not put a new model behind Breakable CUDA Graph (BCG) merely because one forward captures. Diffusion BCG support has three independent admission paths:
BREAKABLE_CUDA_GRAPH_SUPPORTED_MODEL_IDSBREAKABLE_CUDA_GRAPH_SUPPORTED_PIPELINE_CONFIGSruntime/breakable_cuda_graph/model_padders/The third item is model semantics, not a generic shape utility. Reuse
pad_masked_prompt_kwargs only when the model already consumes a real mask and
zero-padding every coupled text tensor leaves attention and RoPE unchanged.
Existing special cases show the common contracts:
DynamicVarlenMaskMeta; stale capture-time varlen indices are incorrect.Keep mask construction active for batch size one. A shortcut such as
if batch > 1 can make eager B=1 appear valid while BCG B=1 attends padded
tokens. Any object whose values depend on live lengths must be rebuilt from
static replay buffers once per replay; do not bake Python lists, varlen
indices, or weakly referenced tensors from warmup into the graph.
BCG validation must prove all of the following:
[Diffusion BCG] capturedserving signature MISSED--warmup-resolutions specifies only width and heightFor a non-bit-exact optimization, integrate through the request-scoped site
framework under sglang.kernels.ops.diffusion.sites. Mark sites during model
construction and let QualityGatedFusion mount them for both
quality="lossless" and quality="high"; quality="lossless" must keep
the original code path. A high-only sparse, caching, or other approximate
path must remain outside this fusion gate.
Eligibility must be all-or-nothing for coupled sites and fail closed on dtype,
shape, layout, backend, BCG, or compile incompatibility. Add clean site-level
guard/parity tests and a model wiring test instead of embedding request-policy
branches throughout the DiT.
Finally, use the benchmark/profile skill's --quality-bcg-matrix to run
same-GPU ABBA pairs for Eager/BCG at exact/lossless/high. Report denoise and saved
request e2e separately, require at least 1.5% repeated mean e2e improvement for
an optimization PR, attach profile and generated-media A/B evidence, then
delete the task-owned checkpoint cache and verify zero residual weight files
in the cleanup ledger.
| Model | Pipeline | BeforeDenoisingStage | PipelineConfig |
|---|---|---|---|
| GLM-Image | runtime/pipelines/glm_image.py | stages/model_specific_stages/glm_image.py | configs/pipeline_configs/glm_image.py |
| Qwen-Image-Layered | runtime/pipelines/qwen_image.py (QwenImageLayeredPipeline) | stages/model_specific_stages/qwen_image_layered.py | configs/pipeline_configs/qwen_image.py (QwenImageLayeredPipelineConfig) |
| Cosmos3 | runtime/pipelines/cosmos3_pipeline.py | stages/model_specific_stages/cosmos3.py | configs/pipeline_configs/cosmos3.py |
| LongCat-Image | runtime/pipelines/longcat_image.py | stages/model_specific_stages/longcat_image.py | configs/pipeline_configs/longcat_image.py |
| ErnieImage | runtime/pipelines/ernie_image.py | stages/model_specific_stages/ernie_image_pe.py | configs/pipeline_configs/ernie_image.py |
| Hunyuan3D | runtime/pipelines/hunyuan3d_pipeline.py | stages/model_specific_stages/hunyuan3d/ | configs/pipeline_configs/hunyuan3d.py |
| SANA-WM | runtime/pipelines/sana_wm_pipeline.py, sana_wm_realtime_pipeline.py | stages/model_specific_stages/sana_wm/ | configs/pipeline_configs/sana_wm.py |
| LingBot World realtime | runtime/pipelines/lingbot_world_causal_dmd_pipeline.py | stages/model_specific_stages/lingbot_world/ | configs/pipeline_configs/lingbot_world.py |
| Krea-2 | runtime/pipelines/krea2.py | stages/model_specific_stages/krea2.py | configs/pipeline_configs/krea2.py |
| Model | Pipeline | Notes |
|---|---|---|
| Qwen-Image (T2I) | runtime/pipelines/qwen_image.py | Uses add_standard_t2i_stages() — standard text encoding + latent prep fits this model |
| Qwen-Image-Edit | runtime/pipelines/qwen_image.py | Uses add_standard_ti2i_stages() — standard image-to-image flow |
| Flux | runtime/pipelines/flux.py | Uses add_standard_t2i_stages() with custom prepare_mu |
| FLUX.2 / FLUX.2 Klein | runtime/pipelines/flux_2.py, flux_2_klein.py | Reuses FLUX.2 stages; Klein differences live in config and sampling params |
| Z-Image | runtime/pipelines/zimage_pipeline.py | Uses standard image pipeline stages plus Z-Image-specific config/model code |
| Ideogram4 | runtime/pipelines/ideogram.py | Uses dedicated text encoding and denoising stages while keeping standard latent prep |
| SANA | runtime/pipelines/sana.py | Spatial image pipeline; reuse the spatial image config pattern |
| SANA-Video | runtime/pipelines/sana_video.py | Native 3D transformer with model-specific text encoding and otherwise standard T2V stages |
| Stable Diffusion 3/3.5 | runtime/pipelines/stable_diffusion_3.py | Spatial image pipeline; compare scheduler, VAE scale, and conditioning layout |
| LTX-2 / LTX-2.3 / LTX-2.5 | runtime/pipelines/ltx_2_pipeline.py | Video pipeline family with one-stage, two-stage, HQ, joint audio/video, and optional LTX-2.5 diffusion-decoder variants; prefer config/loader specialization over a new pipeline |
| Helios | runtime/pipelines/helios_pipeline.py | Video pipeline family with custom denoising and decoding stages |
| FireRed/JoyAI image edit | runtime/pipelines/qwen_image.py, runtime/pipelines/joy_image.py | FireRed reuses Qwen edit-plus config; JoyAI has its own edit pipeline |
| Wan | runtime/pipelines/wan_pipeline.py | Uses add_standard_ti2v_stages() |
| LingBot Video MoE 30B | runtime/pipelines/lingbot_video_moe.py | Uses a model-specific structured-JSON text-encoding stage, then standard latent/timestep preparation, denoising, and decoding |
| Model | Pipeline | Request / stage references |
|---|---|---|
| MiniMax-H3 | runtime/pipelines/minimax_h3_pipeline.py | configs/sample/minimax_h3.py owns the canonical request fields; stages/model_specific_stages/minimax_h3/ owns admission, material I/O, packed video/audio/text denoising, separate video/audio VAE work, and synchronized presentation |
Before submitting, verify:
Common (all styles):
runtime/pipelines/{model_name}.py with EntryClassconfigs/pipeline_configs/{model_name}.pyconfigs/sample/{model_name}.pyruntime/models/dits/{model_name}.pyconfigs/models/dits/{model_name}.pyAutoencoderKL) or create new at runtime/models/vaes/configs/models/vaes/{model_name}.pyregistry.py via register_configs()pipeline_name matches Diffusers model_index.json _class_name_required_config_modules lists all modules from model_index.jsonPipelineConfig callbacks (prepare_pos_cond_kwargs, get_freqs_cis, etc.) match DiT's forward() signatureexisting-fast-paths.md under the benchmark/profile skill)wanvideo.py for TP+SP, qwen_image.py for USPAttention)Hybrid style only:
stages/model_specific_stages/{model_name}.pyBeforeDenoisingStage.forward() populates all fields needed by DenoisingStageNative task-contract style only:
generate and HTTP serving lower through the same validated
request contractbatch.sigmas must be a Python list, not a numpy array. Use .tolist() to convert.batch.prompt_embeds is a list of tensors (one per encoder), not a single tensor. Wrap with [tensor].batch.raw_latent_shape -- DecodingStage uses it to unpack latents.is_neox_style=True = split-half rotation, is_neox_style=False = interleaved. Check the reference model carefully.vae_precision in the PipelineConfig accordingly.After the model produces non-noise output, read
references/testing-and-accuracy.md before
adding GPU cases, component-accuracy skips/hooks, suite entries, or benchmark
claims. That reference tracks the current gpu_cases.py,
DiffusionTestCase.run_component_accuracy_check,
single_test_file/component_accuracy/, and run_suite.py split.
© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model of sgl-project/sglang.
Open the folder on GitHubat commit f620d73
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
Sglang Diffusion Add Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Sglang Diffusion Add Model this skillsgl-project/sglang | 37k | 2 repos | ~9.9k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Adapt New Diffusion Modelintel/auto-round | 1.6k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Add Pipelineverl-project/verl-omni | 1.2k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Clean Startup Logguqiong96/Lsglang | 143 | 1 repos | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Comfyui AnimatoolShiroEirin/comfyui-good-anima | 481 | — | ~4.6k | Automated safety check: Pass | GPL-3.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
intel/auto-round
Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).
verl-project/verl-omni
Router for adding a diffusion or omni pipeline to verl-omni.
guqiong96/Lsglang
Clean up noisy startup warnings and spurious prints in SGLang server logs.
ShiroEirin/comfyui-good-anima
Route ALL Anima image generation: validate Danbooru hard anchors, form visual brief, assemble English prompts and args, then load comfyui-manager for workflow execution.
End2End-Diffusion/diffusion-bench
Add a new HuggingFace-supported VAE to the stage1 tokenizer pipeline.
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
Works with
Categories
A skill your agent uses when adding a new diffusion model or Diffusers pipeline to SGLang. Sglang Diffusion Add Model is an agent skill from sgl-project/sglang. Use when adding a new diffusion model or Diffusers pipeline to SGLang.
Sglang Diffusion Add Model fits situations like: adding a new diffusion model; diffusers pipeline to SGLang.
Run `npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a claude-code`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model in sgl-project/sglang) into .claude/skills/sglang-diffusion-add-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a codex`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-add-model in sgl-project/sglang) into .agents/skills/sglang-diffusion-add-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-diffusion-add-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-diffusion-add-model, .gemini/skills/sglang-diffusion-add-model, .github/skills/sglang-diffusion-add-model and .opencode/skills/sglang-diffusion-add-model in your project.
SKILL.md names no scripts, command-line tools or credentials: Sglang Diffusion Add Model is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Sglang Diffusion Add Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.9k tokens (SKILL.md is roughly 40k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Sglang Diffusion Add Model: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Adapt New Diffusion Model (intel/auto-round, 1.6k stars), Add Pipeline (verl-project/verl-omni, 1.2k stars) and Clean Startup Log (guqiong96/Lsglang, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.