Agent skill

Sglang Diffusion Modelopt Quant

by sgl-project in sgl-project/sglang

A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Sglang Diffusion Modelopt Quant

skills CLI
$ npx skills add sgl-project/sglang --skill sglang-diffusion-modelopt-quant -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang sglang-diffusion-modelopt-quant --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-modelopt-quant .claude/skills/sglang-diffusion-modelopt-quant && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sglang-diffusion-modelopt-quant
GitHub stars
37k
Used in
2 other repos
Token cost
~5k tokens
SKILL.md length
2,027 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

  • Works in 7 steps: Verify The BF16 Baseline First → Quantize With Official ModelOpt → Convert FP8 Exports For SGLang → …
  • Quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8
  • SKILL.md covers Overview, Core Rules, Read First and What SGLang Supports Here, plus 7 more sections
  • Calls python3 and python

What it does

Sglang Diffusion Modelopt Quant is an agent skill from sgl-project/sglang. Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with SGLang, NVIDIA AI Platform and Python. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8
  • NVFP4 checkpoint loadable
  • Benchmarkable in SGLang Diffusion

Example prompts

  • “/sglang-diffusion-modelopt-quant”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Verify The BF16 Baseline First
  2. Quantize With Official ModelOpt
  3. Convert FP8 Exports For SGLang
  4. Load The Quantized Checkpoint In SGLang
  5. Validate Accuracy
  6. Benchmark Correctly
  7. Add Model-Specific Fallbacks Only When Needed

What it can do on your machine

Read from SKILL.md and the folder at commit b7b2975. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sglang Diffusion Modelopt Quant loads about 5k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 2,027 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit b7b2975, republished under its Apache-2.0 licence (© sgl-project). 2,027 words, ~4,980 tokens.

Download SKILL.mdSave it as .claude/skills/sglang-diffusion-modelopt-quant/SKILL.md (or your agent's skills folder).
name
sglang-diffusion-modelopt-quant
description
Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

SGLang Diffusion ModelOpt Quant

Overview

Use this skill when the task is to take a diffusion transformer through the full ModelOpt workflow:

  • quantize it with NVIDIA ModelOpt
  • adapt the exported checkpoint to SGLang Diffusion
  • verify that quality holds up
  • benchmark whether the quantized checkpoint is actually faster

This skill owns the ModelOpt-to-SGLang bridge. It is not a generic kernel-tuning skill.

Core Rules

  • Use ModelOpt's official quantize.py as the PTQ source of truth.
  • Keep the workflow generic. Put model-specific fallback logic in small isolated branches, not in the main conversion path.
  • Benchmark only when BF16 and quantized commands are identical except for the checkpoint override being tested.
  • For diffusion FP8, keep dit_cpu_offload=false. dit_layerwise_offload=true is valid on the fixed path when you want lower DiT residency.
  • For multi-transformer pipelines, use per-component overrides when different components need different checkpoints.
  • For B200 NVFP4 validation, keep backend-sensitive environment variables explicit. The current default is FlashInfer TensorRT-LLM (flashinfer_trtllm); high-resolution Qwen Image can favor SGLANG_DIFFUSION_FLASHINFER_FP4_GEMM_BACKEND=cutlass, while 1024x1024 can remain BF16-faster. Benchmark the exact shape instead of assuming one backend or quantized checkpoint wins.
  • When a branch is missing the validated helper tools, refresh python/sglang/multimodal_gen/tools/build_modelopt_fp8_transformer.py, python/sglang/multimodal_gen/tools/build_modelopt_nvfp4_transformer.py, and python/sglang/multimodal_gen/tools/compare_diffusion_trajectory_similarity.py instead of inventing one-off scripts elsewhere.
  • After validating a new ModelOpt quant path, update the ModelOpt support matrix in docs/docs/sglang-diffusion/quantization.mdx before closing the task.

Read First

Read these sources before changing code:

If you are working on a new model family, inspect the transformer's config and tensor naming before changing the generic converter.

What SGLang Supports Here

This repo now contains:

Validated documentation and CI coverage currently center on these ModelOpt diffusion transformer override families:

  • FP8: FLUX.1-dev, FLUX.2-dev, Wan2.2, HunyuanVideo, Qwen Image, Qwen Image Edit
  • NVFP4: FLUX.1-dev, FLUX.2-dev, Wan2.2, Qwen Image, Qwen Image 2512, Qwen Image Edit, Qwen Image Edit 2511

Treat a new family, a new precision, or a new checkpoint layout as unsupported until it has a documented matrix row and a matching validation story. Current B200 CI also contains an Ideogram4 NVFP4 native load case (ideogram4_nvfp4_t2i via Comfy-Org/Ideogram-4). Treat that as source evidence for an existing NVFP4 path, but do not expand the ModelOpt support matrix to Ideogram4 unless docs/docs/sglang-diffusion/quantization.mdx is updated with the exact checkpoint, loader path, quality check, and benchmark scope. Before writing CLI examples, re-read the active branch's docs/docs/sglang-diffusion/quantization.mdx: FLUX.2 NVFP4 is an official black-forest-labs/* repo rather than a lmsys/* converted repo, and its preferred flag depends on the current documented loader flow. Use --transformer-path for a component override directory with config.json; use --transformer-weights-path when the repo or path should be probed as raw weights.

B200 CI coverage can include loose BF16-vs-quantized quality checks. Inspect the active branch's run_suite.py before assuming they are part of the suite; mainline and feature branches may differ. Those checks are intended to catch blank, corrupted, or obviously divergent images, not exact image parity.

Mainline documentation now tracks thirteen published ModelOpt checkpoints. Twelve live under lmsys/*; the FLUX.2 NVFP4 raw export remains black-forest-labs/FLUX.2-dev-NVFP4. Do not use older BBuf/* examples unless you are explicitly testing a historical branch.

MiniMax-H3 boundary

MiniMax-H3 is current-main evidence for the separate online FP8 path, not a validated ModelOpt PTQ/export family. Its verified B200/B300 serving recipe loads the unquantized root checkpoint with --quantization fp8 and preserves the video/audio patch projections, timestep MLP, and final video/audio heads in FP32. Do not add H3 to the ModelOpt support matrix or run the generic ModelOpt converter until an exact H3 export, loader mapping, accuracy check, and benchmark scope have been validated.

If the user asks for current H3 online quantization, route the command and quality caveats through sglang-diffusion-performance and the MiniMax-H3 cookbook. Online FP8 is approximate and must be compared against eager BF16/FP32 for both video and audio; combining it with Cache-DiT compounds two approximations.

These related SGLang PRs are useful as ModelOpt diffusion support history. Re-check the PR state and the active source tree before treating any item as current behavior, and keep the docs/CI matrix as the support boundary.

  • #23155 added Qwen Image ModelOpt FP8 support.
  • #23199 adds HunyuanVideo ModelOpt FP8 support.
  • #23373 adds a runtime quantization flag; keep PTQ/export workflows separate from runtime quant examples until the CLI behavior is merged.
  • #24024 adds transformer FP8-cast compatibility mode.
  • #24186 re-enables B200 multimodal CI with NVFP4 fixes for FLUX.2 and Wan2.2.

Do not expand the validated matrix beyond the documented rows solely because a related PR exists. Add a row only after the exact checkpoint, loader path, accuracy check, and benchmark scope are validated on the active branch.

Documentation Maintenance

  • Keep the validated ModelOpt support matrix in docs/docs/sglang-diffusion/quantization.mdx.
  • Each row should record the validated scope, the Hugging Face repo or path for the quantized DiT weights, and the key caveats.
  • If the quantized DiT weights are not published yet, write unpublished explicitly instead of leaving the field blank.

FP8 Vs NVFP4

FP8 and NVFP4 are not wired into SGLang in exactly the same way.

FP8:

  • the validated ModelOpt diffusers FP8 export still needs an extra SGLang-side conversion step
  • SGLang expects explicit weight_scale and input_scale
  • the validated path also materializes SGLang-native float8_e4m3fn weights from backbone.pt

NVFP4:

  • the official diffusers export often already contains packed FP4 weights, scale tensors, and enough safetensors metadata for SGLang to rebuild the quant config
  • in that case SGLang mainly needs to detect the checkpoint family and rearrange tensors into the runtime layout
  • this is why NVFP4 often does not need an extra offline conversion pass like FP8 does
  • backend choice matters on B200; record whether the run used the default CUTLASS path or a cuDNN-backed FlashInfer FP4 GEMM path

Important caveat:

  • "often" does not mean "always"
  • the exact load path still depends on the checkpoint family, especially whether a model uses a packed-QKV layout

Generic Workflow

1. Verify The BF16 Baseline First

Before quantizing anything:

  • run the original BF16 model in SGLang
  • fix the prompt, seed, size, step count, and GPU topology
  • save the output and perf.json

Do not start quantization work until the BF16 path is already healthy.

2. Quantize With Official ModelOpt

Use ModelOpt's official script. Generic template:

bash
python quantize.py \
  --model <model-name> \
  --override-model-path <hf-repo-or-local-model> \
  --model-dtype <Half|BFloat16> \
  --format <fp8|fp4> \
  --batch-size 1 \
  --calib-size <calib-size> \
  --n-steps <calib-steps> \
  --quantize-mha \
  --prompts-file <prompt-file> \
  --quantized-torch-ckpt-save-path <out>/ckpt \
  --hf-ckpt-dir <out>/hf

For current ModelOpt diffusion examples, use --format fp4 for NVFP4 exports. Do not assume the checked-out ModelOpt version accepts a literal nvfp4 format string unless you verified it locally.

For multi-transformer models:

  • quantize each backbone deliberately
  • keep each output directory separate
  • save both backbone.pt and the matching hf/<component> export
Show full SKILL.md (868 more words)Show less
3. Convert FP8 Exports For SGLang

FP8 requires an extra conversion step:

bash
PYTHONPATH=python python3 -m sglang.multimodal_gen.tools.build_modelopt_fp8_transformer \
  --modelopt-hf-dir <out>/hf \
  --modelopt-backbone-ckpt <out>/ckpt/backbone.pt \
  --base-transformer-dir <base-model-transformer-dir> \
  --output-dir <out>/sglang_transformer \
  --overwrite

What the converter does:

  • reads weight_quantizer._amax and input_quantizer._amax from backbone.pt
  • writes weight_scale and input_scale
  • materializes eligible FP8 weights as float8_e4m3fn
  • preserves ModelOpt ignore layers as BF16
  • strips stale _quantizer.* tensors and fallback-layer scales that should not survive into the SGLang-native checkpoint

For FLUX.1-dev, the validated fallback set currently keeps these modules in BF16:

  • transformer_blocks.*.norm1.linear
  • transformer_blocks.*.norm1_context.linear
  • transformer_blocks.*.ff.net.0.proj
  • transformer_blocks.*.ff.net.2
  • transformer_blocks.*.ff_context.net.0.proj
  • transformer_blocks.*.ff_context.net.2
  • single_transformer_blocks.*.norm.linear
  • single_transformer_blocks.*.proj_mlp

Use --model-type flux1 to force that profile, or rely on --model-type auto when the export config identifies FluxTransformer2DModel.

HunyuanVideo uses HunyuanVideoTransformer3DModel, so the validated HunyuanVideo FP8 fallback preset keeps these modules in BF16:

  • context_embedder.*
  • x_embedder.proj
  • time_text_embed.(timestep_embedder|guidance_embedder|text_embedder).linear_[12]
  • norm_out.linear
  • proj_out
  • transformer_blocks.*.norm1.linear
  • transformer_blocks.*.norm1_context.linear
  • single_transformer_blocks.*.norm.linear

Use --model-type hunyuan-video to force that profile, or rely on --model-type auto when the export config identifies HunyuanVideoTransformer3DModel.

HunyuanVideo ModelOpt exports use diffusers module names that differ from SGLang runtime names for fused QKV and fused QKV+MLP layers. Keep the diffusers-to-runtime mapping in build_modelopt_fp8_transformer.py in sync with runtime/models/dits/hunyuanvideo.py before trusting converted scale tensors.

Qwen Image and Qwen Image Edit share QwenImageTransformer2DModel, so one ModelOpt FP8 fallback preset covers both. The validated Qwen Image fallback set keeps these modules in BF16:

  • img_in
  • txt_in
  • time_text_embed.timestep_embedder.linear_1
  • time_text_embed.timestep_embedder.linear_2
  • norm_out.linear
  • proj_out
  • transformer_blocks.*.img_mlp.net.2
  • transformer_blocks.*.img_mod
  • transformer_blocks.*.txt_mod

Use --model-type qwen-image to force that profile, or rely on --model-type auto when the export config identifies QwenImageTransformer2DModel.

Qwen modulation weights can appear in safetensors as .img_mod.1.weight and .txt_mod.1.weight. Canonicalize those module names to .img_mod and .txt_mod before fallback matching.

For Qwen Image FP8, explicit BF16 fallback tensors must be written before honoring ModelOpt ignored weights. Otherwise converter stats can report a fallback while the output checkpoint still retains the source FP8 tensor, which causes severe image-quality regressions.

For FLUX.1-dev NVFP4 model families that need a mixed BF16+NVFP4 checkpoint, build the merged transformer explicitly:

bash
PYTHONPATH=python python3 -m sglang.multimodal_gen.tools.build_modelopt_nvfp4_transformer \
  --base-transformer-dir <base-model-transformer-dir> \
  --modelopt-hf-dir <out>/hf/transformer \
  --output-dir <out>/transformer-mixed \
  --pattern-preset flux1-nvfp4

The validated FLUX.1-dev mixed builder also needs to preserve:

  • quant_type: NVFP4 in config.json
  • swap_weight_nibbles: false for the validated diffusers export
4. Load The Quantized Checkpoint In SGLang

Single-transformer example:

bash
sglang generate \
  --model-path <base-model> \
  --transformer-path <quantized-transformer> \
  --prompt "<prompt>" \
  --seed <seed> \
  --save-output

Multi-transformer example:

bash
sglang generate \
  --model-path <base-model> \
  --transformer-path <quantized-transformer> \
  --transformer-2-path <another-transformer-or-bf16-override> \
  --prompt "<prompt>" \
  --seed <seed> \
  --save-output

Full ModelOpt Diffusers repo example (current Qwen Image NVFP4 path):

bash
sglang generate \
  --model-path lmsys/qwen-image-2512-modelopt-nvfp4-sglang \
  --prompt "<prompt>" \
  --seed <seed> \
  --save-output

Guideline:

  • use the global --transformer-path only when the model effectively has one transformer override to apply
  • use per-component overrides when different backbones need different checkpoints
  • use --model-path directly for published full ModelOpt Diffusers repos such as the Qwen Image NVFP4 family; this is different from a transformer-only override
  • the preferred CLI form is --<component>-path
  • config-expanded forms such as --component_paths.transformer_2=... also resolve to the same internal override map
5. Validate Accuracy

Use two levels of validation.

Reduced deterministic validation:

  • keep prompt, seed, resolution, and step count fixed
  • compare BF16 and quantized runs
  • capture denoising trajectories
  • inspect per-step latent cosine similarity plus MAE or RMSE
  • compare final frames with image metrics such as PSNR or MAE

Tool:

bash
PYTHONPATH=python python3 -m sglang.multimodal_gen.tools.compare_diffusion_trajectory_similarity \
  --model-path <base-model> \
  --model-id <optional-native-model-id> \
  --prompt "<prompt>" \
  --width <w> \
  --height <h> \
  --num-inference-steps <steps> \
  --guidance-scale <cfg> \
  --seed <seed> \
  --candidate-transformer-path <quantized-transformer> \
  --output-json <report.json>

Use --model-id FLUX.1-dev when --model-path points to a local directory but the runtime still needs the native FLUX.1 model registration.

Full-output validation:

  • run the same user-facing generation config in BF16 and quantized mode
  • inspect the output visually
  • only claim "quality preserved" for the exact scope you actually checked
6. Benchmark Correctly

Benchmark only when these match between BF16 and quantized:

  • prompt
  • seed
  • width and height
  • frame count
  • inference step count
  • GPU count and topology
  • offload flags
  • compile settings
  • profiler settings

Only the quantized checkpoint path should differ.

Interpretation rule:

  • the primary expected gain is in denoising
  • text-encoding and VAE differences are secondary and should not be over-attributed unless they were quantized too
7. Add Model-Specific Fallbacks Only When Needed

If the generic FP8 path fails on a new model family:

  • inspect which modules are numerically sensitive or loader-incompatible
  • keep fallback patterns small and explicit
  • isolate them in the converter instead of scattering ad-hoc exceptions
  • re-run deterministic trajectory checks after every fallback change

Do not turn one validated model quirk into a generic rule unless another family also needs it.

FP8 Offload Constraint

Current diffusion ModelOpt FP8 support requires:

  • dit_cpu_offload=false
  • dit_layerwise_offload may be enabled when you want lower DiT residency

Reason:

  • the FP8 linear path depends on a CUTLASS-compatible weight layout after loading
  • dit_cpu_offload is still treated conservatively
  • the fixed layerwise offload path now preserves non-contiguous tensor strides instead of flattening and rebuilding FP8 weights into a contiguous layout

Runtime behavior:

  • SGLang still force-disables dit_cpu_offload when it detects modelopt_fp8
  • benchmark commands should still pin offload flags explicitly so the command line itself makes the comparison rule obvious

Claim Discipline

When documenting results:

  • claim only scopes that were actually validated end to end
  • do not collapse "single-transformer FP8 override" into "full-model FP8"
  • do not call a practical deployment comparison a benchmark if BF16 and quantized commands used different offload behavior

Current Code Areas

FileRole
runtime/layers/quantization/__init__.pyregisters diffusion quant methods
runtime/layers/quantization/modelopt_fp8.pystatic per-tensor ModelOpt FP8 path used by flat quant_method=modelopt exports
runtime/layers/quantization/modelopt_quant.pyModelOpt FP8 and NVFP4 runtime loading
runtime/utils/quantization_utils.pyresolves flat ModelOpt configs and reconstructs NVFP4 config from metadata
runtime/loader/transformer_load_utils.pyguards incompatible FP8 offload modes
runtime/models/dits/flux_2.pypacked-QKV handling for the packed FLUX.2 NVFP4 family
tools/build_modelopt_fp8_transformer.pyBuild an SGLang-loadable FP8 transformer from a ModelOpt export
tools/build_modelopt_nvfp4_transformer.pyBuild mixed BF16+NVFP4 transformer directories when a family needs preserved BF16 layers
tools/compare_diffusion_trajectory_similarity.pyreduced deterministic BF16-vs-quantized validation
docs/docs/sglang-diffusion/quantization.mdxpublic ModelOpt support matrix and CLI examples
python/sglang/multimodal_gen/test/server/testcase_configs.pyreusable ModelOpt testcase constants, thresholds, and helpers
python/sglang/multimodal_gen/test/server/gpu_cases.pyconcrete GPU and B200 ModelOpt CI case lists

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-modelopt-quant of sgl-project/sglang.

Open the folder on GitHubat commit b7b2975

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sglang Diffusion Modelopt Quant next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sglang Diffusion Modelopt Quant compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sglang Diffusion Modelopt Quant this skillsgl-project/sglang37k2 repos~5kAutomated safety check: PassApache-2.0
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Make Op VerifyCVCUDA/CV-CUDA2.7k—~433Automated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
One EvalOpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Make Op Verify

    CVCUDA/CV-CUDA

    Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).

    2.7k GitHub stars~433 tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • One Eval

    OpenDCAI/One-Eval

    驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

    165 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 31 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Questions about Sglang Diffusion Modelopt Quant

What does Sglang Diffusion Modelopt Quant do?

A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion. Sglang Diffusion Modelopt Quant is an agent skill from sgl-project/sglang. Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

When should I use Sglang Diffusion Modelopt Quant?

Sglang Diffusion Modelopt Quant fits situations like: quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8; NVFP4 checkpoint loadable; benchmarkable in SGLang Diffusion.

How do I install Sglang Diffusion Modelopt Quant in Claude Code?

Run `npx skills add sgl-project/sglang --skill sglang-diffusion-modelopt-quant -a claude-code`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-modelopt-quant in sgl-project/sglang) into .claude/skills/sglang-diffusion-modelopt-quant in your project. Claude Code loads it when a task matches its description.

How do I install Sglang Diffusion Modelopt Quant in Codex?

Run `npx skills add sgl-project/sglang --skill sglang-diffusion-modelopt-quant -a codex`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-modelopt-quant in sgl-project/sglang) into .agents/skills/sglang-diffusion-modelopt-quant in your project. Codex loads it when a task matches its description.

Can I use Sglang Diffusion Modelopt Quant in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-diffusion-modelopt-quant -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-diffusion-modelopt-quant, .gemini/skills/sglang-diffusion-modelopt-quant, .github/skills/sglang-diffusion-modelopt-quant and .opencode/skills/sglang-diffusion-modelopt-quant in your project.

What does Sglang Diffusion Modelopt Quant need to run?

Going by SKILL.md and its folder, Sglang Diffusion Modelopt Quant needs the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.

Does Sglang Diffusion Modelopt Quant access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sglang Diffusion Modelopt Quant safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sglang Diffusion Modelopt Quant use?

Sglang Diffusion Modelopt Quant is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sglang Diffusion Modelopt Quant use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sglang Diffusion Modelopt Quant?

Skills that share tags, products or a category with Sglang Diffusion Modelopt Quant: SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Make Op Verify (CVCUDA/CV-CUDA, 2.7k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sglang Diffusion Modelopt Quant?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,851 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 8, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.