Agent skill

Sglang Diffusion Benchmark Profile

by sgl-project in sgl-project/sglang

A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Sglang Diffusion Benchmark Profile

skills CLI
$ npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang sglang-diffusion-benchmark-profile --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile .claude/skills/sglang-diffusion-benchmark-profile && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sglang-diffusion-benchmark-profile
GitHub stars
37k
Used in
2 other repos
Token cost
~2.4k tokens
SKILL.md length
1,186 words
Files
5 (incl. scripts)
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

  • Benchmarking denoise latency
  • SKILL.md covers Preflight, Native Backend Gate, Main Reference and Opportunity Discovery Rule
  • Runs Python scripts from its folder; needs HF_TOKEN
  • Profiling a diffusion bottleneck in SGLang

What it does

Sglang Diffusion Benchmark Profile is an agent skill from sgl-project/sglang. Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `benchmark-and-profile.md`, `existing-fast-paths.md` and `scripts/bench_diffusion_denoise.py`).

It sits in AI & LLM Engineering, covering Performance optimization. It works with SGLang. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Benchmarking denoise latency
  • Profiling a diffusion bottleneck in SGLang

Example prompts

  • “/sglang-diffusion-benchmark-profile”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sglang Diffusion Benchmark Profile loads about 2.4k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 1,186 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 1,186 words, ~2,359 tokens.

Download SKILL.mdSave it as .claude/skills/sglang-diffusion-benchmark-profile/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
sglang-diffusion-benchmark-profile
description
Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

SGLang Diffusion Benchmark and Profile

Use this skill when measuring denoise performance, finding the slow op, checking whether an existing fast path can solve it, or verifying that a hotspot is real before any kernel work in sglang.multimodal_gen.

This skill is diagnosis-first. It owns:

  • checked-in denoise benchmark presets
  • same-GPU quality/BCG applicability checks with repeated exact, lossless, and high rows
  • perf dump collection and before/after comparison
  • torch.profiler trace capture and quick hotspot ranking
  • mapping hot kernels back to known fast paths and fusion families
  • packaging confirmed kernel work with enough evidence for the appropriate kernel, Nsight, or framework-specific optimization workflow

This skill does not own low-level kernel authoring or standalone Nsight workflows.

Preflight

Before running any benchmark, profiler, or kernel-validation command:

  • use scripts/diffusion_skill_env.py to derive the repo root from sglang.__file__
  • verify the repo is writable
  • export HF_TOKEN before using gated Hugging Face models such as black-forest-labs/FLUX.*
  • export FLASHINFER_DISABLE_VERSION_CHECK=1
  • set SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 when comparing stage-level denoise/decode timings; the preset helper sets it by default unless the caller explicitly overrides it
  • for downloaded checkpoints, use the preset helper's task-owned --model-cache-root together with --cleanup-model-cache; verify the JSONL ledger reports zero residual weight files before moving to the next model
  • choose idle GPU(s) before starting perf work; for a comparison matrix, hold the same GPU set and verify it has no foreign process at every run boundary

Native Backend Gate

All diffusion benchmark and profiling results owned by this skill must come from the native SGLang diffusion backend.

Treat any of the following as a hard stop condition:

  • Falling back to diffusers backend
  • Using diffusers backend
  • Loaded diffusers pipeline

If any benchmark, perf-dump, or torch.profiler command prints one of those signals:

  • stop the workflow immediately
  • do not keep the generated numbers or traces as SGLang benchmark evidence
  • do not continue to hotspot classification or kernel work
  • first fix model resolution, pipeline selection, overlay/materialization, or other backend-selection issues so the model runs on the native SGLang diffusion path

Main Reference

  • benchmark-and-profile.md — canonical denoise benchmark, perf dump, and torch.profiler workflow; uses checked-in nightly-aligned presets plus current-source extras such as LongCat image/edit, Qwen base edit/layered, SD3.5, SANA-Video/SANA-WM, LingBot Video/World, Cosmos3 Edge/Super I2V/distilled and the explicit Super TP2 x CFG2 comparator, LTX-2.5 and its diffusion decoder, MiniMax-H3, FLUX.2 Klein, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.1/2.2, the Blackwell-only Wan2.2 NVFP4 comparator, LTX-2.3, HunyuanVideo, MOVA, Helios, image edit, Hunyuan3D shape, and a separate Pi0.5 action-policy lane
  • existing-fast-paths.md — map bottlenecks to existing fused kernels, MoE routing, packed QKV paths, fused QK norm + RoPE, distributed overlap patterns, and open optimization PRs before proposing new code
  • scripts/diffusion_skill_env.py — preflight helper: repo root discovery from the skill's owning checkout before falling back to sglang.__file__, write-access probe, benchmark/profile output directories, idle GPU selection
  • scripts/bench_diffusion_denoise.py — end-to-end denoise benchmark preset runner via sglang generate; defaults to eager/lossless, supports explicit quality and BCG comparators plus a same-GPU applicability matrix, rejects invalid BCG capture/fallback logs and late high-quality DiT fusion mounts, forces H3 to its eager consistency mode, enables synchronized stage attribution, validates nightly preset drift, and can clean one isolated model cache after the full matrix in a finally block with a JSONL ledger

Opportunity Discovery Rule

Before calling a diffusion hotspot "new", first classify it with existing-fast-paths.md.

Always rule out these existing families first:

  • HunyuanVideo VAE GroupNorm+SiLU
  • LTX upsampler GroupNorm+SiLU
  • Z-Image bf16-native Triton RMSNorm scale/tanh-residual modulation
  • SANA packed self-attention Q/K/V and cross-attention K/V GEMMs
  • SANA-Video's packed projections and request-scoped BF16-input linear attention at quality=lossless or quality=high; keep the second attention GEMM in FP32 and compare against quality=exact before changing its precision further
  • SANA-Video reuse of SANA's bit-exact bias/activation, residual-gate, and LayerNorm-modulation fast paths before adding video-only kernels
  • MiniMax-H3 indexed modulation, fused QK norm + RoPE, packed Ulysses QKV, USP relayout, and batched TP AdaLN collectives
  • bit-exact diffusion adaLN modulation and fused LayerNorm + modulation for FLUX.1, GLM-Image, and SANA
  • request-scoped DiT and VAE fast paths at quality=lossless or quality=high
  • LingBot Video's default-on fused group-limited top-k expert selection before treating its router's topk/mask/gather chain as a new hotspot
  • Wan causal-VAE cache/padding and DupUp3D data-movement fusions
  • fused diffusion QK norm + RoPE
  • LTX2 split RoPE
  • LTX2 residual-gate add
  • LTX-2.5 diffusion-decoder NATTEN selection before interpreting a FlexAttention fallback trace
  • varlen USP attention pack/scatter
  • NVFP4 / Nunchaku packed QKV
  • Nunchaku fused GELU MLP
  • Ulysses / USP attention overlap
  • turbo-layer async all-to-all overlap
  • torch.compile compute / communication reorder
  • breakable CUDA graph capture for supported fixed-resolution pipelines
  • dual-stream diffusion execution
Show full SKILL.md (469 more words)Show less

The checked-in helper defaults to eager. Use --torch-compile only for a controlled comparator, never for the eager ground truth. The legacy --no-torch-compile spelling remains accepted but is redundant.

For kernel/BCG discovery, run --quality-bcg-matrix. It executes Eager/BCG as A-B-B-A at exact, then repeats the pair at lossless and high, on one locked GPU set and one isolated checkpoint cache. The lossless/high+BCG rows are applicability checks, not presumed-valid performance cells. A BCG row is invalid unless the log contains [Diffusion BCG] captured and contains no support-disable, capture-failure, serving-signature-miss, or late quality-fusion marker. In particular, a request-scoped DiT fusion mounted after exact warmup capture would be bypassed by replay; reject that row even when capture and signature checks pass. For video presets, the helper declares both the request resolution and --warmup-num-frames so the synthetic BCG warmup captures the requested temporal shape. Treat any remaining temporal or conditioning signature miss as Eager fallback, not as a valid BCG measurement.

A zero process exit is not sufficient evidence: every accepted row must also contain its requested perf dump and a generated image, video, audio, or 3D mesh file. The helper gives every cell a unique output name and rejects missing artifacts.

On machines with a read-only Hugging Face cache, combine --model-cache-root <task-owned-dir> with one or more --seed-model-cache-root <read-only-HF-home-or-hub> options. The helper exposes cached repos through a task-owned copy-on-write directory overlay, downloads misses only into the isolated cache, and removes links plus new downloads in its normal cleanup finally block without modifying the seed cache.

Keep prompt, negative prompt, seed, shape, steps, guidance, dtype, topology, and residency fixed. Lossless comparisons require byte-identical artifacts. For quality=lossless and quality=high, report aggregate and worst-frame SSIM/PSNR; the repository defaults are 0.95/28 dB for images and 0.92/24 dB for video unless the model's checked-in consistency metadata defines a different threshold. A performance PR needs repeated saved-request e2e improvement of at least 1.5%, a representative profile, and before/after image or video evidence.

MiniMax-H3 is always an eager consistency case on current main. Use --model minimax-h3-t2va; its preset writes the H3 request fields through a generated config and suppresses the helper's global compile default. Do not turn the model's nominal BCG support gate into a performance claim: prompt- dependent packed-sequence host boundaries can differ between warmup and the serving request. A valid H3 BCG experiment must prove that every captured segment replays, keeps the MP4 byte-identical, and does not trade latency for the extra graph memory.

For FLUX-family manual profiling runs with a quantized transformer override:

  • use sglang generate directly
  • pass the override as --transformer-path <dir>
  • prefer --prompt-path <file> when also fixing --output-file-name
  • if the base model is already cached locally and the machine has unreliable HF access, use the local cached --model-path plus HF_HUB_OFFLINE=1
  • remember that --profile changes latency substantially; use the non-profile perf dump for the real before/after benchmark claim

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile of sgl-project/sglang.

  • SKILL.md
  • benchmark-and-profile.md
  • existing-fast-paths.md
  • scripts/bench_diffusion_denoise.py
  • scripts/diffusion_skill_env.py

Open the folder on GitHubat commit f620d73

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sglang Diffusion Benchmark Profile next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sglang Diffusion Benchmark Profile compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sglang Diffusion Benchmark Profile this skillsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~3.9kAutomated safety check: PassNone
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Serving Framework BenchmarkBBuf/AI-Infra-Auto-Driven-SKILLS925—~7.5kAutomated safety check: PassNone
Magpie Kernel Evaluatoramd/skills406—~2.3kAutomated safety check: PassMIT

Similar skills

  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    925 GitHub stars~3.9k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Serving Framework Benchmark

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

    925 GitHub stars~7.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    406 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Rl Msprobe

    ascend-ai-coding/awesome-ascend-skills

    自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +…

    174 GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Works with

Questions about Sglang Diffusion Benchmark Profile

What does Sglang Diffusion Benchmark Profile do?

A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang. Sglang Diffusion Benchmark Profile is an agent skill from sgl-project/sglang. Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

When should I use Sglang Diffusion Benchmark Profile?

Sglang Diffusion Benchmark Profile fits situations like: benchmarking denoise latency; profiling a diffusion bottleneck in SGLang.

How do I install Sglang Diffusion Benchmark Profile in Claude Code?

Run `npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a claude-code`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile in sgl-project/sglang) into .claude/skills/sglang-diffusion-benchmark-profile in your project. Claude Code loads it when a task matches its description.

How do I install Sglang Diffusion Benchmark Profile in Codex?

Run `npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a codex`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile in sgl-project/sglang) into .agents/skills/sglang-diffusion-benchmark-profile in your project. Codex loads it when a task matches its description.

Can I use Sglang Diffusion Benchmark Profile in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-diffusion-benchmark-profile, .gemini/skills/sglang-diffusion-benchmark-profile, .github/skills/sglang-diffusion-benchmark-profile and .opencode/skills/sglang-diffusion-benchmark-profile in your project.

What does Sglang Diffusion Benchmark Profile need to run?

Going by SKILL.md and its folder, Sglang Diffusion Benchmark Profile needs Python for the scripts in its folder and credentials named HF_TOKEN. Our summary lists: Python 3.

Does Sglang Diffusion Benchmark Profile access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sglang Diffusion Benchmark Profile safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sglang Diffusion Benchmark Profile use?

Sglang Diffusion Benchmark Profile is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sglang Diffusion Benchmark Profile use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sglang Diffusion Benchmark Profile?

Skills that share tags, products or a category with Sglang Diffusion Benchmark Profile: LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Serving Framework Benchmark (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sglang Diffusion Benchmark Profile?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.