LLM Torch Profiler Trace Analysis
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-benchmark-profile --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile .claude/skills/sglang-diffusion-benchmark-profile && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "sglang-diffusion-benchmark-profile" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile into .claude/skills/sglang-diffusion-benchmark-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-benchmark-profile", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profileType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-benchmark-profile --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile .agents/skills/sglang-diffusion-benchmark-profile && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "sglang-diffusion-benchmark-profile" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile into .agents/skills/sglang-diffusion-benchmark-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-benchmark-profile", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-benchmark-profile --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile .cursor/skills/sglang-diffusion-benchmark-profile && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "sglang-diffusion-benchmark-profile" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile into .cursor/skills/sglang-diffusion-benchmark-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-benchmark-profile", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-benchmark-profile --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile .gemini/skills/sglang-diffusion-benchmark-profile && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "sglang-diffusion-benchmark-profile" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile into .gemini/skills/sglang-diffusion-benchmark-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-benchmark-profile", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang sglang-diffusion-benchmark-profileInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile .github/skills/sglang-diffusion-benchmark-profile && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "sglang-diffusion-benchmark-profile" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile into .github/skills/sglang-diffusion-benchmark-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-benchmark-profile", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang sglang-diffusion-benchmark-profile --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile .opencode/skills/sglang-diffusion-benchmark-profile && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "sglang-diffusion-benchmark-profile" agent skill from https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile into .opencode/skills/sglang-diffusion-benchmark-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sglang-diffusion-benchmark-profile", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
sglang-diffusion-benchmark-profileA skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.
Sglang Diffusion Benchmark Profile is an agent skill from sgl-project/sglang. Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `benchmark-and-profile.md`, `existing-fast-paths.md` and `scripts/bench_diffusion_denoise.py`).
It sits in AI & LLM Engineering, covering Performance optimization. It works with SGLang. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Sglang Diffusion Benchmark Profile loads about 2.4k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 1,186 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 1,186 words, ~2,359 tokens.
.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Use this skill when measuring denoise performance, finding the slow op, checking whether an existing fast path can solve it, or verifying that a hotspot is real before any kernel work in sglang.multimodal_gen.
This skill is diagnosis-first. It owns:
torch.profiler trace capture and quick hotspot rankingThis skill does not own low-level kernel authoring or standalone Nsight workflows.
Before running any benchmark, profiler, or kernel-validation command:
scripts/diffusion_skill_env.py to derive the repo root from sglang.__file__HF_TOKEN before using gated Hugging Face models such as black-forest-labs/FLUX.*FLASHINFER_DISABLE_VERSION_CHECK=1SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 when comparing stage-level
denoise/decode timings; the preset helper sets it by default unless the
caller explicitly overrides it--model-cache-root together with --cleanup-model-cache; verify the JSONL
ledger reports zero residual weight files before moving to the next modelAll diffusion benchmark and profiling results owned by this skill must come from the native SGLang diffusion backend.
Treat any of the following as a hard stop condition:
Falling back to diffusers backendUsing diffusers backendLoaded diffusers pipelineIf any benchmark, perf-dump, or torch.profiler command prints one of those signals:
torch.profiler workflow; uses checked-in nightly-aligned presets plus current-source extras such as LongCat image/edit, Qwen base edit/layered, SD3.5, SANA-Video/SANA-WM, LingBot Video/World, Cosmos3 Edge/Super I2V/distilled and the explicit Super TP2 x CFG2 comparator, LTX-2.5 and its diffusion decoder, MiniMax-H3, FLUX.2 Klein, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.1/2.2, the Blackwell-only Wan2.2 NVFP4 comparator, LTX-2.3, HunyuanVideo, MOVA, Helios, image edit, Hunyuan3D shape, and a separate Pi0.5 action-policy laneQK norm + RoPE, distributed overlap patterns, and open optimization PRs before proposing new codesglang.__file__, write-access probe, benchmark/profile output directories, idle GPU selectionsglang generate; defaults to eager/lossless, supports explicit quality and BCG comparators plus a same-GPU applicability matrix, rejects invalid BCG capture/fallback logs and late high-quality DiT fusion mounts, forces H3 to its eager consistency mode, enables synchronized stage attribution, validates nightly preset drift, and can clean one isolated model cache after the full matrix in a finally block with a JSONL ledgerBefore calling a diffusion hotspot "new", first classify it with existing-fast-paths.md.
Always rule out these existing families first:
quality=lossless or quality=high; keep the second attention GEMM in FP32 and
compare against quality=exact before changing its precision furtherquality=lossless or quality=hightopk/mask/gather chain as a new hotspotQK norm + RoPEtorch.compile compute / communication reorderThe checked-in helper defaults to eager. Use --torch-compile only for a
controlled comparator, never for the eager ground truth. The legacy
--no-torch-compile spelling remains accepted but is redundant.
For kernel/BCG discovery, run --quality-bcg-matrix. It executes Eager/BCG as
A-B-B-A at exact, then repeats the pair at lossless and high, on
one locked GPU set and one isolated checkpoint cache. The lossless/high+BCG
rows are applicability checks, not presumed-valid performance cells. A BCG row is invalid unless the log
contains [Diffusion BCG] captured and contains no support-disable,
capture-failure, serving-signature-miss, or late quality-fusion marker. In
particular, a request-scoped DiT fusion mounted after exact warmup capture
would be bypassed by replay; reject that row even when capture and signature
checks pass. For video presets, the helper declares both the request resolution
and --warmup-num-frames so the synthetic BCG warmup captures the requested
temporal shape. Treat any remaining temporal or conditioning signature miss as
Eager fallback, not as a valid BCG measurement.
A zero process exit is not sufficient evidence: every accepted row must also contain its requested perf dump and a generated image, video, audio, or 3D mesh file. The helper gives every cell a unique output name and rejects missing artifacts.
On machines with a read-only Hugging Face cache, combine
--model-cache-root <task-owned-dir> with one or more
--seed-model-cache-root <read-only-HF-home-or-hub> options. The helper exposes
cached repos through a task-owned copy-on-write directory overlay, downloads
misses only into the isolated cache, and removes links plus new downloads in
its normal cleanup finally block without modifying the seed cache.
Keep prompt, negative prompt, seed, shape, steps, guidance, dtype, topology,
and residency fixed. Lossless comparisons require byte-identical artifacts.
For quality=lossless and quality=high, report aggregate and worst-frame SSIM/PSNR; the repository
defaults are 0.95/28 dB for images and 0.92/24 dB for video unless the model's
checked-in consistency metadata defines a different threshold. A performance
PR needs repeated saved-request e2e improvement of at least 1.5%, a
representative profile, and before/after image or video evidence.
MiniMax-H3 is always an eager consistency case on current main. Use
--model minimax-h3-t2va; its preset writes the H3 request fields through a
generated config and suppresses the helper's global compile default. Do not
turn the model's nominal BCG support gate into a performance claim: prompt-
dependent packed-sequence host boundaries can differ between warmup and the
serving request. A valid H3 BCG experiment must prove that every captured
segment replays, keeps the MP4 byte-identical, and does not trade latency for
the extra graph memory.
For FLUX-family manual profiling runs with a quantized transformer override:
sglang generate directly--transformer-path <dir>--prompt-path <file> when also fixing --output-file-name--model-path plus HF_HUB_OFFLINE=1--profile changes latency substantially; use the non-profile perf dump for the real before/after benchmark claim© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts) in python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile of sgl-project/sglang.
Open the folder on GitHubat commit f620d73
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
Sglang Diffusion Benchmark Profile next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Sglang Diffusion Benchmark Profile this skillsgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~2.8k | Automated safety check: Pass | None | |
| LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~3.9k | Automated safety check: Pass | None | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Serving Framework BenchmarkBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~7.5k | Automated safety check: Pass | None | |
| Magpie Kernel Evaluatoramd/skills | 406 | — | ~2.3k | Automated safety check: Pass | MIT |
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
BBuf/AI-Infra-Auto-Driven-SKILLS
Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
ascend-ai-coding/awesome-ascend-skills
自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +…
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
Works with
Categories
A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang. Sglang Diffusion Benchmark Profile is an agent skill from sgl-project/sglang. Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.
Sglang Diffusion Benchmark Profile fits situations like: benchmarking denoise latency; profiling a diffusion bottleneck in SGLang.
Run `npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a claude-code`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile in sgl-project/sglang) into .claude/skills/sglang-diffusion-benchmark-profile in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a codex`. Or copy the skill folder (python/sglang/multimodal_gen/.agents/skills/sglang-diffusion-benchmark-profile in sgl-project/sglang) into .agents/skills/sglang-diffusion-benchmark-profile in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-diffusion-benchmark-profile -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-diffusion-benchmark-profile, .gemini/skills/sglang-diffusion-benchmark-profile, .github/skills/sglang-diffusion-benchmark-profile and .opencode/skills/sglang-diffusion-benchmark-profile in your project.
Going by SKILL.md and its folder, Sglang Diffusion Benchmark Profile needs Python for the scripts in its folder and credentials named HF_TOKEN. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Sglang Diffusion Benchmark Profile is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Sglang Diffusion Benchmark Profile: LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Serving Framework Benchmark (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.