Agent skill

Diffusion Perf Opt

by vllm-project in vllm-project/vllm-omni

Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Diffusion Perf Opt

skills CLI
$ npx skills add vllm-project/vllm-omni --skill diffusion-perf-opt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-omni diffusion-perf-opt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/diffusion-perf-opt .claude/skills/diffusion-perf-opt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
diffusion-perf-opt
GitHub stars
7.1k
Token cost
~7.5k tokens
SKILL.md length
3,628 words
Files
4 (incl. scripts, references)
Skills in repo
20
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

  • Works in 9 steps: Freeze the measurement protocol and… → Collect a real baseline. → Model the parallel strategy before… → …
  • Codex is asked to analyze profiling traces
  • SKILL.md covers First Questions, Workflow, Priority Rules and Optimization Layers, plus 2 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Diffusion Perf Opt is an agent skill from vllm-project/vllm-omni. Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. Use when Codex is asked to analyze profiling traces, choose parallel strategies, inspect torch profiler trace.json or trace.json.gz timelines, estimate optimization ROI, investigate GPU idle/free bubbles, compare USP/CFG/HSDP/VAE parallelism, or design operator/host/quantization optimizations for vLLM Omni.

Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/optimization-playbook.md` and `scripts/trace_analyzer.py`).

It sits in AI & LLM Engineering, covering LLM inference and serving, Performance optimization and AI video generation. It works with vLLM and Qwen. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.

When your agent uses it

  • Codex is asked to analyze profiling traces
  • Choose parallel strategies
  • Inspect torch profiler trace.json
  • Trace.json.gz timelines

Example prompts

  • “/diffusion-perf-opt”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Freeze the measurement protocol and commands.
  2. Collect a real baseline.
  3. Model the parallel strategy before testing.
  4. Search for the best parallel configuration.
  5. Run targeted A/B tests.
  6. Enforce the quality and precision gate for every optimization.
  7. Collect diagnostic trace only after narrowing hypotheses.
  8. Analyze host, communication, and operators.
  9. Produce an optimization plan.

What it can do on your machine

Read from SKILL.md and the folder at commit c548a11. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Diffusion Perf Opt loads about 7.5k tokens when it runs, and up to ~8.4k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 3,628 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~7.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from vllm-project/vllm-omni at commit c548a11, republished under its Apache-2.0 licence (© vllm-project). 3,628 words, ~7,504 tokens.

Download SKILL.mdSave it as .claude/skills/diffusion-perf-opt/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
diffusion-perf-opt
description
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. Use when Codex is asked to analyze profiling traces, choose parallel strategies, inspect torch profiler trace.json or trace.json.gz timelines, estimate optimization ROI, investigate GPU idle/free bubbles, compare USP/CFG/HSDP/VAE parallelism, or design operator/host/quantization optimizations for vLLM Omni.

vLLM Omni Performance Optimization

Use this skill to run a disciplined optimization loop for vLLM Omni diffusion workloads. Keep two ideas separate: real performance baselines are collected with low overhead, while torch profiler traces are diagnostic artifacts and may distort latency.

First Questions

Before proposing changes, ask for the optimization scene if it is not already known:

  • GPU model, card count, topology, and whether NVLink is present.
  • Model and pipeline, for example Wan2.2 I2V A14B.
  • User workload: resolution, frames, steps, batch/concurrency, CFG scales, prompt/image inputs.
  • Three runnable commands for each model/workload, or scripts that generate them:
    • Server startup command: exact vllm serve command, environment variables, model path, port, parallelism flags, profiler flags, and precision/compile settings.
    • User request command: exact single-request client command (for example curl or Python client) with fixed prompt/media/seed/size/steps. Use this to validate correctness and collect per-request stage timings.
    • User benchmark command: exact repeatable benchmark command or script with warmup count, measured iteration count, concurrency/batch policy, output directory, and summary format.
  • Current enabled strategies: USP/SP, CFG parallel, HSDP/FSDP, VAE patch parallel, torch.compile, profiler options.
  • Optimization target: latency, throughput, memory, cost, or quality-preserving speed.
  • Precision/quality tolerance: bf16/fp8/quantization/sparsity/approximate attention allowed or not.
  • Quality validation method: exact-output comparison when feasible, or metrics such as SSIM, PSNR, LPIPS, cosine similarity, MAE, MSE, temporal flicker checks, baseline self-run variance, and reviewable output artifacts.

Workflow

  1. Freeze the measurement protocol and commands.

    • Before analyzing parallel strategies or traces, establish the three commands for the target model/workload: server startup, single user request, and user benchmark.
    • Prefer checked-in scripts or one-off shell scripts over free-form commands in chat. The scripts should make workload variables explicit, including model path, port/base URL, prompt/media inputs, resolution, frames, steps, seed, warmup, measured iterations, concurrency, and output path.
    • Measurement reports must preserve the concrete process, not just the final numbers. For each tested configuration, write down the server startup command, the user request command, the user benchmark command or polling loop, the final server response metrics, and the client-side timing/HTTP result.
    • Present all duration and latency values in milliseconds (ms) in tables and summaries. If an API returns seconds, convert to ms before reporting. Keep raw units only inside quoted raw logs or source snippets.
    • Keep profiler modes separate in the command set:
      • Baseline/benchmark commands should avoid torch profiler and stack collection.
      • Diagnostic commands may enable torch profiler after a baseline identifies a bottleneck.
    • If any of the three commands is missing, ask for it or create a proposed script before proceeding.
    • Make clear which command is authoritative for correctness validation and which command is authoritative for performance numbers.
  2. Collect a real baseline.

    • Disable PyTorch profiler and stack collection.
    • Disable or fix torch.compile state so A/B is fair.
    • Avoid --enforce-eager for production-speed baselines unless eager is the target.
    • Prefer --log-stats and --enable-diffusion-pipeline-profiler for low-overhead stage timing. PR 3069 has the relevant metrics/log-stats changes; if local code does not include them, fetch or cherry-pick the minimal metrics changes rather than merging unrelated PR drift.
    • Run warmup requests and exclude first-request lazy init.
  3. Model the parallel strategy before testing.

    • Estimate compute and communication for self-attention, cross-attention, FFN, CFG, VAE encode/decode, and HSDP.
    • Select a small candidate matrix rather than exhaustively testing everything.
    • Typical 2-card candidates: USP=2, CFG=2, USP=1/HSDP on-off, VAE parallel on-off if memory allows.
    • Typical 4-card candidates: USP=4, CFG=2 x USP=2, USP=2 x HSDP, VAE parallel on-off.
    • Typical 8-card candidates for official CFG-enabled video diffusion models:
      • Primary candidate: CFG=2 x USP=4 with VAE patch parallel across all 8 ranks.
      • Compare against: CFG=1 x USP=8 to test whether larger sequence parallel groups beat CFG branch parallelism.
      • Isolate VAE parallelism: keep the DiT strategy fixed and compare VAE patch parallel on/off or different VAE patch sizes.
      • If long-video diffuse remains dominant and Ulysses all-to-all is suspected, test a hybrid sequence strategy such as CFG=2 x USP=2 x Ring=2.
      • Test HSDP off only if the model fits without it; treat HSDP primarily as a memory strategy until A/B proves otherwise.
    • Prefer USP/SP for long video token sequences; prefer CFG parallel when CFG doubles transformer forwards and sequence length is modest.
    • Convert workload shape to latent/patch token counts before choosing candidates. For video models, record frames, latent frames, latent HxW, patch size, approximate token count, and which stages should scale with the token count.
  4. Search for the best parallel configuration.

    • Start with a small matrix that answers one question per comparison:
      • CFG parallel vs larger USP: compare CFG=2 x USP=world/2 against CFG=1 x USP=world for CFG-enabled workloads.
      • VAE patch parallel value: compare the best DiT strategy with VAE patch parallel enabled and disabled.
      • HSDP cost: compare HSDP on/off only when both configurations fit memory.
      • Ulysses vs Ring: test a Ring/Ulysses hybrid only after long-sequence diffuse is confirmed dominant.
    • For each configuration, create or record a stable config id, for example A_cfg2_usp4_vaepp8_hsdp_tiling.
    • For each config id, capture:
      • exact server command and environment variables,
      • observed distributed setup from logs, such as SP groups, CFG groups, HSDP shard/replicate sizes, VAE patch size,
      • exact request command for every scenario,
      • final server response metrics and client-side elapsed time,
      • output artifact paths and any failed/empty responses.
    • Use one warmup request per scenario, then at least three measured repeats for the shortlisted config. For exploratory matrix pruning, one measured request is acceptable only if the margin is large; label it as one-shot.
    • Compare configurations by stage, not only end-to-end latency. A configuration can improve diffuse while hurting vae.decode; record both effects.
    • Select the best config only after checking the target metric, dominant stages, memory headroom, and output correctness.
  5. Run targeted A/B tests.

    • Change one variable per test.
    • Keep model, input, seed, request parameters, GPU placement, and warmup policy fixed.
    • Record latency, stage timings, memory, output quality, and logs.
    • Report comparison tables in ms. Include at least end-to-end client time, server inference_time, server stage generation time if available, vae.encode, diffuse, vae.decode, and peak memory.
  6. Enforce the quality and precision gate for every optimization.

    • Do not mark an optimization ready until quality is checked on the same workload: model, prompt/media, seed, shape, steps, scheduler, dtype, backend, parallelism, and output encoding.
    • For math-preserving changes, require exact/near-exact agreement or show that differences are within baseline self-run variance.
    • For precision, quantization, approximate attention, sparse/custom kernels, or backend changes, include reviewable artifacts and metrics such as SSIM, PSNR, LPIPS, cosine similarity, MAE, MSE, and temporal flicker checks.
    • If quality tolerance is not stated, default to quality-preserving behavior. Failed or inconclusive quality validation blocks a ready-to-merge claim.
  7. Collect diagnostic trace only after narrowing hypotheses.

    • Use torch profiler for a small number of requests.
    • Run two separate diagnostic traces instead of mixing concerns:
      • Operator/shape trace: enable torch_profiler_record_shapes=True and keep stack collection disabled. Use this to rank CUDA kernels, NCCL collectives, attention/MLP/norm/RoPE work, and shape-specific hot operators.
      • Host-stack trace: enable torch_profiler_with_stack=True and normally keep shape collection disabled. Use this to map CPU/Python host gaps, synchronization points, scheduler paths, and request handling overhead.
    • Keep both trace commands and reports separate from baseline/benchmark commands. Torch profiler latency is diagnostic only and must not be used as the final latency claim.
    • Prefer profiling only the narrowed dominant scenario, for example the highest-resolution/video-length case where diffuse dominates.
    • If profiler endpoints are available, run one warmup request first, then call /start_profile, run one profiled request, and call /stop_profile. This keeps model initialization and warmup out of the diagnostic trace.
    • Start by analyzing rank 0 only. Expand to more ranks only if rank 0 suggests imbalance, unclear GPU idle/free bubbles, high NCCL wait, server timing mismatch, CFG branch imbalance, or USP group stragglers.
    • Diagnostic reports must be written to disk and preserve: server command, profiler config, warmup/request/polling commands, trace artifact paths, rank analyzed, analyzer output or summary, and the decision about whether additional ranks are necessary.
    • Analyze both rank-level balance and device-level free bubbles when additional ranks are opened.
  8. Analyze host, communication, and operators.

    • Find GPU idle/free intervals and map each large gap to the enclosing CPU/Python code.
    • Separate real GPU idle from profiler overhead such as CUPTI Command Buffer Full.
    • Compare NCCL kernel time to user annotations; annotations can overcount nested intervals.
    • Rank operator work by total CUDA time and by repeated small-kernel launch count.
  9. Produce an optimization plan.

    • Classify candidates as P0/P1/P2.
    • For each candidate, state necessity, expected benefit, implementation path, validation plan, and quality risk.
    • Include a concrete quality/precision validation gate for each candidate.
    • Do not implement high-risk operator rewrites before proving the operator is a bottleneck for the target shapes.
    • End the plan with a user-facing candidate selection table. The assistant should not automatically choose a risky optimization just because it is technically possible. Present the options clearly and let the user decide which item is worth implementing based on latency target, engineering budget, memory headroom, and quality tolerance.
    • Organize the plan by optimization layer, not by a flat list of ideas:
      • host/runtime optimization,
      • measurement/benchmark reliability,
      • parallelism and communication,
      • VAE encode/decode and media pre/post processing,
      • operator fusion and layout cleanup,
      • attention main-path optimization,
      • algorithmic, precision, or approximation changes.
    • For each layer, tie every candidate back to evidence from baseline metrics, diagnostic trace, source code, or output quality requirements. Do not list generic optimizations without a trace or workload reason.

Priority Rules

  • P0: low risk, likely useful, or required for trustworthy measurement. Examples: real baseline, warmup, targeted parallel A/B, disabling avoidable empty_cache, scheduler coefficient caching.
  • P1: meaningful code changes with contained risk. Examples: cross-attention KV caching, VAE gather/broadcast reduction, AdaLayerNorm/RMSNorm/RoPE fusion after trace evidence.
  • P2: high implementation or quality risk. Examples: FA to LA replacement, custom Triton/CUDA fused kernels, FP8/quantization, sparsity/Rainfusion-style acceleration.

Every implemented optimization needs A/B validation and a passing quality/precision gate, including math-preserving changes such as padding trim, layout cleanup, cache reuse, or host/runtime cleanup. Use objective metrics when available, keep reviewable artifacts, and compare against baseline self-run variance for precision, quantization, approximate kernels, or backend changes. If quality validation fails or is inconclusive, do not present the optimization as ready to merge.

Optimization Layers

After baseline, parallel-search, and diagnostic traces, summarize optimization opportunities by layer. This is the core of the performance analysis: the goal is to connect evidence to a scoped implementation and a validation plan.

Host and Runtime Optimization

Purpose: remove CPU/Python stalls, synchronization points, allocator overhead, and request-path overhead that leave GPU lanes empty.

Evidence to look for:

  • High idle_pct or large GAP blocks in trace_analyzer.py.
  • Host-stack trace lines such as torch.cuda.empty_cache, cudaStreamSynchronize, cudaDeviceSynchronize, Python locks, scheduler waits, image/video preprocessing, or repeated small allocation paths.
  • Difference between client wall-clock time and server inference_time_s.

Typical candidates:

  • Make avoidable torch.cuda.empty_cache() optional or guard it by memory headroom.
  • Cache scheduler coefficients, timesteps, masks, or other tiny repeated CPU computations when the request shape/steps are fixed.
  • Remove avoidable host-device synchronizations and blocking logging/stat calls.
  • Move expensive preprocessing out of the critical path or cache fixed prompt, image, and transform work for benchmark scenarios.
  • Ensure benchmark scripts record client-side elapsed time, HTTP status, output path, and server response metrics.

Priority guidance:

  • Usually P0 when the change is measurement reliability or an obvious removable synchronization.
  • Usually P1 when it changes scheduling, memory lifetime, or request execution order.

Validation:

  • Re-run non-profiler baseline with same workload and seed.
  • Confirm peak memory headroom if disabling cache cleanup.
  • Confirm generated output exists and quality/seed behavior is unchanged.
Parallelism and Communication Optimization

Purpose: choose the right decomposition for CFG branches, sequence tokens, model weights, VAE tiles, and rank topology.

Evidence to look for:

  • Baseline A/B across CFG, USP/SP, Ring, HSDP/FSDP, and VAE patch parallelism.
  • Stage timing shifts: diffuse, vae.encode, vae.decode, and server end-to-end.
  • NCCL kernel time from trace, not only user_annotation time.
  • Rank imbalance across SP group ranks or CFG branch ranks.
  • Memory headroom and OOM risk.

Typical candidates:

  • CFG=2 x USP=world/2 versus CFG=1 x USP=world for CFG-enabled models.
  • VAE patch parallel on/off or patch size tuning.
  • HSDP/FSDP on/off only if both configurations fit memory.
  • Ulysses versus Ulysses+Ring only after long-sequence diffuse is confirmed dominant and all-to-all is suspected.
  • Rank mapping/topology changes if all-rank traces show stragglers or NCCL wait.
  • Buffer reuse or preallocation for FSDP/HSDP all-gather paths.

Priority guidance:

  • P0/P1 for configuration-only changes with strong measured wins.
  • P1 for buffer reuse or rank mapping changes.
  • P2 for invasive distributed algorithm changes.

Validation:

  • Measure by stage and memory, not only end-to-end.
  • Use one-variable A/B with identical prompt/media/seed/shape/steps.
  • When communication is suspected, compare rank0-3 in one USP group and rank0 versus rank4 across CFG branches for CFG=2 x USP=4.
VAE and Media Pipeline Optimization

Purpose: reduce encode/decode, tiling, split/gather, and media conversion time.

Evidence to look for:

  • Large vae.encode or vae.decode in low-overhead stage timings.
  • Host-stack gaps in VAE tile split, gather, merge, broadcast, or image/video transforms.
  • VAE kernels or cuDNN convolution in operator trace.
  • Whether every rank needs the final decoded tensor.

Typical candidates:

  • Keep VAE patch parallel enabled when it has clear measured benefit.
  • Reduce VAE gather/broadcast to only ranks that need the final media output.
  • Reuse tile metadata, split buffers, or gather buffers.
  • Evaluate bf16/autocast behavior for VAE only with visual quality checks.
  • Avoid redundant image conversion, resize, or tensor construction in repeated benchmark runs.

Priority guidance:

  • P0/P1 if VAE is a large share of the target workload or if a host gap is obvious and low risk.
  • Lower priority when diffuse dominates and VAE is already patch-parallelized.

Validation:

  • Compare vae.encode, vae.decode, server end-to-end, and peak memory.
  • Check output video integrity, artifacts, flicker, and seed stability.
Show full SKILL.md (1,412 more words)Show less
Operator Fusion and Layout Cleanup

Purpose: reduce high-frequency small kernels, memory bandwidth pressure, layout conversions, and launch overhead in transformer and VAE blocks.

Evidence to look for:

  • Top operator tables showing many aten::copy_, aten::cat, split_with_sizes_copy, aten::add, aten::mul, aten::div, norm, activation, RoPE, or reshape/layout kernels.
  • ops_rankN.xlsx by_shape sheet showing repeated small shapes inside the same block path.
  • Trace lanes showing many short kernels between larger GEMM/attention kernels.
  • Source code patterns with repeated elementwise chains or layout conversions.

Typical fusion targets:

  • AdaLayerNorm / RMSNorm / LayerNorm plus scale/shift fusion.
  • RoPE fusion with Q/K layout preparation when shapes are stable.
  • Residual add, scale, gate, and elementwise chains.
  • MLP gate/up/down cleanup, such as fusing activation and multiply around GEMM outputs when feasible.
  • QKV projection and reshape/split/cat path cleanup.
  • Attention pre/post layout cleanup to avoid unnecessary copies, cats, and splits around sequence parallel all-to-all.

Priority guidance:

  • P1 when implemented with existing PyTorch/Triton/local helper patterns and validated against exact outputs or tolerances.
  • P2 when it requires custom CUDA/Triton kernels, changes numerics, or touches attention math directly.

Validation:

  • First prove the operator family is material for the target shape.
  • Use non-profiler A/B for latency and stage timing.
  • Use quality regression checks for generated video stability.
  • Check compile behavior and graph breaks if using torch.compile.
Attention Main-Path Optimization

Purpose: address the dominant self-attention cost when FlashAttention or other attention kernels dominate CUDA time.

Evidence to look for:

  • Operator trace where FlashAttention/SDPA kernels dominate total CUDA time.
  • Attention shape from ops_rankN.xlsx by_shape, model code, or trace metadata.
  • Whether attention cost scales with latent frames, latent H/W, patch size, or CFG duplication.
  • Layout/copy/all-to-all work around attention.

Typical candidates:

  • Verify the attention backend and shape are on the intended fast kernel path.
  • Compare supported attention backends only with identical workload and quality settings.
  • Reduce attention input size by safe model/config choices when allowed: latent resolution, frame count, patching, boundary ratio, or windowing.
  • Remove avoidable layout conversions before/after attention.
  • Reuse condition-side KV or other static inputs if the model structure allows.
  • Consider custom kernels, sparse/window/linear attention, or approximation only after quality risk is accepted.

Priority guidance:

  • P1 for backend/config/layout changes with preserved math.
  • P2 for approximate attention, sparsity, custom kernels, or any change that can alter quality/temporal consistency.

Validation:

  • Always include output quality and seed behavior checks.
  • Compare diffuse, server end-to-end, peak memory, and attention kernel time in diagnostic traces if needed.
Algorithmic, Precision, and Approximation Optimization

Purpose: reduce mathematical work or precision cost beyond local code cleanup.

Evidence to look for:

  • A single operator family dominates even after low-risk runtime, parallel, and fusion work.
  • Memory bandwidth or compute utilization suggests precision or quantization could matter.
  • The user explicitly allows quality-preserving or approximate methods.

Typical candidates:

  • FP8/quantization for transformer or selected projections.
  • Sparsity or Rainfusion-style acceleration.
  • Reduced steps, scheduler changes, distillation, or caching across frames.
  • Approximate attention or linear attention.

Priority guidance:

  • Usually P2 because quality, numerics, and implementation risk are high.

Validation:

  • Requires strict A/B, visual quality review, temporal flicker checks, seed stability, and possibly human evaluation.
Interpolation, Super-Resolution, and E2E Pipeline Optimization

Purpose: optimize the whole user-visible video product, not only the base diffusion invocation. Some deployments trade base-model latency against post-processing, interpolation, or super-resolution stages.

Evidence to look for:

  • E2E latency breakdown across base generation, interpolation, super-resolution, encoding, storage, and response streaming.
  • Fast/slow GPU or fast/slow stage analysis across multiple cards and pipeline stages.
  • User quality target: resolution, FPS, temporal smoothness, and acceptable post-processing artifacts.

Typical candidates:

  • Add or optimize a frame interpolation stage when it reduces required base model frames for the same perceived FPS.
  • Add or optimize a super-resolution model when generating lower base resolution plus SR is faster for the target quality.
  • Analyze E2E2 pipeline behavior: client request, service scheduling, diffusion, VAE/media, post-process, file write, and response.
  • Identify fast/slow cards or stages and rebalance pipeline placement.

Priority guidance:

  • P1 when using proven interpolation/SR components without changing diffusion math.
  • P2 when quality risk is high or the pipeline adds significant operational complexity.

Validation:

  • Measure E2E wall-clock, per-stage server timings, output FPS/resolution, artifacts, flicker, and user-visible quality.
Optimization Candidate Library

Use this table as a compact menu, not as automatic recommendations. Pick items only when baseline metrics, trace evidence, source inspection, or quality tolerance supports them.

LayerCandidateEvidencePriorityValidation focus
MeasurementFreeze server/request/benchmark commandsMissing or drifting commandsP0Repeatable non-profiler A/B
MeasurementSeparate baseline and diagnostic profiler runsProfiler used for latency claimsP0Low-overhead stage timings
Host/runtimeGuard or remove avoidable empty_cacheHost-stack gaps or sync stallsP0/P1Latency, peak memory, OOM safety
Host/runtimeCache scheduler coefficients/timestepsRepeated tiny CPU/GPU workP0/P1Same seed/output, stage timing
Host/runtimeReduce framework scheduling overheadClient time exceeds server timeP1E2E latency, throughput
ParallelCFG=2 x USP=world/2 vs USP=worldCFG doubles forward workP0/P1diffuse, NCCL, memory
ParallelTune VAE patch parallelismVAE encode/decode is materialP0/P1VAE time, output correctness
ParallelHSDP on/off or buffer reuseHSDP affects memory/all-gatherP1Memory, latency, OOM risk
ParallelUlysses vs Ulysses+RingLong sequence all-to-all suspectedP1/P2Rank balance, NCCL kernels
Cross-attnDisable SP for short condition tokensCross-attn comm exceeds computeP1diffuse, correctness
VAE/mediaReduce VAE gather/broadcastRank traces show VAE waitP1Rank balance, output file
VAE/mediaReuse tile metadata/buffersTile split/merge host gapsP1vae.encode/decode, memory
VAE/mediaVAE bf16/autocastVAE float kernels are slowP1/P2Artifacts, flicker, seed stability
Operator fusionAdaLayerNorm/LayerNorm fusionNorm plus scale/shift kernelsP1Numeric tolerance, latency
Operator fusionRMSNorm fusionMany small RMSNorm kernelsP1Numeric tolerance, latency
Operator fusionRoPE cache/fuse/layout cleanupRoPE copy/reshape kernelsP1/P2Kernel count, correctness
Operator layoutQKV or attention layout cleanupCopy/cat/split around attentionP1Copy kernels, compile behavior
AttentionVerify backend fast pathFA/SDPA dominates traceP1diffuse, attention kernels
AttentionFA to LA or selected-head LAAttention remains dominantP2Quality, temporal stability
PrecisionTransformer FP8/quantizationCompute/bandwidth bound and allowedP2Quality, speed, stability
SparsityRainfusion-style accelerationDiT compute remains dominantP2Prompt diversity, quality
PipelineFrame interpolationFewer base frames can meet FPSP1/P2E2E latency, motion artifacts
PipelineSuper-resolutionLower base res plus SR may winP1/P2Detail quality, artifacts
E2EFast/slow-card analysisMulti-card stragglersP0/P1Per-rank/stage wall-clock
Optimization Plan Template

Use this table shape when reporting the next work items:

PriorityLayerCandidateEvidenceExpected benefitImplementation pathValidationQuality risk
P0Host/runtimeGuard empty_cacheHost-stack gap points to torch.cuda.empty_cacheSmall latency reduction, less idleAdd config/env guardNon-profiler A/B, memory checkLow
P1Operator fusionRMSNorm/AdaLayerNorm fusionHigh-frequency norm/elementwise kernelsLower launch/bandwidth overheadUse existing fusion helper or targeted TritonA/B + output checkMedium
P1/P2AttentionAttention layout/backend investigationFA kernel dominates CUDA timePotentially largeInspect shapes/backend and remove layout copiesA/B + trace + qualityMedium/high

Then present a short selection prompt using the same rows:

text
Which candidate should we implement next?

1. P0 Host/runtime: guard empty_cache
   - Expected benefit: small but low-risk latency reduction.
   - Risk: possible memory increase/OOM if memory headroom is insufficient.

2. P1 Operator fusion: inspect by_shape and implement first norm/RoPE/layout fusion
   - Expected benefit: medium if high-frequency small kernels are confirmed.
   - Risk: numerical/compile/quality validation needed.

3. P1/P2 Attention: FA/LA/backend/layout investigation
   - Expected benefit: potentially large.
   - Risk: high quality and implementation risk.

If the user has not chosen an item, default to explaining tradeoffs and asking which candidate to execute. Only proceed autonomously on low-risk P0 measurement or instrumentation fixes.

Analysis Helpers

PyTorch profiler traces are Chrome/Perfetto-compatible JSON files, usually trace_rankN.json or trace_rankN.json.gz. They normally contain a top-level traceEvents list, though some exporters emit the raw event list directly.

Use the checked-in analyzer from the repository root:

bash
python3 .claude/skills/diffusion-perf-opt/scripts/trace_analyzer.py \
  vllm_profile/.../trace_rank0.json.gz \
  --min-gap-ms 5 \
  --topn 20

For rank imbalance or communication questions, pass all relevant rank traces in one command. For host gaps, lower --min-gap-ms to 1 and use a host-stack trace. Read gpu_span_s, busy_union_s, idle_union_s, idle_pct, GAP blocks, Top GPU/operator events by total duration, and Top NCCL-like events by category. Treat cat=user_annotation NCCL ranges as enclosing annotations; prefer cat=kernel or cat=gpu_user_annotation for real device work.

The analyzer summarizes timing only. It does not parse tensor shapes, attribute overlap to individual streams, prove quality, or provide final latency claims. Use ops_rankN.xlsx or PyTorch key averages for shape analysis, and re-test any optimization with non-profiler baseline commands.

Read references/optimization-playbook.md when drafting the optimization table or comparing candidate techniques.

vLLM Omni Heuristics

  • Cross-attention usually should not use USP/SP when text/image condition token count is much smaller than latent video tokens. Confirm via trace; in Wan2.2 I2V, self-attention dominates cross-attention.
  • VAE bf16/autocast is often worthwhile but requires visual quality checks.
  • VAE patch parallel can help decode/encode but may add gather/merge/broadcast overhead. Check whether all ranks need the final decoded tensor.
  • HSDP/FSDP is primarily a memory strategy. If the model fits without it, run an on/off latency comparison.
  • Scheduler work can create small host/device gaps; cache tiny solve coefficients when timesteps/order are known.
  • torch.cuda.empty_cache() can prevent OOM but creates synchronization/idle. Make it optional if memory headroom is sufficient.
  • Command Buffer Full in profiler output is profiler overhead, not a model optimization target.

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in .claude/skills/diffusion-perf-opt of vllm-project/vllm-omni.

  • SKILL.md
  • agents/openai.yaml
  • references/optimization-playbook.md
  • scripts/trace_analyzer.py

Open the folder on GitHubat commit c548a11

Compare with similar skills

Diffusion Perf Opt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Diffusion Perf Opt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Diffusion Perf Opt this skillvllm-project/vllm-omni7.1k—~7.5kAutomated safety check: PassApache-2.0
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~3.9kAutomated safety check: PassNone
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT

Similar skills

  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    911 GitHub stars~3.9k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    333 GitHub stars~3.8k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-omni

All 20 skills in this repo
  • Precheck PR

    vllm-project/vllm-omni

    Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Quantization

    vllm-project/vllm-omni

    Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Review PR

    vllm-project/vllm-omni

    Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.

    7.1k GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • H3 Prompt Writing

    vllm-project/vllm-omni

    Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.

    7.1k GitHub starsUsed in 6 repos~744 tokens
    Auto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    Auto-check passed
  • Add Recipe

    vllm-project/vllm-omni

    Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Questions about Diffusion Perf Opt

What does Diffusion Perf Opt do?

Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. Diffusion Perf Opt is an agent skill from vllm-project/vllm-omni. Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

When should I use Diffusion Perf Opt?

Diffusion Perf Opt fits situations like: Codex is asked to analyze profiling traces; choose parallel strategies; inspect torch profiler trace.json; trace.json.gz timelines.

How do I install Diffusion Perf Opt in Claude Code?

Run `npx skills add vllm-project/vllm-omni --skill diffusion-perf-opt -a claude-code`. Or copy the skill folder (.claude/skills/diffusion-perf-opt in vllm-project/vllm-omni) into .claude/skills/diffusion-perf-opt in your project. Claude Code loads it when a task matches its description.

How do I install Diffusion Perf Opt in Codex?

Run `npx skills add vllm-project/vllm-omni --skill diffusion-perf-opt -a codex`. Or copy the skill folder (.claude/skills/diffusion-perf-opt in vllm-project/vllm-omni) into .agents/skills/diffusion-perf-opt in your project. Codex loads it when a task matches its description.

Can I use Diffusion Perf Opt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill diffusion-perf-opt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diffusion-perf-opt, .gemini/skills/diffusion-perf-opt, .github/skills/diffusion-perf-opt and .opencode/skills/diffusion-perf-opt in your project.

What does Diffusion Perf Opt need to run?

Going by SKILL.md and its folder, Diffusion Perf Opt needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Diffusion Perf Opt access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Diffusion Perf Opt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Diffusion Perf Opt use?

Diffusion Perf Opt is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Diffusion Perf Opt use?

About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 930 tokens, read only when the agent opens those files.

What are the alternatives to Diffusion Perf Opt?

Skills that share tags, products or a category with Diffusion Perf Opt: LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars) and Add Model (guoqingbao/xinfer, 333 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Diffusion Perf Opt?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,072 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.

Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.