Kernel Microbenchmark
guqiong96/Lvllm
Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…
Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.
$ npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-omni production-add-diffusion-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/production-add-diffusion-model .claude/skills/production-add-diffusion-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "production-add-diffusion-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/production-add-diffusion-model into .claude/skills/production-add-diffusion-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "production-add-diffusion-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/production-add-diffusion-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-omni production-add-diffusion-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/production-add-diffusion-model .agents/skills/production-add-diffusion-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "production-add-diffusion-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/production-add-diffusion-model into .agents/skills/production-add-diffusion-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "production-add-diffusion-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-omni production-add-diffusion-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/production-add-diffusion-model .cursor/skills/production-add-diffusion-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "production-add-diffusion-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/production-add-diffusion-model into .cursor/skills/production-add-diffusion-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "production-add-diffusion-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-omni.git --path .claude/skills/production-add-diffusion-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-omni production-add-diffusion-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/production-add-diffusion-model .gemini/skills/production-add-diffusion-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "production-add-diffusion-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/production-add-diffusion-model into .gemini/skills/production-add-diffusion-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "production-add-diffusion-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-omni production-add-diffusion-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/production-add-diffusion-model .github/skills/production-add-diffusion-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "production-add-diffusion-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/production-add-diffusion-model into .github/skills/production-add-diffusion-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "production-add-diffusion-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-omni production-add-diffusion-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/production-add-diffusion-model .opencode/skills/production-add-diffusion-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "production-add-diffusion-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/production-add-diffusion-model into .opencode/skills/production-add-diffusion-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "production-add-diffusion-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
production-add-diffusion-modelProductionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.
Production Add Diffusion Model is an agent skill from vllm-project/vllm-omni. Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. Use when work requires official API parity and input limits, feature-combination evidence, online FP8, distributed layerwise offload, per-request Cache-DiT, CUDA/ROCm/NPU/XPU recipes, faster or sparse attention, fused operators, USP or disaggregation analysis, continuous batching and abort handling, long-running RPS stability, or accuracy/performance/reliability CI. For initial architecture porting, registry wiring, and basic weight…
Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `agents/openai.yaml`, `references/api-and-recipes.md` and `references/feature-patterns.md`).
It sits in AI & LLM Engineering, covering Diffusion and image models and LLM inference and serving. It works with vLLM and CUDA. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c548a11. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Production Add Diffusion Model loads about 5.5k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 147 tokens; SKILL.md has 2,183 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-omni at commit c548a11, republished under its Apache-2.0 licence (© vllm-project). 2,183 words, ~5,460 tokens.
.claude/skills/production-add-diffusion-model/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Use this skill to move a working model integration from Day-0 support to a
measured, fail-closed production deployment. It is deliberately separate from
add-diffusion-model; do not modify or copy
that skill when productionizing a model. If registry, pipeline, or basic loader
work is incomplete, read that skill first and finish the Day-0 vertical slice.
MiniMax-H3 PR #5691
is the merged Day-0 case study. Its open optimization
issue #5700 is an
actively maintained roadmap and evidence map, not a list of already supported
features. It can lag merged implementation: read its current revision, cited
PRs, target source, and maintained recipes, then leave every unverified row
not tested until the target revision has scoped evidence.
Read these references before the corresponding gate:
Create a matrix whose row key is at least:
task x API mode x execution mode/capacity x shape/schedule x attention backend x
packed layout x cache policy x quantization x offload mode x topology x
hardware x dtype x output representation/transportUse exactly these states:
| State | Meaning |
|---|---|
validated | The exact row passed correctness and measured deployment tests, with reproducible artifacts. |
limited | A bounded subset passed; the limitation and rejection/fallback behavior are explicit. |
unsupported | A stable architectural/platform restriction is proven and fails early with an actionable error. |
not tested | No adequate evidence exists. This is the default. |
Never infer compatibility from import success, server startup, another task,
another card, or another topology. Record the model/checkpoint revision,
vLLM-Omni commit, commands, request assets and hashes, seed/schedule/shape,
raw output artifacts, environment, and result for every validated row.
Follow the gates in order. Keep a dense BF16 single-device path as the correctness oracle until every advertised fast path has parity evidence.
Pin the official implementation and checkpoint revision. Build a matrix for:
Normalize offline, sync, and async requests through one validation contract. Reject an invalid task/source combination before downloads, upload persistence, temporary files, async job creation, or engine submission. Use 400 for invalid semantics and 413 for payload limits; clean partial resources on every failure. See API and hardware recipes.
cu_seqlens, padding exclusion,
modality/CFG ownership, split/gather boundaries, and one output per request.Prefer explicit, testable vLLM-Omni/vLLM operators with native fallbacks.
torch.compile remains a separately validated regional optimization; it is
not a substitute for correct operator selection or lifecycle-safe boundaries.
Use the shared cross-platform operators, preserving the official Q/K norm and RoPE order:
from vllm_omni.diffusion.layers.norm import RMSNorm
from vllm_omni.diffusion.layers.rope import (
RotaryEmbedding,
apply_rope_to_qk,
)
self.norm_q = RMSNorm(head_dim, eps=eps)
self.norm_k = RMSNorm(head_dim, eps=eps)
self.rope = RotaryEmbedding(is_neox_style=False)
query = self.norm_q(query) # [B, S, Hq, D] or packed [T, Hq, D]
key = self.norm_k(key) # [B, S, Hkv, D] or packed [T, Hkv, D]
query, key = apply_rope_to_qk(
self.rope,
query,
key,
(cos, sin),
)Verify is_neox_style, half_head_dim, partial rotary dimensions, cos/sin
layout, and packed row ownership against the official model. False means
interleaved/GPT-J style; do not assume the default is NeoX. Shared RMSNorm and
RoPE are two dispatched operators, not automatically one fused kernel.
For packed non-interleaved RoPE, use the shared fused boundary when its contract matches the model:
from vllm_omni.diffusion.layers.fused_qk_norm_rope import fused_qk_norm_rope
query, key = fused_qk_norm_rope(
query, # [T, Hq, D]
key, # [T, Hkv, D]
self.norm_q.weight, # [D]
self.norm_k.weight, # [D]
rope_table, # [T, rotary_dim] = [cos | sin]
self.norm_q.variance_epsilon,
)The current CUDA fast path specializes BF16 head_dim=128, rotary_dim=96;
the public function keeps an eager fallback for unsupported inputs. Verify the
packed row/frequency contract and trace the actual fast path. Do not reshape an
incompatible official RoPE layout merely to enter this kernel.
MiniMax-H3-style partial NeoX RoPE rotates a prefix and passes the suffix through:
self.q_norm = RMSNorm(head_dim, eps=eps, dtype=torch.bfloat16)
self.k_norm = RMSNorm(head_dim, eps=eps, dtype=torch.bfloat16)
self.rope = RotaryEmbedding(is_neox_style=True, half_head_dim=False)
def apply_partial_rope(self, x, freqs, rot_dim):
x_rot, x_pass = x[..., :rot_dim], x[..., rot_dim:]
cos, sin = torch.cos(freqs).to(x.dtype), torch.sin(freqs).to(x.dtype)
return torch.cat((self.rope(x_rot, cos, sin), x_pass), dim=-1)The partial-RoPE snippet assumes import torch and lives on the owning
attention module; adapt names and validated frequency shapes rather than
copying it at module scope.
Call the modules normally so CustomOp selects CUDA/HIP/NPU/XPU/native
implementations. Do not call forward_cuda() directly. Compare the shared
path with native BF16 at operator, block, denoise-trajectory, and artifact
levels before claiming either accuracy or speed.
Use local TP head counts and role-aware shared attention:
from vllm.model_executor.layers.linear import QKVParallelLinear
from vllm_omni.diffusion.attention.layer import Attention
self.qkv_proj = QKVParallelLinear(
hidden_size=hidden_size,
head_size=head_dim,
total_num_heads=num_heads,
total_num_kv_heads=num_kv_heads,
bias=False,
quant_config=quant_config,
prefix=f"{prefix}.qkv_proj",
)
self.attn = Attention(
num_heads=self.qkv_proj.num_heads,
num_kv_heads=self.qkv_proj.num_kv_heads,
head_size=head_dim,
softmax_scale=head_dim**-0.5,
causal=False,
qkv_layout="BSND",
role="self",
prefix=prefix,
)Prove the selected backend from runtime logs/trace and compare it with dense SDPA. Packed metadata must express real sample boundaries; do not build masks that the selected backend ignores. A backend that is faster for one shape or platform is not a global default. See performance patterns for QKV, SwiGLU, AdaLN, sparse attention, redundancy, and fusion examples.
Route the active config to each component, then pass a stable runtime prefix and config into every quantizable vLLM linear:
def resolve_component_quant_config(quant_config, component):
return quant_config.resolve(component) if hasattr(quant_config, "resolve") else quant_config
transformer_quant_config = resolve_component_quant_config(
od_config.quantization_config,
"transformer",
)
self.transformer = MyDiT(
od_config,
quant_config=transformer_quant_config,
prefix="transformer",
)
self.qkv_proj = QKVParallelLinear(
hidden_size=hidden_size,
head_size=head_dim,
total_num_heads=num_heads,
total_num_kv_heads=num_kv_heads,
quant_config=quant_config,
prefix=f"{prefix}.qkv_proj",
)Use build_quant_config() for direct construction/testing:
from vllm_omni.quantization import build_quant_config
quant_config = build_quant_config({
"transformer": {"method": "fp8"},
"text_encoder": None,
"vae": None,
"default": None,
})Preserve each parameter's vLLM weight_loader; perform checkpoint layout
conversion before invoking it. Account for every QKV/gate-up source shard.
Keep precision-sensitive norm/modulation/embedders in BF16 unless separately
proven, using the current safe_quant_config pattern where applicable.
Validate resident DiT FP8 first: prove which modules are quantized, HBM saved,
BF16-vs-FP8 trajectory/artifact quality, latency, and throughput. Text encoder,
VAE, other hardware, TP/cache/DLO combinations are separate rows.
Pre-quantized, pruned, rotated, or adapter-modified checkpoints are separate
model/loading lanes; compare them with the released BF16 model without calling
the result a pure quantization ablation. Read
../quantization/SKILL.md if present and prefer the
target revision's implementation/docs when sibling guidance differs.
Declare topology instead of relying on heuristic discovery:
from typing import ClassVar
import torch.nn as nn
from vllm_omni.diffusion.models.interface import SupportsComponentDiscovery
from vllm_omni.diffusion.offloader import OffloadPlan
class MyPipeline(nn.Module, SupportsComponentDiscovery):
_dit_modules: ClassVar[list[str]] = ["transformer"]
_encoder_modules: ClassVar[list[str]] = ["text_encoder"]
_vae_modules: ClassVar[list[str]] = ["vae"]
_resident_modules: ClassVar[list[str]] = []
_offload_plan: ClassVar[OffloadPlan] = OffloadPlan(
on_demand_component_paths=frozenset({"text_encoder", "vae"}),
block_attrs={"transformer": ("blocks",)},
offload_submodules={"token_refiner": "blocks"},
resident_dit_paths=frozenset({"transformer"}),
encoder_block_attrs={"text_encoder": ("encoder.layers",)},
)Add a test that resolves every dotted component/block path and fails if it is missing or is not the expected module/container; discovery may otherwise warn and skip. Validate ordinary layerwise offload and distributed layerwise offload (DLO) independently.
Keep component lifecycle and host-weight storage separate. OffloadPlan
declares what can be streamed; the diffusion loader owns any HostWeightPlan
and hands the exact prevalidated plan to DLO. Do not make the offloader rescan
checkpoint files or duplicate loader name/shape/dtype decisions. Direct mmap is
a fail-closed optimization: current preflight requires TP1, no HSDP, no online
quantization, complete bindings, and represented transforms. Otherwise the
ordinary loader remains authoritative.
If the target revision supports Host Weight Runtime (HWR), treat it as a
separate, exact-identity startup/cache path rather than generic offload or
zero-copy execution. Qualify preferred population/fallback and required
consume-only failure independently; verify immutable manifests, final-layout
restore, lease cleanup, corruption/quarantine, and node-local sharing. Current
model contracts may be BF16 no-AllGather-only, so reject AllGather, online
quantization, HSDP, LoRA/adapted weights, and unrecognized load formats unless
the exact target revision explicitly adds them.
On small-HBM cards, record load/materialization peak, resident peak, encode,
each denoise wave, decode, transient per-rank HBM, host PSS, H2D, and collective
time. Test DLO AllGather and --dlo-no-use-allgather as different deployments.
Sweep the supported resident-block policies and publish the non-dominated
latency-HBM frontier; a single minimum-memory or fastest point is not the DLO
trade-off. Keep request-specific preparation replica-local. Only weight
materialization may enter the declared DLO collective group; never use WORLD
for per-request work in a DP deployment.
Direct checkpoint mmap preflight is limited to TP1 without HSDP or online
quantization; TP>1 falls back to the ordinary TP-aware loader and can still feed
DLO AllGather or no-AllGather after scoped E2E. HSDP remains incompatible with
DLO AllGather. A target revision may allow finalized per-tensor online FP8
through the ordinary loader and AllGather; keep every other online quantizer
fail-closed, and require model/card/topology E2E before promotion.
Use deploy YAML for DP under --omni; vLLM DP CLI flags are rejected.
Server-wide Cache-DiT support does not imply a per-request quality API. Only
add request policy after defining calibrated quality tiers and a safe hook
lifecycle using SupportsRequestScopedCacheDiT, CacheDiTRequestSpec, and
RequestScopedCacheDiTRuntime. Validate alternating and concurrent quality
tiers, real cache hits, refresh/reset, success, error, disconnect, abort, and
the next uncached request. Do not batch different quality policies unless the
batch compatibility key and hook ownership make that safe.
Then test Cache-DiT against FP8, DLO, TP/SP, compile, and batching one combination at a time. See production feature patterns for the lifecycle skeleton and compatibility matrix.
Freeze one canonical workload manifest before profiling. Keep lossless runtime A/Bs, accelerated paths with numerical/quality trade-offs, and production topology studies in separate result lanes. Reconcile the client boundary as queue + encode + denoise + video/audio decode + output transport/codec + residual; do not hide a material residual in another stage.
Profile stage time, GPU kernels, CPU/GPU synchronization, allocations, H2D, output payload/copies, and collectives before choosing work. Apply and A/B one change at a time:
.item() and
unused masks/allocations.SiluAndMul, RMSNorm/RoPE, AdaLN, and
layout/residual operators with guarded native fallbacks. Preserve every
materialized dtype/rounding boundary consumed by the unfused graph.skip_sequence_parallel. Qualify regular and accelerated Ulysses transport
independently; activation logs, JIT/readiness time, workspace growth,
stream ownership, maximum-shape warmup, and numerical drift are gates.
Preserve the checkpoint's Q:KV head ratio when padding GQA for Ulysses, and
validate strided QKV staging rather than assuming packed contiguity.Do not combine independent optimizations into one performance claim. Preserve raw before/after runs with fixed workload and warmup policy.
Treat request-mode batching and step-wise execution as separate capabilities:
supports_request_batch = True and
DiffusionRequestBatch -> list[DiffusionOutput], exactly one output per
request;SupportsStepExecution contract:
prepare_encode, denoise_step, step_scheduler, and post_decode.Keep all mutable scheduler, RNG, latent, mask, cache, and step state request
scoped. First validate --step-execution --max-num-seqs 1 and step-boundary
abort; only then test heterogeneous multi-request waves. Step continuous
batching is experimental and is not automatically compatible with Cache-DiT.
Before recommending it, predeclare a useful case and success threshold, then
A/B it against request mode under the same arrival process. Structural support
or a generic bridge is not latency, throughput, or cancellation evidence for a
model.
Test queued and in-flight abort, client disconnect, timeout, OOM, worker error,
and restart. Require one terminal result, idempotent cleanup, no post-decode
after abort, no cache/temp/VRAM/request-ID leak, and a successful next request.
Inject late worker results/exceptions after cancellation and prove the result
pump stays alive instead of raising InvalidStateError. For large batched or
long-window video outputs, freeze dtype, range, layout, contiguity, ownership,
payload size, serializer/codec limits, and offline-versus-HTTP semantics.
Validate device-side output preparation, D2H/IPC or shared handles, remote
encoding, event-loop responsiveness, and client materialization as separate
stages, including abort and consumer failure cleanup.
A model/VAE callback that publishes ordered chunks is only a producer contract;
it is not transport, backpressure, cancellation, public streaming, or
time-to-first-frame evidence. Distinguish complete-response faster-than-playback
(client E2E / output duration <= 1) from streaming and report first-chunk,
steady cadence, finalization, and complete-artifact latency separately.
Run below-, near-, and above-saturation mixed-RPS soaks and report success/error
rate, throughput, queue time, memory slope, temp growth, and worker health.
Report p50/p95/p99 only from a declared arrival model with enough measured
requests for the claimed percentile. See
serving and CI validation.
For every target vendor/card/topology, publish exact environment, serve/deploy
command, one complete JSON or multipart curl per supported task, expected
output checks, benchmark command, warmup/repeat/concurrency policy, raw metrics,
per-rank HBM/host PSS, accuracy artifact, and limitations. Mark each
CUDA/ROCm/NPU/XPU row validated, unsupported, or not tested; never inherit
support from the platform abstraction or another card. Re-run DLO on each
small-HBM recipe.
Maintain separate CI gates:
A model is production-ready only when:
not tested and are not presented as supported.Day-0 support is still a valid milestone, but title, docs, and matrix must say which representative task passed and which production gates remain.
© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in .claude/skills/production-add-diffusion-model of vllm-project/vllm-omni.
Open the folder on GitHubat commit c548a11
Production Add Diffusion Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Production Add Diffusion Model this skillvllm-project/vllm-omni | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | |
| Kernel Microbenchmarkguqiong96/Lvllm | 464 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~2.8k | Automated safety check: Pass | None | |
| Vllm Deploy Simplevllm-project/vllm-skills | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 |
guqiong96/Lvllm
Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
vllm-project/vllm-skills
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
vllm-project/vllm-omni
Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.
vllm-project/vllm-omni
Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.
vllm-project/vllm-omni
Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.
vllm-project/vllm-omni
Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.
vllm-project/vllm-omni
Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…
Categories
Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. Production Add Diffusion Model is an agent skill from vllm-project/vllm-omni. Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.
Production Add Diffusion Model fits situations like: work requires official API parity and input limits; feature-combination evidence; distributed layerwise offload; per-request Cache-DiT.
Run `npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a claude-code`. Or copy the skill folder (.claude/skills/production-add-diffusion-model in vllm-project/vllm-omni) into .claude/skills/production-add-diffusion-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a codex`. Or copy the skill folder (.claude/skills/production-add-diffusion-model in vllm-project/vllm-omni) into .agents/skills/production-add-diffusion-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill production-add-diffusion-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/production-add-diffusion-model, .gemini/skills/production-add-diffusion-model, .github/skills/production-add-diffusion-model and .opencode/skills/production-add-diffusion-model in your project.
SKILL.md names no scripts, command-line tools or credentials: Production Add Diffusion Model is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Production Add Diffusion Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Production Add Diffusion Model: Kernel Microbenchmark (guqiong96/Lvllm, 464 stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,072 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.
Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.