Agent skill

Add Diffusion Model

by vllm-project in vllm-project/vllm-omni

Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add Diffusion Model

skills CLI
$ npx skills add vllm-project/vllm-omni --skill add-diffusion-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-omni add-diffusion-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-diffusion-model .claude/skills/add-diffusion-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-diffusion-model
GitHub stars
7.1k
Token cost
~7k tokens
SKILL.md length
2,468 words
Files
7 (incl. references)
Skills in repo
20
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

  • Reviewing a new diffusion model
  • SKILL.md covers Overview, Prerequisites, Step 0: Classify the Migration… and Path A: Diffusers-Based Model, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Porting a Diffusers pipeline

What it does

Add Diffusion Model is an agent skill from vllm-project/vllm-omni. Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT, offload, and parallelism support (TP, SP/USP, CFG-Parallel, HSDP). Use when integrating or reviewing a new diffusion model, porting a Diffusers pipeline or custom model repository, creating a DiT adapter, reusing shared examples, or qualifying multi-GPU and memory optimizations.

Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/cache-dit-patterns.md`, `references/custom-model-patterns.md` and `references/native-model-integration-checklist.md`).

It sits in AI & LLM Engineering, covering Diffusion and image models, AI video generation and LLM inference and serving. It works with vLLM. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.

When your agent uses it

  • Reviewing a new diffusion model
  • Porting a Diffusers pipeline
  • Custom model repository
  • Creating a DiT adapter

Example prompts

  • “/add-diffusion-model”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c548a11. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Diffusion Model loads about 7k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 2,468 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-omni at commit c548a11, republished under its Apache-2.0 licence (© vllm-project). 2,468 words, ~6,999 tokens.

Download SKILL.mdSave it as .claude/skills/add-diffusion-model/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
add-diffusion-model
description
Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT, offload, and parallelism support (TP, SP/USP, CFG-Parallel, HSDP). Use when integrating or reviewing a new diffusion model, porting a Diffusers pipeline or custom model repository, creating a DiT adapter, reusing shared examples, or qualifying multi-GPU and memory optimizations.

Adding a Diffusion Model to vLLM-Omni

Overview

This skill guides you through adding a new diffusion model to vLLM-Omni. The model may come from HuggingFace Diffusers (structured pipeline) or from a private/custom repo. The workflow differs significantly depending on the source.

Prerequisites

Before starting, determine:

  1. Model category: Text-to-Image, Text-to-Video, Image-to-Video, Image Editing, Text-to-Audio, or Omni
  2. Reference source: Diffusers pipeline, custom repo, or a combination
  3. Model HuggingFace ID or local checkpoint path
  4. Architecture: Scheduler, text encoder, VAE, transformer/backbone

Step 0: Classify the Migration Path

Check the model's HF repo for model_index.json. This determines your path:

ScenarioHow to identifyMigration path
Already supported_class_name in model_index.json matches a key in _DIFFUSION_MODELS in registry.pySkip implementation, then validate model-specific examples, tests, and docs as needed
Diffusers-basedHas standard model_index.json with _diffusers_version, subfolders for transformer/, vae/, etc.Follow Path A below
Native non-Diffusers modelNo Diffusers index, non-standard checkpoint hierarchy, or custom architecture in a separate repoFollow Path B below; port the runtime natively unless an external adapter was explicitly requested
HybridHas some diffusers components (VAE) but custom transformer/fusionMix of Path A and Path B

Before coding, write a short integration contract covering runtime ownership, checkpoint discovery, reference revision, I/O geometry, attention semantics, CFG behavior, auxiliary components, target hardware, and the default deployment. For Path B, hybrid models, or any optimization work, read references/native-model-integration-checklist.md and use its phase gates.

Path A: Diffusers-Based Model

For models with a standard diffusers layout. See references/transformer-adaptation.md for detailed code patterns.

A1. Analyze model_index.json

Identify components: transformer, scheduler, vae, text_encoder, tokenizer.

A2. Create model directory
vllm_omni/diffusion/models/your_model_name/
├── __init__.py
├── pipeline_your_model.py
└── your_model_transformer.py
A3. Adapt transformer
  1. Copy from diffusers source. Remove mixins (ModelMixin, ConfigMixin, AttentionModuleMixin).
  2. Replace attention with vllm_omni.diffusion.attention.layer.Attention (QKV shape: [B, seq, heads, head_dim]).
  3. Add od_config: OmniDiffusionConfig | None = None to __init__.
  4. Add load_weights() method mapping diffusers weight names to vllm-omni names.
  5. Add class attributes for acceleration features such as _repeated_blocks and _layerwise_offload_blocks_attrs (see references/transformer-adaptation.md for examples).
A4. Adapt pipeline

Inherit from nn.Module. The key contract:

python
class YourPipeline(nn.Module):
    def __init__(self, *, od_config: OmniDiffusionConfig, prefix: str = ""):
        # Load VAE, text encoder, tokenizer via from_pretrained()
        # Instantiate transformer (weights loaded later via weights_sources)
        self.weights_sources = [
            DiffusersPipelineLoader.ComponentSource(
                model_or_path=od_config.model, subfolder="transformer",
                prefix="transformer.", fall_back_to_pt=True)]

    def forward(self, req: OmniDiffusionRequest) -> DiffusionOutput:
        # Encode prompt → prepare latents → denoise loop → VAE decode
        return DiffusionOutput(output=output)

    def load_weights(self, weights):
        return AutoWeightsLoader(self).load_weights(weights)

Add post/pre-process functions in the same pipeline file. Register them in registry.py.

For pipelines with a standard denoising loop, prefer the existing progress bar pattern instead of hand-rolled logging.

python
from vllm_omni.diffusion.models.progress_bar import ProgressBarMixin

class YourPipeline(nn.Module, ProgressBarMixin):
    def forward(self, req: OmniDiffusionRequest) -> DiffusionOutput:
        # ... prepare timesteps / latents ...
        with self.progress_bar(total=len(timesteps)) as progress_bar:
            for i, t in enumerate(timesteps):
                # predict noise / scheduler step
                latents = ...
                progress_bar.update()

        return DiffusionOutput(output=output)

For custom loop structures, follow vllm_omni/diffusion/models/progress_bar.py and existing pipelines using ProgressBarMixin.

A5. Register, test, docs → continue at Step 4 below.

Path B: Native / Non-Diffusers Model

For models without a Diffusers pipeline—weights in custom formats and model code in another public or private repository. Treat that repository as a pinned correctness oracle. A request for native support means the vLLM-Omni runtime must not import the reference implementation; an external adapter is appropriate only when the requested scope explicitly permits that dependency. See references/custom-model-patterns.md for concrete integration patterns.

B1. Understand the reference repo

Study the original model's code to identify:

  • Model architecture files (transformers, fusion modules, embeddings)
  • Weight format (safetensors, .pth, custom checkpoint structure)
  • Weight loading helpers (custom init functions, checkpoint loaders)
  • Pre/post-processing (image/audio transforms, tokenization, VAE encode/decode)
  • External dependencies (packages not on PyPI)
  • Config format (JSON config files, hardcoded dicts)
B2. Decide what lives WHERE

This is the key design decision for custom models. Follow these placement rules:

Code typeWhere to placeExample
Pipeline orchestration (init, forward, denoise loop)vllm_omni/diffusion/models/<name>/pipeline_<name>.pyAlways required
Custom transformer/backbone (ported and adapted to vllm-omni)vllm_omni/diffusion/models/<name>/<name>_transformer.py or similarwan2_2.py, fusion.py, bagel_transformer.py
Custom sub-models (VAE, fusion, autoencoder)vllm_omni/diffusion/models/<name>/ as separate filesautoencoder.py, fusion.py
Reference-only codeKeep outside the runtime; use a pinned revision for golden outputs and architecture analysisReference inference script
Explicit external adapter dependencyExternal package, only when the requested scope and maintainer direction allow itCompatibility adapter, not native support
Hardcoded model configsModule-level dicts in pipeline fileVIDEO_CONFIG, AUDIO_CONFIG dicts
Download/setup scriptexamples/offline_inference/<name>/download_<name>.pydownload_<name>.py
Custom model_index.jsonGenerated by download script, placed at model rootMinimal: {"_class_name": "YourPipeline", ...}
B3. Handle external dependencies

If the model's code lives in a separate git repo, first decide whether the requested deliverable is native support or an external adapter. Do not silently choose the adapter path.

Option 1: Port the code directly (default for native support)

Copy the essential model files into vllm_omni/diffusion/models/<name>/ and adapt them to shared vLLM-Omni contracts. Keep checkpoint loading strict and use the pinned reference only to generate parity evidence.

Option 2: Import with graceful fallback (adapter scope only)

python
try:
    from external_model.utils import init_vae, load_checkpoint
except ImportError:
    raise ImportError(
        "Failed to import from dependency 'external_model'. "
        "Please run the download script first."
    )

Use an external runtime dependency only when the user explicitly requested an adapter or a maintainer approved the exception. Document the dependency, pinned revision, installation path, and unsupported native features.

B4. Handle custom weight loading

Custom models have two common patterns for weight loading:

Pattern 1: Bypass standard loader (eager custom init)

When the original model has complex custom init functions that load weights in __init__:

python
class CustomPipeline(nn.Module):
    def __init__(self, *, od_config, prefix=""):
        super().__init__()
        model = od_config.model
        # Load everything eagerly in __init__ using custom helpers
        self.vae = custom_init_vae(model, device=self.device)
        self.text_encoder = custom_init_text_encoder(model, device=self.device)
        self.transformer = CustomFusionModel(CONFIG)
        load_custom_checkpoint(
            self.transformer,
            checkpoint_path=os.path.join(model, "model.safetensors"),
        )
        # NO weights_sources defined — bypasses standard loader

    def load_weights(self, weights):
        pass  # No-op — all weights loaded in __init__

Pattern 2: Use standard loader with custom load_weights (BAGEL style)

When weights are in safetensors format but need name remapping:

python
from vllm.model_executor.model_loader.weight_utils import default_weight_loader

class CustomPipeline(nn.Module):
    def __init__(self, *, od_config, prefix=""):
        super().__init__()
        # Instantiate model architecture without weights
        self.bagel = BagelModel(config)
        self.vae = AutoEncoder(ae_params)

        # Point loader at the safetensors in the model root
        self.weights_sources = [
            DiffusersPipelineLoader.ComponentSource(
                model_or_path=od_config.model,
                subfolder=None,  # weights at root, not in subfolder
                prefix="",
                fall_back_to_pt=False,
            )
        ]

    def load_weights(self, weights):
        # Custom name remapping for non-diffusers weight names
        params = dict(self.named_parameters())
        loaded = set()
        for name, tensor in weights:
            # Remap original weight names to vllm-omni module names
            name = self._remap_weight_name(name)
            if name in params:
                param = params[name]
                weight_loader = getattr(param, "weight_loader", default_weight_loader)
                weight_loader(param, tensor)
                loaded.add(name)
        return loaded
B5. Create the model_index.json

Prefer a model_index.json at the model root when vLLM-Omni owns an assembled checkpoint directory. For custom models, this is minimal:

json
{
    "_class_name": "YourModelPipeline",
    "custom_key": "path/to/custom_weights.safetensors"
}

The _class_name must match a key in _DIFFUSION_MODELS in registry.py. Additional keys are model-specific (accessed via od_config.model_config).

If the released repository is immutable and has neither a root config.json nor a Diffusers index, add it to a generic native-checkpoint signature resolver: match the exact Hub ID or a distinctive set of local files and return the pipeline class. Do not use model-name substrings or add parallel one-off predicates in CLI, config, and serving consumers.

If the model's weights come from multiple HF repos, write a download script that:

  1. Downloads from each repo
  2. Assembles into a single directory
  3. Generates model_index.json
  4. Installs any external dependencies (git clone + .pth file)

Place at: examples/offline_inference/<name>/download_<name>.py

B6. Handle multi-modal inputs

If the model accepts images, audio, or other multi-modal inputs, implement the protocol classes from vllm_omni/diffusion/models/interface.py:

python
from vllm_omni.diffusion.models.interface import SupportImageInput, SupportAudioInput

class MyPipeline(nn.Module, SupportImageInput, SupportAudioInput):
    # Protocol markers — the engine uses these to enable proper input routing
    pass

Preprocessing for custom models is typically done inside forward() rather than via registered pre-process functions, since the logic is often tightly coupled to the model.

B7. Continue at Step 4 below.

Common Steps (Both Paths)

Step 4: Register Model in registry.py

Edit vllm_omni/diffusion/registry.py:

python
_DIFFUSION_MODELS = {
    "YourModelPipeline": ("your_model_name", "pipeline_your_model", "YourModelPipeline"),
}
_DIFFUSION_POST_PROCESS_FUNCS = {
    "YourModelPipeline": "get_your_model_post_process_func",  # if applicable
}
_DIFFUSION_PRE_PROCESS_FUNCS = {
    "YourModelPipeline": "get_your_model_pre_process_func",  # if applicable
}

The registry key is the _class_name from model_index.json. The tuple is (folder_name, module_file, class_name).

Create __init__.py exporting the pipeline class and any factory functions.

Step 5: Run, Test, Debug

Use the appropriate existing example script:

CategoryScript
Text-to-Imageexamples/offline_inference/text_to_image/text_to_image.py
Text-to-Videoexamples/offline_inference/text_to_video/text_to_video.py
Image-to-Videoexamples/offline_inference/image_to_video/image_to_video.py
Image-to-Imageexamples/offline_inference/image_to_image/image_edit.py
Text-to-Audioexamples/offline_inference/text_to_audio/text_to_audio.py

Reuse these shared scripts even for custom models when their request and output contracts fit. Create a dedicated model script only when the shared category cannot represent the protocol, and document that gap in the PR.

Validation: No errors, output is meaningful, quality matches reference implementation.

See references/troubleshooting.md for common errors.

Step 6: Add Example Scripts

Only when the shared category scripts cannot represent the model, create:

  • examples/offline_inference/your_model_name/ — offline script + README
  • examples/online_serving/your_model_name/ — server script + client
  • Download script if weights require assembly from multiple sources
Step 7: Update Documentation

Follow the add-recipe skill to add or update the model-family recipe and its recipes/README.md row with verified specifications, hardware, commands, feature links, and qualification evidence.

Required updates:

  1. docs/user_guide/diffusion/parallelism/overview.md — parallelism support overview/table
  2. docs/user_guide/diffusion/cpu_offload.md — if CPU offload supported (add to supported models table)
  3. docs/user_guide/diffusion/cache_acceleration/teacache.md — if TeaCache supported
  4. docs/user_guide/diffusion/cache_acceleration/cache_dit.md — if Cache-DiT supported
  5. Offline example docs under examples/offline_inference/<name>/ (README.md or category-specific .md)
  6. examples/online_serving/<name>/README.md — online serving docs
Step 8: Add E2E Tests

Follow the vllm-omni-test skill for markers, file naming, Buildkite wiring, and run commands. Also read l4_functionality_tests.inc.md, test_system_overview.md, and test_writing_guide.md.

Classify the model's CI priority first:

PriorityRequired test levelsFiles & markers
High (listed in #1832 or on the diffusion hot path)L1 · L2 online · L3 online + offline · L4 feature + performanceSee table below
Medium (normal priority in L4 docs)L3 online + offline · L4 feature onlyFewer L4 parametrized rows
LowL4 feature onlyOne or two *_expansion.py cases

Per-level deliverables (diffusion / pytest.mark.diffusion):

LevelLocationMarkerCI pipelineNotes
L1tests/diffusion/models/{slug}/, tests/diffusion/cache/, transformer unit testscore_model + cputest-ready.ymlWeight remap, _sp_plan, cache enabler registration, shape contracts
L2tests/e2e/online_serving/test_{slug}.py (and offline if the category is offline-first)core_model + advanced_model (both on baseline smoke) + diffusion + @hardware_test / hardware_markstest-ready.ymlDefault deploy smoke — minimal num_inference_steps, single prompt
L3tests/e2e/online_serving/test_{slug}.py and tests/e2e/offline_inference/test_{slug}.py when offline mattersBaseline smoke: core_model + advanced_model; heavier cases: advanced_model only (+ diffusion)test-merge.yml or merged into nightly diffusion function jobReal weights, streaming/API paths, LoRA/offload smoke
L4tests/e2e/online_serving/test_{slug}_expansion.py (+ offline expansion if needed)full_model + diffusiontest-nightly.yml (X2I / X2V / X2A function groups)Feature combos per #1832; perf → tests/dfx/perf/tests/test_{model}_vllm_omni.json with per-case mark (hardware_marks + full_model + diffusion)

L2 & L3 online — same file, dual marks on the baseline smoke: The first / simplest case in test_{slug}.py (default deploy, minimal steps, single prompt) should carry both @pytest.mark.core_model and @pytest.mark.advanced_model on the same function so L2 (test-ready.yml) and L3 (test-merge.yml) share one smoke test. Heavier deploy variants or API paths in the same file use advanced_model only. When L3 moves to nightly, migrate those heavier cases into test_{slug}_expansion.py with full_model and remove the dedicated test-merge.yml job (see test_longcat_image_expansion.py, test_qwen_image_expansion.py).

L4 design (high priority): Combine multiple supported features (Cache-DiT, TP, USP, CFG, HSDP, CPU offload, quantization) into few parametrized OmniServerParams rows so each feature appears in at least one case without exploding GPU jobs. Shard single-GPU vs multi-GPU cases across the nightly X2I/X2V function steps (cards_1 vs not cards_1).

L4 design (medium / low): One or two parametrized rows covering the best quality/perf trade-off; skip perf JSON unless the model is high priority.

Reference implementations: tests/e2e/online_serving/test_qwen_image_edit_expansion.py, tests/e2e/online_serving/test_longcat_image_expansion.py, tests/e2e/online_serving/test_hunyuan_video_15_expansion.py.

Keep model-specific code inside test modules — not tests/helpers/{slug}.py: deploy constants, prompts, sampling dicts, and inline request_config / form_data belong in each test_{slug}.py and test_{slug}_expansion.py. Do not add per-model files under tests/helpers/; reuse only repo-wide harness (mark, media, runtime, stage_config, assertions). L2+ online/offline e2e: reuse or add send_*_request in tests/helpers/runtime.py — tests call the handler, not raw omni.generate / HTTP. See vllm-omni-test skill § Runtime send helpers.

Keep the model suite proportional using the six distinct failure owners in the native integration checklist. Combine supported features into a few parametrized E2E rows instead of creating one test per optimization.

Show full SKILL.md (825 more words)Show less
Step 9: Add Cache-DiT Acceleration

Add caching only after uncached single-device correctness. Read references/cache-dit-patterns.md, use the automatic single-block-list path when possible, and add a registered BlockAdapter only for genuinely custom block topology.

Verify a real cache hit and compare quality with the uncached baseline. Make a speed claim only at a realistic step count where warmup permits hits; an all-warmup smoke proves integration, not acceleration.


Step 10: Add Parallelism Support

After the model works on a single GPU, add multi-GPU parallelism. Add each type incrementally, testing after each addition.

See references/parallelism-patterns.md for detailed code patterns and API reference.

Recommended order: TP → SP/USP → CFG Parallel → HSDP

10a. Tensor Parallelism (TP)

Replace compatible projections with vLLM parallel linears, preserve checkpoint fusion/loading, and use local head counts. Require query/KV head divisibility and compare a multi-rank forward with the one-rank oracle.

10b. Sequence Parallelism (SP / USP)

Prefer the declarative _sp_plan. For packed variable-length attention or learned-sink LSE correction, keep model math explicit and reuse shared exchange utilities. Validate uneven sequence splits, RoPE coordinates, and outputs against the one-rank oracle.

10c. CFG Parallel

Confirm the model uses CFG. Then reuse CFGParallelMixin, overriding prediction or recombination only for non-standard or multi-output pipelines. Distinguish a packed positive/negative implementation from two independent branches and validate the two-rank result against the packed one-rank oracle.

10d. HSDP (Hybrid Sharded Data Parallel)

Declare layer shard conditions and ignored rank-local modules; preserve mixed checkpoint dtypes. HSDP cannot combine with TP. Measure parameter loading, FSDP materialization, warm HBM, and host PSS rather than assuming sharding saves peak memory. Keep the resident layout as default if HSDP is worse.

10e. Update parallelism documentation

After adding parallelism support, update:

  1. docs/user_guide/diffusion/parallelism/overview.md — add your model to the support overview/table
  2. Record which parallelism methods are supported (USP, Ring, CFG, TP, HSDP, VAE-Patch)
Step 11: Add CPU Offload Support

Implement SupportsComponentDiscovery on your pipeline class to enable --enable-cpu-offload and --enable-layerwise-offload. The protocol declares which submodules the offloader should manage:

python
from typing import ClassVar
from vllm_omni.diffusion.models.interface import SupportsComponentDiscovery

class YourPipeline(nn.Module, SupportsComponentDiscovery):
    _dit_modules: ClassVar[list[str]] = ["transformer"]
    _encoder_modules: ClassVar[list[str]] = ["text_encoder"]
    _vae_modules: ClassVar[list[str]] = ["vae"]
    _resident_modules: ClassVar[list[str]] = []  # optional
  • _dit_modules: denoising submodules (kept on GPU during diffusion loop)
  • _encoder_modules: encoder/vision submodules (offloaded to CPU during diffusion loop)
  • _vae_modules: VAE(s) (handled by both sequential and layerwise backends)
  • _resident_modules: additional modules to pin on GPU during layerwise offloading (e.g. embedders, connectors). Only used by the layerwise backend. Optional — defaults to [].

All attribute names support dotted paths for nested submodules (e.g. "pipe.transformer", "bagel.time_embedder").

Pipelines without SupportsComponentDiscovery fall back to scanning well-known attribute names (transformer, text_encoder, vae, etc.), which fails for non-standard names.

Keep model-specific checkpoint paths, nested block aliases, layout transforms, and component lifecycles in the model package. Change a shared offloader only for a general contract, demonstrate another consumer or a framework-level bug, and add one focused shared regression. Avoid if ModelName branches in shared backends.

Step 12: Performance Profiling

After verifying correctness and implementing parallelism/caching, profile the model's performance to identify bottlenecks and ensure optimal execution.

See the Profiling Single-Stage Diffusion guide for detailed instructions on:

  1. Using the PyTorch profiler (profiler: "torch") to capture detailed CPU/CUDA traces.
  2. Using Nsight Systems (nsys) with profiler: "cuda" for low-overhead CUDA traces.
  3. Controlling profiling via omni.start_profile() and omni.stop_profile().

Report cold E2E, warm user latency, steady-state wave time, peak allocated and reserved HBM, host PSS for the full process tree, and stage boundaries. State whether prompt encoding and VAE/audio decoding are included. For multi-device layouts, plot user latency against throughput per device and keep different denoising-step counts on separate Pareto frontiers.


Pre-commit conventions

New library files must pass the local gates in docs/contributing/README.md. That page is the full hook list (SPDX, forbidden imports including Hugging Face Hub / Triton / pickle, torch.cuda, mypy, test marks, markdownlint, Buildkite, shellcheck). In particular:

  • SPDX copyright is vLLM-Omni project (stale vLLM project is rewritten).
  • Use import regex as re and pybase64 in vllm_omni/; do not import stdlib re or base64. Hugging Face Hub downloads go through vllm.transformers_utils.repo_utils.
  • Do not add torch.cuda.* call sites; use current_omni_platform.
  • New tests/**/test_*.py files need a CI level mark and a hardware mark.
  • Do not expand CHECK_IMPORTS[*].allowed_files or ALLOWED_FILES without review.
  • GitHub Actions skips SPDX/shellcheck/mypy-3.10/test-marks/markdownlint; run pre-commit locally.

Iterative Development Tips

  1. Start minimal: Basic generation first, no parallelism/caching
  2. Use --enforce-eager: Disable torch.compile during debugging
  3. Use small models: Test with smaller variants first
  4. Check tensor shapes: Most errors are reshape mismatches in attention
  5. Add features incrementally: Single GPU → TP → SP → CFG → HSDP → Cache-DiT
  6. For custom models: Run the pinned reference separately, then port the runtime natively; do not ship temporary imports from the reference implementation
  7. Cache-DiT before parallelism tuning: Cache-DiT is lossy — verify quality at baseline before combining with parallelism
  8. Combine lossless + lossy: e.g., TP + SP + Cache-DiT for maximum throughput

Reference Files

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in .claude/skills/add-diffusion-model of vllm-project/vllm-omni.

  • SKILL.md
  • references/cache-dit-patterns.md
  • references/custom-model-patterns.md
  • references/native-model-integration-checklist.md
  • references/parallelism-patterns.md
  • references/transformer-adaptation.md
  • references/troubleshooting.md

Open the folder on GitHubat commit c548a11

Compare with similar skills

Add Diffusion Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Diffusion Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Diffusion Model this skillvllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4642 repos~349Automated safety check: PassApache-2.0
Adapt New Diffusion Modelintel/auto-round1.6k—~2.8kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    464 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

    1.6k GitHub stars~2.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Metax Model Upgrade

    MetaX-MACA/vLLM-metax

    Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

    179 GitHub stars~2.6k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-omni

All 20 skills in this repo
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    Auto-check passed
  • Precheck PR

    vllm-project/vllm-omni

    Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Quantization

    vllm-project/vllm-omni

    Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Review PR

    vllm-project/vllm-omni

    Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.

    7.1k GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • H3 Prompt Writing

    vllm-project/vllm-omni

    Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.

    7.1k GitHub starsUsed in 6 repos~744 tokens
    Auto-check passed
  • Add Recipe

    vllm-project/vllm-omni

    Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Questions about Add Diffusion Model

What does Add Diffusion Model do?

Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…. Add Diffusion Model is an agent skill from vllm-project/vllm-omni. Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT, offload, and parallelism support (TP, SP/USP, CFG-Parallel, HSDP).

When should I use Add Diffusion Model?

Add Diffusion Model fits situations like: reviewing a new diffusion model; porting a Diffusers pipeline; custom model repository; creating a DiT adapter.

How do I install Add Diffusion Model in Claude Code?

Run `npx skills add vllm-project/vllm-omni --skill add-diffusion-model -a claude-code`. Or copy the skill folder (.claude/skills/add-diffusion-model in vllm-project/vllm-omni) into .claude/skills/add-diffusion-model in your project. Claude Code loads it when a task matches its description.

How do I install Add Diffusion Model in Codex?

Run `npx skills add vllm-project/vllm-omni --skill add-diffusion-model -a codex`. Or copy the skill folder (.claude/skills/add-diffusion-model in vllm-project/vllm-omni) into .agents/skills/add-diffusion-model in your project. Codex loads it when a task matches its description.

Can I use Add Diffusion Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill add-diffusion-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-diffusion-model, .gemini/skills/add-diffusion-model, .github/skills/add-diffusion-model and .opencode/skills/add-diffusion-model in your project.

What does Add Diffusion Model need to run?

SKILL.md names no scripts, command-line tools or credentials: Add Diffusion Model is instructions for the agent only. Our summary lists: Python 3.

Does Add Diffusion Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add Diffusion Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Diffusion Model use?

Add Diffusion Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Diffusion Model use?

About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Add Diffusion Model?

Skills that share tags, products or a category with Add Diffusion Model: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 464 stars) and Adapt New Diffusion Model (intel/auto-round, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Diffusion Model?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,072 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.

Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.