Official agent skill

Adapt New Diffusion Model

by intel in intel/auto-round

Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Adapt New Diffusion Model

skills CLI
$ npx skills add intel/auto-round --skill adapt-new-diffusion-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intel/auto-round adapt-new-diffusion-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/adapt-new-diffusion-model .claude/skills/adapt-new-diffusion-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
adapt-new-diffusion-model
GitHub stars
1.6k
Token cost
~2.8k tokens
SKILL.md length
731 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

  • Works in 7 steps: Diagnose the Problem → Ensure Model Detection → Register Transformer Block Output Config → …
  • A new diffusion model fails quantization
  • SKILL.md covers Overview, Step 0: Diagnose the Problem, Step 1: Ensure Model Detection and Step 2: Register Transformer…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Adapt New Diffusion Model is an agent skill from intel/auto-round, published by the product's own GitHub organization. Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT). Use when a new diffusion model fails quantization, needs custom output configs, requires a custom pipeline function, or is a hybrid architecture with both autoregressive and diffusion components.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Diffusion and image models and LLM inference and serving. The repository describes itself as: A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包. The licence is Apache-2.0.

When your agent uses it

  • A new diffusion model fails quantization
  • Needs custom output configs
  • Requires a custom pipeline function
  • Is a hybrid architecture with both autoregressive and diffusion components

Example prompts

  • “/adapt-new-diffusion-model”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Diagnose the Problem
  2. Ensure Model Detection
  3. Register Transformer Block Output Config
  4. Handle Non-Standard Pipeline API
  5. Add Hybrid AR+DiT Support
  6. Add Custom Calibration Dataset (Optional)
  7. Test

What it can do on your machine

Read from SKILL.md and the folder at commit ae21ef9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Adapt New Diffusion Model loads about 2.8k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 731 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intel/auto-round at commit ae21ef9, republished under its Apache-2.0 licence (© intel). 731 words, ~2,825 tokens.

Download SKILL.mdSave it as .claude/skills/adapt-new-diffusion-model/SKILL.md (or your agent's skills folder).
name
adapt-new-diffusion-model
description
Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT). Use when a new diffusion model fails quantization, needs custom output configs, requires a custom pipeline function, or is a hybrid architecture with both autoregressive and diffusion components.

Adapting AutoRound for a New Diffusion Model Architecture

Overview

AutoRound's new diffusion path uses auto_round/compressors/diffusion_mixin.py, auto_round/calibration/diffusion.py, and the quantizer implementations under auto_round/algorithms/quantization/. This skill covers what code changes are needed when a new diffusion model doesn't work out-of-the-box. Common reasons for adaptation:

  • Transformer block type not registered in DIFFUSION_OUTPUT_CONFIGS
  • Non-standard pipeline API (not compatible with pipe(prompts, ...))
  • Hybrid architecture with both AR and diffusion components
  • Model not detected as a diffusion model

Step 0: Diagnose the Problem

python
from auto_round import AutoRound

ar = AutoRound(
    "your-org/your-diffusion-model",
    scheme="W4A16",
    iters=2,
    nsamples=2,
    num_inference_steps=5,
)
ar.quantize_and_save(output_dir="./test_output", format="fake")
Error / SymptomRoot CauseFix Section
"using LLM mode" instead of DiffusionModel not detected as diffusionStep 1
assert len(output_config) == len(tmp_output)Block output config mismatchStep 2
Pipeline call failsNon-standard inference APIStep 3
Hybrid model only quantizes DiTAR component not handledStep 4

Step 1: Ensure Model Detection

AutoRound detects diffusion models by checking for model_index.json in the model directory:

python
# auto_round/utils/model.py
def is_diffusion_model(model_or_path):
    # Checks for model_index.json presence

If your model doesn't have model_index.json, either create one in the model directory or pass diffusion-specific options through new-architecture AutoRound kwargs:

python
ar = AutoRound(
    model,
    num_inference_steps=5,
)
Pipeline Loading

diffusion_load_model() uses AutoPipelineForText2Image.from_pretrained() and extracts pipe.transformer as the quantizable model. If your model uses a different attribute (e.g., pipe.unet), this needs adjustment in auto_round/utils/model.py.

Step 2: Register Transformer Block Output Config

This is the most common adaptation needed. DIFFUSION_OUTPUT_CONFIGS maps transformer block class names to their output tensor names. Without this, calibration crashes because AutoRound doesn't know how to collect activations.

Find your block class name
python
import diffusers

pipe = diffusers.AutoPipelineForText2Image.from_pretrained("your-model")
for name, module in pipe.transformer.named_modules():
    if hasattr(module, "forward") and "block" in name.lower():
        print(f"{name}: {type(module).__name__}")
Register in DIFFUSION_OUTPUT_CONFIGS

Edit auto_round/algorithms/quantization/base.py:

python
class BaseQuantizers:
    DIFFUSION_OUTPUT_CONFIGS = {
        "FluxTransformerBlock": ["encoder_hidden_states", "hidden_states"],
        "FluxSingleTransformerBlock": ["encoder_hidden_states", "hidden_states"],
        # Add your block type:
        "YourTransformerBlock": ["hidden_states"],  # output tensor names in order
    }

The list must match the exact order of tensors returned by the block's forward() method.

How to determine output tensor names
  1. Read the block's forward() method in diffusers source code
  2. Identify what tensors it returns (usually hidden_states, sometimes also encoder_hidden_states)
  3. List them in the order they're returned

Example: If forward() returns (hidden_states, encoder_hidden_states):

python
BaseQuantizers.DIFFUSION_OUTPUT_CONFIGS["YourBlock"] = ["hidden_states", "encoder_hidden_states"]

Example: If forward() returns just hidden_states:

python
BaseQuantizers.DIFFUSION_OUTPUT_CONFIGS["YourBlock"] = ["hidden_states"]

Step 3: Handle Non-Standard Pipeline API

If your model's inference API differs from the standard pipe(prompts, guidance_scale=..., num_inference_steps=...), provide a custom pipeline function.

Option A: Add a custom pipeline dispatch in DiffusionCalibrator

Update auto_round/calibration/diffusion.py so DiffusionCalibrator.calib() dispatches through a small helper instead of calling pipe(...) directly:

python
class DiffusionCalibrator(LLMCalibrator):
    ...

    def _run_pipeline(self, pipe, prompts, generator):
        if getattr(pipe, "_autoround_pipeline_fn", None) is not None:
            pipe._autoround_pipeline_fn(
                pipe,
                prompts,
                guidance_scale=self.compressor.guidance_scale,
                num_inference_steps=self.compressor.num_inference_steps,
                generator=generator,
            )
            return
        pipe(
            prompts,
            guidance_scale=self.compressor.guidance_scale,
            num_inference_steps=self.compressor.num_inference_steps,
            generator=generator,
        )
Option B: Attach a model-specific function during model loading

For a known model family, attach _autoround_pipeline_fn in auto_round/utils/model.py or auto_round/special_model_handler.py:

python
pipe._autoround_pipeline_fn = your_model_pipeline_fn
Option C: Add a dedicated branch in DiffusionCalibrator

For full control, update auto_round/calibration/diffusion.py so DiffusionCalibrator.calib() dispatches through your custom pipeline function:

python
class DiffusionCalibrator(LLMCalibrator):
    ...

    def _run_pipeline(self, pipe, prompts):
        c = self.compressor
        generator = (
            None if c.generator_seed is None else torch.Generator(device=pipe.device).manual_seed(c.generator_seed)
        )
        pipe.your_custom_generate(
            prompts,
            steps=c.num_inference_steps,
            cfg=c.guidance_scale,
            generator=generator,
        )

Step 4: Add Hybrid AR+DiT Support

For models with both autoregressive and diffusion components (e.g., GLM-Image).

4a. Register AR component

Add hybrid routing through the new architecture. Start with auto_round/autoround.py, auto_round/compressors/entry.py, and auto_round/compressors/diffusion_mixin.py. If a reusable AR-component registry is needed, place it near the new routing code:

python
HYBRID_AR_COMPONENTS = [
    "vision_language_encoder",  # GLM-Image
    "your_ar_component",  # Your model's AR attribute name
]

The attribute name must match what exists on the diffusers pipeline object (i.e., pipe.your_ar_component).

Show full SKILL.md (287 more words)Show less
4b. Register DiT block output config

Add the DiT-specific output config in BaseQuantizers.DIFFUSION_OUTPUT_CONFIGS:

python
BaseQuantizers.DIFFUSION_OUTPUT_CONFIGS["YourDiTBlock"] = ["hidden_states", "encoder_hidden_states"]
4c. Register AR block handler

In auto_round/special_model_handler.py, add a block handler for the AR component so AutoRound knows which layers to quantize:

python
def _get_your_hybrid_multimodal_block(model, quant_vision=False):
    block_names = []
    if quant_vision and hasattr(model, "vision_encoder"):
        block_names.append([f"vision_encoder.blocks.{i}" for i in range(len(model.vision_encoder.blocks))])
    block_names.append([f"language_model.layers.{i}" for i in range(len(model.language_model.layers))])
    return block_names


SPECIAL_MULTIMODAL_BLOCK["your_model_type"] = _get_your_hybrid_multimodal_block
Hybrid quantization flow

The new hybrid flow should run two phases:

  1. Phase 1 (AR): Quantizes the AR component using text calibration data (MLLM-style)
  2. Phase 2 (DiT): Quantizes the DiT component using diffusion pipeline calibration
python
ar = AutoRound(
    "your-hybrid-model",
    dataset="coco2014",  # DiT calibration
    ar_dataset="NeelNanda/pile-10k",  # AR calibration
    quant_ar=True,
    quant_dit=True,
)

Step 5: Add Custom Calibration Dataset (Optional)

If your model needs a specific dataset format:

Edit the diffusion calibration path used by the new architecture:

  • auto_round/calibration/diffusion.py for how diffusion prompts are loaded and consumed
  • auto_round/calib_dataset.py for reusable dataset registration helpers
python
def get_diffusion_dataloader(dataset_name, nsamples, ...):
    # Add handling for your dataset format
    if dataset_name == "your_custom_dataset":
        return _load_your_dataset(dataset_name, nsamples)
    ...

The default coco2014 dataset works for most text-to-image models. Custom datasets need a TSV file with id and caption columns.

Step 6: Test

python
def test_your_diffusion_model():
    ar = AutoRound(
        "your-org/your-diffusion-model",
        scheme="W4A16",
        iters=2,
        nsamples=4,
        num_inference_steps=5,
        guidance_scale=7.5,
    )
    compressed_model, layer_config = ar.quantize()
    assert len(layer_config) > 0, "No layers quantized"
    ar.save_quantized(output_dir="./test_output", format="fake")

For hybrid models, test both phases:

python
ar = AutoRound(
    "your-hybrid-model",
    quant_ar=True,
    quant_dit=True,
    iters=2,
    nsamples=4,
)

Checklist

  • is_diffusion_model() detects model
  • Transformer block class name identified
  • DIFFUSION_OUTPUT_CONFIGS entry added with correct output tensor names and order
  • Pipeline runs without errors during calibration
  • Custom pipeline dispatch added in DiffusionCalibrator if non-standard API
  • For hybrid: AR component registered in HYBRID_AR_COMPONENTS
  • For hybrid: AR block handler in SPECIAL_MULTIMODAL_BLOCK
  • For hybrid: DiT output config in BaseQuantizers.DIFFUSION_OUTPUT_CONFIGS
  • Quantization produces valid layers (check "Quantized X/Y layers" log)
  • Export to fake format works
  • README.md + README_CN.md updated

Key Files

FilePurpose
auto_round/algorithms/quantization/base.pyBaseQuantizers.DIFFUSION_OUTPUT_CONFIGS
auto_round/calibration/diffusion.pyDiffusionCalibrator, pipeline-driving calibration logic
auto_round/compressors/diffusion_mixin.pyDiffusion compressor mixin and calibrator routing
auto_round/compressors/entry.pyNew-architecture AutoRoundCompatible factory routing
auto_round/utils/model.pyis_diffusion_model(), diffusion_load_model()
auto_round/special_model_handler.pyAR block handlers for hybrid models
auto_round/autoround.pyModel type routing (diffusion vs hybrid vs LLM)

Reference: Existing Adaptations

ModelTypeWhat Was Adapted
FLUX.1-devPure DiTDIFFUSION_OUTPUT_CONFIGS for FluxTransformerBlock/FluxSingleTransformerBlock
GLM-ImageHybrid AR+DiTAR routing + SPECIAL_MULTIMODAL_BLOCK + DiT DIFFUSION_OUTPUT_CONFIGS
NextStepCustom pipelinemodel-specific pipeline function attached by model handler / loader

© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/adapt-new-diffusion-model of intel/auto-round.

Open the folder on GitHubat commit ae21ef9

Compare with similar skills

Adapt New Diffusion Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Adapt New Diffusion Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Adapt New Diffusion Model this skillintel/auto-round1.6k—~2.8kAutomated safety check: PassApache-2.0
Add Diffusion Modelvllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0
Integrate Modeltryonlabs/opentryon551—~1.1kAutomated safety check: PassCustom licence
Production Add Diffusion Modelvllm-project/vllm-omni7.1k—~5.5kAutomated safety check: PassApache-2.0
Local LLM Freeartokun/comfyui-mcp795—~897Automated safety check: NotesMIT
Flux2 Lora TrainingAnastasiyaW/codex-claude-code-config154—~4.5kAutomated safety check: PassMIT

Similar skills

  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Integrate Model

    tryonlabs/opentryon

    Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner).

    551 GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Production Add Diffusion Model

    vllm-project/vllm-omni

    Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.

    7.1k GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Local LLM Free

    artokun/comfyui-mcp

    Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama.

    795 GitHub stars~897 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Flux2 Lora Training

    AnastasiyaW/codex-claude-code-config

    Plan or review LoRA and edit-training work specifically for FLUX.2 Klein or Qwen-Image-Edit, including paired datasets, trainer-version contracts, and held-out fidelity checks.

    154 GitHub stars~4.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed

More from intel/auto-round

  • Adapt New LLM

    intel/auto-round

    Official

    Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

    1.6k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Add Export Format

    intel/auto-round

    Official

    Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

    1.6k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Add Inference Backend

    intel/auto-round

    Official

    Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Official

    Add a new quantization data type to AutoRound (e.g., INT, FP8, MXFP, NVFP, GGUF variants).

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Add Vlm Model

    intel/auto-round

    Official

    Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

    1.6k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Review PR

    intel/auto-round

    Official

    Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files…

    1.6k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about Adapt New Diffusion Model

What does Adapt New Diffusion Model do?

Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT). Adapt New Diffusion Model is an agent skill from intel/auto-round, published by the product's own GitHub organization. Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

When should I use Adapt New Diffusion Model?

Adapt New Diffusion Model fits situations like: A new diffusion model fails quantization; needs custom output configs; requires a custom pipeline function; is a hybrid architecture with both autoregressive and diffusion components.

How do I install Adapt New Diffusion Model in Claude Code?

Run `npx skills add intel/auto-round --skill adapt-new-diffusion-model -a claude-code`. Or copy the skill folder (.claude/skills/adapt-new-diffusion-model in intel/auto-round) into .claude/skills/adapt-new-diffusion-model in your project. Claude Code loads it when a task matches its description.

How do I install Adapt New Diffusion Model in Codex?

Run `npx skills add intel/auto-round --skill adapt-new-diffusion-model -a codex`. Or copy the skill folder (.claude/skills/adapt-new-diffusion-model in intel/auto-round) into .agents/skills/adapt-new-diffusion-model in your project. Codex loads it when a task matches its description.

Can I use Adapt New Diffusion Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/auto-round --skill adapt-new-diffusion-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adapt-new-diffusion-model, .gemini/skills/adapt-new-diffusion-model, .github/skills/adapt-new-diffusion-model and .opencode/skills/adapt-new-diffusion-model in your project.

What does Adapt New Diffusion Model need to run?

SKILL.md names no scripts, command-line tools or credentials: Adapt New Diffusion Model is instructions for the agent only. Our summary lists: Python 3.

Does Adapt New Diffusion Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Adapt New Diffusion Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Adapt New Diffusion Model use?

Adapt New Diffusion Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Adapt New Diffusion Model use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Adapt New Diffusion Model?

Skills that share tags, products or a category with Adapt New Diffusion Model: Add Diffusion Model (vllm-project/vllm-omni, 7.1k stars), Integrate Model (tryonlabs/opentryon, 551 stars), Production Add Diffusion Model (vllm-project/vllm-omni, 7.1k stars) and Local LLM Free (artokun/comfyui-mcp, 795 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Adapt New Diffusion Model?

intel (a GitHub organization, an official publisher) maintains it in intel/auto-round, which has 1,628 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.

Source: intel/auto-round on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.