Official agent skill

Add Vlm Model

by intel in intel/auto-round

Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Add Vlm Model

skills CLI
$ npx skills add intel/auto-round --skill add-vlm-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intel/auto-round add-vlm-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-vlm-model .claude/skills/add-vlm-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-vlm-model
GitHub stars
1.6k
Token cost
~2.4k tokens
SKILL.md length
564 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

  • Works in 7 steps: Add Multimodal Block Handler → Wire MLLM Calibration → Add Calibration Template → …
  • Integrating a new VLM like LLaVA
  • SKILL.md covers Overview, Prerequisites, Step 1: Add Multimodal Block… and Step 2: Wire MLLM Calibration, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Add Vlm Model is an agent skill from intel/auto-round, published by the product's own GitHub organization. Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling. Use when integrating a new VLM like LLaVA, Qwen2-VL, GLM-Image, Phi-Vision, or similar multi-modal models for quantization.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving, Performance reviews and Computer vision. The repository describes itself as: A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包. The licence is Apache-2.0.

When your agent uses it

  • Integrating a new VLM like LLaVA
  • Similar multi-modal models for quantization

Example prompts

  • “/add-vlm-model”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Add Multimodal Block Handler
  2. Wire MLLM Calibration
  3. Add Calibration Template
  4. Handle Special Forward Pass (If Needed)
  5. Add Custom Calibration Dataset (Optional)
  6. Test
  7. Update Documentation

What it can do on your machine

Read from SKILL.md and the folder at commit ae21ef9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Vlm Model loads about 2.4k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 564 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intel/auto-round at commit ae21ef9, republished under its Apache-2.0 licence (© intel). 564 words, ~2,438 tokens.

Download SKILL.mdSave it as .claude/skills/add-vlm-model/SKILL.md (or your agent's skills folder).
name
add-vlm-model
description
Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling. Use when integrating a new VLM like LLaVA, Qwen2-VL, GLM-Image, Phi-Vision, or similar multi-modal models for quantization.

Adding a New Vision-Language Model to AutoRound

Overview

This skill guides you through adding support for a new Vision-Language Model (VLM) to AutoRound. VLMs require special handling because they typically have separate vision encoder and language model components, and calibration may need multi-modal data.

The integration involves three parts:

  1. Multimodal Block Handler — Tell AutoRound how to find quantizable blocks
  2. MLLM Calibration Path — Ensure MLLMCalibrator can build and feed calibration samples
  3. Special Model Handler — Handle model-specific forward pass quirks

Prerequisites

Before starting, determine:

  1. Model architecture: What sub-modules exist? (vision encoder, projector, language model, audio tower, etc.)
  2. Model type: The model_type string from config.json
  3. Block structure: Where are the transformer layers? (e.g., model.layers, thinker.model.layers, language_model.layers)
  4. Text-only support: Can the model be calibrated with text-only data?
  5. Batch size limitations: Does the VLM have restrictions on batch size?

Step 1: Add Multimodal Block Handler

Edit auto_round/special_model_handler.py:

1a. Create a block discovery function
python
def _get_your_vlm_multimodal_block(model, quant_vision=False):
    """Get block names for YourVLM model.

    YourVLM structure:
    - model.vision_encoder.blocks: vision encoder
    - model.projector.layers: vision-language projector
    - model.language_model.layers: text decoder

    By default, only the text decoder is quantized. Set quant_vision=True
    to include vision encoder and projector blocks.
    """
    block_names = []

    if quant_vision:
        if hasattr(model, "model") and hasattr(model.model, "vision_encoder"):
            if hasattr(model.model.vision_encoder, "blocks"):
                block_names.append(
                    [f"model.vision_encoder.blocks.{i}" for i in range(len(model.model.vision_encoder.blocks))]
                )
        # Add projector if it has quantizable layers
        if hasattr(model, "model") and hasattr(model.model, "projector"):
            if hasattr(model.model.projector, "layers"):
                block_names.append([f"model.projector.layers.{i}" for i in range(len(model.model.projector.layers))])

    # Language model layers (always quantized)
    if hasattr(model, "model") and hasattr(model.model, "language_model"):
        if hasattr(model.model.language_model, "layers"):
            block_names.append(
                [f"model.language_model.layers.{i}" for i in range(len(model.model.language_model.layers))]
            )

    return block_names
1b. Register in the SPECIAL_MULTIMODAL_BLOCK dict

Find the SPECIAL_MULTIMODAL_BLOCK dictionary (in special_model_handler.py) and add your model:

python
SPECIAL_MULTIMODAL_BLOCK["your_vlm"] = _get_your_vlm_multimodal_block

The key must match the model_type from the model's config.json.

1c. Add to support lists
python
# If your VLM supports text-only calibration (most do):
SUPPORT_ONLY_TEXT_MODELS.append("your_vlm")

# If your VLM has batch size limitations:
mllms_with_limited_bs = (
    ...,
    "your_vlm",
)

Step 2: Wire MLLM Calibration

The new architecture routes multimodal calibration through:

  • auto_round/compressors/mllm_mixin.py for compressor construction and calibrator selection
  • auto_round/calibration/mllm.py for template selection, dataloader creation, and calibration forward calls
  • auto_round/special_model_handler.py for multimodal block discovery and special forwards

If your model works with an existing template/processor, prefer passing template=..., processor=..., or image_processor=... directly through AutoRound kwargs instead of adding compressor code.

Step 3: Add Calibration Template

The built-in MLLM template and processor registries live in auto_round/compressors/mllm/ and are consumed by the new architecture through MLLMCalibrator. When adding a new built-in template, keep the new-architecture caller in mind: auto_round/calibration/mllm.py will load it via get_template().

3a. Create template JSON

Create a template JSON file in auto_round/compressors/mllm/templates/:

json
{
    "model_type": "your_vlm",
    "format_user": "<|user|>\n{content}\n",
    "format_assistant": "<|assistant|>\n{content}\n",
    "format_system": "<|system|>\n{content}\n",
    "format_observation": "",
    "system": "",
    "separator": "",
    "stop_words": ["<|end|>"]
}

Adjust the template fields to match your model's chat format. Check the model's tokenizer_config.json or documentation for the correct chat template.

3b. Register the template

Register it in the MLLM template registry loaded by auto_round/calibration/mllm.py:

python
_register_template(
    "your_vlm",
    default_dataset="liuhaotian/llava_conv_58k",  # or appropriate dataset
    processor=PROCESSORS["default"],  # or a custom processor
)
Show full SKILL.md (229 more words)Show less
3c. Add a custom processor (if needed)

If your model requires special image/prompt processing for calibration, create a processor in auto_round/compressors/mllm/processor.py, which is used by MLLMCalibrator:

python
def _your_vlm_processor(raw_data, model_path, seqlen, processor=None, **kwargs):
    """Process calibration data for YourVLM.

    Args:
        raw_data: Dataset samples
        model_path: Path to the model
        seqlen: Sequence length for calibration
        processor: The model's processor

    Returns:
        list: Processed samples ready for calibration
    """
    # Build prompts with images and text
    ...

Register it:

python
PROCESSORS["your_vlm"] = _your_vlm_processor

Step 4: Handle Special Forward Pass (If Needed)

If your VLM's forward() method is non-standard (e.g., requires special kwargs, has multiple model components that need separate handling), add a custom forward wrapper in special_model_handler.py:

python
def _your_vlm_forward(model, **kwargs):
    """Custom forward pass for YourVLM during calibration."""
    # Handle special input processing
    # Route inputs to correct sub-models
    return model.language_model(**kwargs)

Register it in _handle_special_model():

python
def _handle_special_model(model):
    ...
    if hasattr(model, "config") and model.config.model_type == "your_vlm":
        from functools import partial

        model.forward = partial(_your_vlm_forward, model)
    return model

Step 5: Add Custom Calibration Dataset (Optional)

If your model needs a specialized calibration dataset loader, create one in auto_round/calib_dataset.py using the @register_dataset decorator:

python
@register_dataset("your_vlm_dataset")
class YourVLMDataset:
    def __init__(self, dataset_name, model_path, seqlen, **kwargs): ...

    def __len__(self):
        return len(self.data)

    def __iter__(self):
        for sample in self.data:
            yield sample

Step 6: Test

python
def test_your_vlm_quantization():
    model_name = "your-org/your-vlm-small"
    ar = AutoRound(
        model_name,
        bits=4,
        group_size=128,
        iters=2,
        nsamples=2,
        quant_nontext_module=False,  # text-only quantization
    )
    compressed_model, _ = ar.quantize()
    ar.save_quantized(output_dir="./tmp_your_vlm", format="auto_round")

Test with vision quantization:

python
ar = AutoRound(
    model_name,
    bits=4,
    group_size=128,
    quant_nontext_module=True,  # also quantize vision encoder
)

Step 7: Update Documentation

  1. Add your model to the supported VLM list in README.md
  2. Update README_CN.md with the same changes (Chinese translation required)
  3. Add example quantization script if the model has special usage patterns

Reference: Existing VLM Implementations

Model TypeBlock HandlerTemplateSpecial Forward
llava_get_llava_multimodal_blockllava templateNo
qwen2_vl_get_qwen2_vl_multimodal_blockqwen2_vl templateNo
qwen2_5_omni_get_qwen2_5_omni_multimodal_blockqwen2_5_omni templateYes (_qwen2_5_omni_forward)
qwen3_omni_moe_get_qwen3_omni_moe_multimodal_blockqwen3_omni_moe templateYes (_qwen3_omni_moe_forward)
deepseek_vl_v2_get_deepseek_vl2_multimodal_blockdeepseek_vl_v2 templateYes (_deepseek_vl2_forward)
glm_image_get_glm_image_multimodal_blockglm_image templateNo
phi3_vvia generic handlerphi3_v templateNo

Key Registration Points

WhatWhereMechanism
Block handlerspecial_model_handler.pySPECIAL_MULTIMODAL_BLOCK[model_type]
Text-only supportspecial_model_handler.pySUPPORT_ONLY_TEXT_MODELS list
Batch limitspecial_model_handler.pymllms_with_limited_bs tuple
MLLM routingcompressors/mllm_mixin.py_get_calibrator_kind() -> "mllm"
MLLM calibrationcalibration/mllm.pyMLLMCalibrator.calib()
Templatecompressors/mllm/template.py_register_template()
Processorcompressors/mllm/processor.pyPROCESSORS dict
Custom forwardspecial_model_handler.py_handle_special_model()
Dataset loadercalib_dataset.py@register_dataset()

© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/add-vlm-model of intel/auto-round.

Open the folder on GitHubat commit ae21ef9

Compare with similar skills

Add Vlm Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Vlm Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Vlm Model this skillintel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0
Quark Onnx Ptq Workflowamd/Quark181—~4.5kAutomated safety check: PassMIT
Astreawarpfront/hipfire653—~2.6kAutomated safety check: PassCustom licence
Visiongridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0
Robot Perceptionarpitg1304/robotics-agent-skills368—~15kAutomated safety check: PassApache-2.0
Quark Onnx Autosearch Proamd/Quark181—~3.4kAutomated safety check: PassMIT

Similar skills

  • End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output.

    181 GitHub stars~4.5k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Astrea

    warpfront/hipfire

    A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…

    653 GitHub stars~2.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Robot Perception

    arpitg1304/robotics-agent-skills

    Comprehensive best practices for robot perception systems covering cameras, LiDARs, depth sensors, IMUs, and multi-sensor setups.

    368 GitHub stars~15k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Tao Train Rtdetr

    NVIDIA/skills

    Official

    RT-DETR (Real-Time DEtection TRansformer) for 2D object detection.

    3.5k GitHub stars~4.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from intel/auto-round

  • Official

    Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

    1.6k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Adapt New LLM

    intel/auto-round

    Official

    Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

    1.6k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Add Export Format

    intel/auto-round

    Official

    Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

    1.6k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Add Inference Backend

    intel/auto-round

    Official

    Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Official

    Add a new quantization data type to AutoRound (e.g., INT, FP8, MXFP, NVFP, GGUF variants).

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Review PR

    intel/auto-round

    Official

    Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files…

    1.6k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about Add Vlm Model

What does Add Vlm Model do?

Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling. Add Vlm Model is an agent skill from intel/auto-round, published by the product's own GitHub organization. Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

When should I use Add Vlm Model?

Add Vlm Model fits situations like: integrating a new VLM like LLaVA; similar multi-modal models for quantization.

How do I install Add Vlm Model in Claude Code?

Run `npx skills add intel/auto-round --skill add-vlm-model -a claude-code`. Or copy the skill folder (.claude/skills/add-vlm-model in intel/auto-round) into .claude/skills/add-vlm-model in your project. Claude Code loads it when a task matches its description.

How do I install Add Vlm Model in Codex?

Run `npx skills add intel/auto-round --skill add-vlm-model -a codex`. Or copy the skill folder (.claude/skills/add-vlm-model in intel/auto-round) into .agents/skills/add-vlm-model in your project. Codex loads it when a task matches its description.

Can I use Add Vlm Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/auto-round --skill add-vlm-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-vlm-model, .gemini/skills/add-vlm-model, .github/skills/add-vlm-model and .opencode/skills/add-vlm-model in your project.

What does Add Vlm Model need to run?

SKILL.md names no scripts, command-line tools or credentials: Add Vlm Model is instructions for the agent only. Our summary lists: Python 3.

Does Add Vlm Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add Vlm Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Vlm Model use?

Add Vlm Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Vlm Model use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Vlm Model?

Skills that share tags, products or a category with Add Vlm Model: Quark Onnx Ptq Workflow (amd/Quark, 181 stars), Astrea (warpfront/hipfire, 653 stars), Vision (gridaco/grida, 2.7k stars) and Robot Perception (arpitg1304/robotics-agent-skills, 368 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Vlm Model?

intel (a GitHub organization, an official publisher) maintains it in intel/auto-round, which has 1,628 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.

Source: intel/auto-round on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.