Agent skill

Quantization

by vllm-project in vllm-project/vllm-omni

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Quantization

skills CLI
$ npx skills add vllm-project/vllm-omni --skill quantization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-omni quantization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/quantization .claude/skills/quantization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quantization
GitHub stars
7.1k
Token cost
~1.4k tokens
SKILL.md length
560 words
Files
6 (incl. references)
Skills in repo
20
Repo updated
First seen
Licence
Apache-2.0

At a glance

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

  • Works in 6 steps: Read the local method doc under… → Inspect the active factory and config… → Inspect the target model pipeline and… → …
  • Adding methods such as fp8
  • SKILL.md covers First Checks, Quick Decision, Current Runtime Shape and Ownership Boundary, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quantization is an agent skill from vllm-project/vllm-omni. Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. Use when choosing or adding methods such as fp8, int8, gguf, mxfp8, mxfp4, mxfp4dualscale, ModelOpt, AutoRound, INC, msModelSlim, awq, or gptq; debugging quantized loading; or validating memory, speed, and output quality.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/adding-models.md`, `references/diffusion.md` and `references/methods.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and llama.cpp. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.

When your agent uses it

  • Adding methods such as fp8
  • Debugging quantized loading
  • Validating memory

Example prompts

  • “/quantization”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read the local method doc under docs/user_guide/quantization/.
  2. Inspect the active factory and config class in vllm_omni/quantization/.
  3. Inspect the target model pipeline and transformer for stable prefixes and
  4. For pre-quantized checkpoints, inspect transformer/config.json and the
  5. Validate with the same prompt, seed, size, scheduler, steps, dtype, eager or
  6. Report quality, latency, and memory separately. Do not call a method

What it can do on your machine

Read from SKILL.md and the folder at commit c548a11. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quantization loads about 1.4k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 560 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-omni at commit c548a11, republished under its Apache-2.0 licence (© vllm-project). 560 words, ~1,354 tokens.

Download SKILL.mdSave it as .claude/skills/quantization/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
quantization
description
Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. Use when choosing or adding methods such as fp8, int8, gguf, mxfp8, mxfp4, mxfp4_dualscale, ModelOpt, AutoRound, INC, msModelSlim, awq, or gptq; debugging quantized loading; or validating memory, speed, and output quality.

vLLM-Omni Quantization

Use this skill for quantization work in vllm-omni. Start from the local code and docs, then pick the smallest validated path for the model, method, hardware, and quality target.

First Checks

Before changing code or recommending a command, identify:

  • Model and task: AR text, omni/TTS, text-to-image, text-to-video, or multi-stage diffusion.
  • Hardware backend: CUDA generation, Ascend NPU, Intel XPU, or another platform.
  • Quantization mode: online from BF16/FP16 weights, pre-quantized checkpoint, runtime KV/FA quantization, or per-component routing.
  • Scope: transformer/DiT, thinker language model, text encoder, VAE, audio/vision encoder, talker, or a specific stage.
  • Validation target: memory reduction, latency/throughput, output quality, or all three.

Do not assume a method is supported just because the CLI accepts the string. Check docs/user_guide/quantization/ and the closest implementation first.

Quick Decision

TaskStart With
Choose a method or commandreferences/methods.md and references/modality-compat.md
Use build_quant_config() or per-component routingreferences/methods.md
Work on diffusion quantizationreferences/diffusion.md
Add quantization to a new modelreferences/adding-models.md
Convert or load ModelOpt FP8 checkpointsreferences/modelopt-fp8.md
Debug ModelOpt, GGUF, AutoRound, MXFP, or serialized Int8 loadingreferences/diffusion.md and method docs under docs/user_guide/quantization/

Current Runtime Shape

The unified entrypoint is:

python
from vllm_omni.quantization import build_quant_config

It supports method strings, flat method dictionaries, per-component dictionaries, existing QuantizationConfig objects, and None. The factory delegates generic methods to upstream vllm and keeps vLLM-Omni overrides for diffusion or omni-specific routing:

  • gguf
  • int8
  • mxfp8
  • mxfp4
  • mxfp4_dualscale
  • inc, auto-round, auto_round
  • ModelOpt auto-detected configs: modelopt, modelopt_fp4, modelopt_mixed

Use ComponentQuantizationConfig when only one stage or component should be quantized. Pre-quantized ModelOpt-style checkpoints should not spill into vision/audio encoders that have no corresponding scale tensors.

Ownership Boundary

  • Upstream vllm owns generic quantization configs, kernels, loader semantics, hardware capability rules, and generic AR quantization methods.
  • vllm-omni owns unified config routing, diffusion-specific wrappers, component scoping, GGUF or ModelOpt checkpoint adapters, model-specific prefix mapping, docs, examples, and validation.

If a new method needs missing generic kernels or loader behavior, fix upstream vllm first. In vllm-omni, add thin integration and model wiring.

Show full SKILL.md (243 more words)Show less

Working Loop

  1. Read the local method doc under docs/user_guide/quantization/.
  2. Inspect the active factory and config class in vllm_omni/quantization/.
  3. Inspect the target model pipeline and transformer for stable prefixes and whether every relevant vLLM linear layer receives quant_config.
  4. For pre-quantized checkpoints, inspect transformer/config.json and the checkpoint tensor names before running full generation.
  5. Validate with the same prompt, seed, size, scheduler, steps, dtype, eager or non-eager mode, and parallelism as the BF16 baseline.
  6. Report quality, latency, and memory separately. Do not call a method supported until quality is checked.

Common Mistakes

SymptomLikely CauseFix
--quantization has no visible effectWrong scope or unsupported model pathCheck component routing and method docs
Some layers stay BF16 unexpectedlyquant_config was not threaded into all vLLM linear layersAudit transformer constructors and prefixes
Quality collapses but loading succeedsToo many sensitive layers were quantizedAdd model-specific ignored_layers and compare to BF16
ModelOpt checkpoint loads but output is corruptedPrefixes, packed-module mapping, scale routing, or BF16 fallback is wrongUse references/modelopt-fp8.md
GGUF shape or tensor mismatchMissing architecture-specific adapterAdd explicit adapter mapping; avoid generic fallback
Online and offline MXFP4 disagreeWrong mxfp4 vs mxfp4_dualscale modeCheck docs/user_guide/quantization/mxfp4.md
Quantized path is slower than BF16Kernel path, compile mode, dtype, or shape mismatchCompare eager-to-eager and non-eager-to-non-eager

References

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in .claude/skills/quantization of vllm-project/vllm-omni.

  • SKILL.md
  • references/adding-models.md
  • references/diffusion.md
  • references/methods.md
  • references/modality-compat.md
  • references/modelopt-fp8.md

Open the folder on GitHubat commit c548a11

Compare with similar skills

Quantization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quantization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quantization this skillvllm-project/vllm-omni7.1k—~1.4kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k3 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
Add Export Formatintel/auto-round1.6k—~1.9kAutomated safety check: PassApache-2.0
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 3 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Export Format

    intel/auto-round

    Official

    Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

    1.6k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Outlines Structured Generation

    Orchestra-Research/AI-Research-SKILLs

    Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.

    13k GitHub starsUsed in 10 repos~4k tokens
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-omni

All 20 skills in this repo
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    Auto-check passed
  • Precheck PR

    vllm-project/vllm-omni

    Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Review PR

    vllm-project/vllm-omni

    Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.

    7.1k GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • H3 Prompt Writing

    vllm-project/vllm-omni

    Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.

    7.1k GitHub starsUsed in 6 repos~744 tokens
    Auto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    Auto-check passed
  • Add Recipe

    vllm-project/vllm-omni

    Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Questions about Quantization

What does Quantization do?

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. Quantization is an agent skill from vllm-project/vllm-omni. Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

When should I use Quantization?

Quantization fits situations like: adding methods such as fp8; debugging quantized loading; validating memory.

How do I install Quantization in Claude Code?

Run `npx skills add vllm-project/vllm-omni --skill quantization -a claude-code`. Or copy the skill folder (.claude/skills/quantization in vllm-project/vllm-omni) into .claude/skills/quantization in your project. Claude Code loads it when a task matches its description.

How do I install Quantization in Codex?

Run `npx skills add vllm-project/vllm-omni --skill quantization -a codex`. Or copy the skill folder (.claude/skills/quantization in vllm-project/vllm-omni) into .agents/skills/quantization in your project. Codex loads it when a task matches its description.

Can I use Quantization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill quantization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quantization, .gemini/skills/quantization, .github/skills/quantization and .opencode/skills/quantization in your project.

What does Quantization need to run?

SKILL.md names no scripts, command-line tools or credentials: Quantization is instructions for the agent only. Our summary lists: Python 3.

Does Quantization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quantization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quantization use?

Quantization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quantization use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.

What are the alternatives to Quantization?

Skills that share tags, products or a category with Quantization: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), Add Export Format (intel/auto-round, 1.6k stars) and Add Model (guoqingbao/xinfer, 333 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quantization?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,072 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 8, 2026.

Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.