Official agent skill

Add Export Format

by intel in intel/auto-round

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Add Export Format

skills CLI
$ npx skills add intel/auto-round --skill add-export-format -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intel/auto-round add-export-format --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-export-format .claude/skills/add-export-format && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-export-format
GitHub stars
1.6k
Token cost
~1.9k tokens
SKILL.md length
306 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

  • Works in 6 steps: Create Export Module Directory → Implement the Export Logic → Register the Format → …
  • Implementing a new quantized model serialization format
  • SKILL.md covers Overview, Prerequisites, Step 1: Create Export Module… and Step 2: Implement the Export…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Add Export Format is an agent skill from intel/auto-round, published by the product's own GitHub organization. Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor). Use when implementing a new quantized model serialization format, adding a new packing method, or extending export compatibility for deployment frameworks like vLLM, SGLang, or llama.cpp.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp, SGLang and vLLM. The repository describes itself as: A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包. The licence is Apache-2.0.

When your agent uses it

  • Implementing a new quantized model serialization format
  • Adding a new packing method
  • Extending export compatibility for deployment frameworks like vLLM

Example prompts

  • “/add-export-format”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Create Export Module Directory
  2. Implement the Export Logic
  3. Register the Format
  4. Update SUPPORTED_FORMATS
  5. Wire Up Backend Info (If Needed)
  6. Test

What it can do on your machine

Read from SKILL.md and the folder at commit 6afaecd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Export Format loads about 1.9k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 306 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intel/auto-round at commit 6afaecd, republished under its Apache-2.0 licence (© intel). 306 words, ~1,906 tokens.

Download SKILL.mdSave it as .claude/skills/add-export-format/SKILL.md (or your agent's skills folder).
name
add-export-format
description
Add a new model export format to AutoRound (e.g., auto_round, auto_gptq, auto_awq, gguf, llm_compressor). Use when implementing a new quantized model serialization format, adding a new packing method, or extending export compatibility for deployment frameworks like vLLM, SGLang, or llama.cpp.

Adding a New Export Format to AutoRound

Overview

This skill guides you through adding a new export format for saving quantized models. An export format defines how quantized weights, scales, and zero-points are packed and serialized for deployment. Each format is registered via the @OutputFormat.register() decorator in auto_round/formats.py.

Prerequisites

Before starting, determine:

  1. Target deployment framework: vLLM, llama.cpp, Transformers, SGLang, etc.
  2. Packing scheme: How quantized weights are packed (e.g., INT32 packing, safetensors, GGUF binary)
  3. Supported quantization schemes: Which bit-widths, data types, and configs are compatible
  4. Config format: How quantization metadata is stored (e.g., quantize_config.json, GGUF metadata)

Step 1: Create Export Module Directory

Create a new directory:

auto_round/export/export_to_yourformat/
├── __init__.py
└── export.py

Step 2: Implement the Export Logic

In export.py, implement two core functions:

pack_layer()

Packs a single quantized layer's weights, scales, and zero-points:

python
def pack_layer(layer_name, model, backend, output_dtype=torch.float16):
    """Pack a quantized layer for serialization.

    Args:
        layer_name: Full module path (e.g., "model.layers.0.self_attn.q_proj")
        model: The quantized model
        backend: Backend configuration string
        output_dtype: Output tensor dtype

    Returns:
        dict: Packed tensors ready for serialization
    """
    layer = get_module(model, layer_name)
    device = layer.weight.device

    # Get quantization parameters from layer
    bits = layer.bits
    group_size = layer.group_size
    scale = layer.scale
    zp = layer.zp
    weight = layer.weight

    # Pack weights according to your format
    packed_weight = _pack_weights(weight, bits, group_size)

    return {
        f"{layer_name}.qweight": packed_weight,
        f"{layer_name}.scales": scale,
        f"{layer_name}.qzeros": zp,
    }
save_quantized_as_yourformat()

Saves the complete quantized model:

python
def save_quantized_as_yourformat(output_dir, model, tokenizer, layer_config, serialization_dict=None, **kwargs):
    """Save quantized model in your format.

    Args:
        output_dir: Directory to save to
        model: The quantized model
        tokenizer: Model tokenizer
        layer_config: Per-layer quantization configuration
        serialization_dict: Pre-packed layer tensors (optional)
        **kwargs: Additional format-specific arguments
    """
    import os
    from safetensors.torch import save_file

    os.makedirs(output_dir, exist_ok=True)

    # 1. Pack all quantized layers (if not pre-packed)
    if serialization_dict is None:
        serialization_dict = {}
        for layer_name, config in layer_config.items():
            serialization_dict.update(pack_layer(layer_name, model, ...))

    # 2. Save weights
    save_file(serialization_dict, os.path.join(output_dir, "model.safetensors"))

    # 3. Save quantization config
    quant_config = {
        "quant_method": "yourformat",
        "bits": ...,
        "group_size": ...,
        # format-specific metadata
    }
    # Write config to output_dir

    # 4. Save tokenizer
    tokenizer.save_pretrained(output_dir)

Step 3: Register the Format

Create the OutputFormat subclass in auto_round/formats.py:

python
@OutputFormat.register("yourformat")
class YourFormat(OutputFormat):
    format_name = "yourformat"
    support_schemes = ["W4A16", "W8A16"]  # List supported scheme names

    def __init__(self, format: str, ar):
        super().__init__(format, ar)

    @classmethod
    def check_scheme_args(cls, scheme: QuantizationScheme) -> bool:
        """Check if a QuantizationScheme is compatible with this format."""
        return scheme.bits in [4, 8] and scheme.data_type == "int" and scheme.act_bits >= 16

    def pack_layer(self, layer_name, model, output_dtype=torch.float16):
        from auto_round.export.export_to_yourformat.export import pack_layer

        return pack_layer(layer_name, model, self.get_backend_name(), output_dtype)

    def save_quantized(self, output_dir, model, tokenizer, layer_config, serialization_dict=None, **kwargs):
        from auto_round.export.export_to_yourformat.export import save_quantized_as_yourformat

        return save_quantized_as_yourformat(
            output_dir, model, tokenizer, layer_config, serialization_dict=serialization_dict, **kwargs
        )

Step 4: Update SUPPORTED_FORMATS

Update the supported-format registry in auto_round/utils/common.py so your format appears in CLI help and validation.

In this repository, SUPPORTED_FORMATS is a SupportedFormats object, not a plain list. Add your format string to the _support_format tuple inside SupportedFormats.__init__():

python
class SupportedFormats:
    def __init__(self):
        self._support_format = (
            "auto_round",
            "auto_gptq",
            # ...
            "yourformat",  # Add your format here
        )

SUPPORTED_FORMATS = SupportedFormats() is then built from that tuple (plus GGUF-derived formats), so contributors should modify the registry definition, not treat SUPPORTED_FORMATS itself as a mutable list.

Step 5: Wire Up Backend Info (If Needed)

If your format requires specific inference backends, register them in auto_round/inference/backend.py:

python
BackendInfos["auto_round:yourformat"] = BackendInfo(
    device=["cuda"],
    sym=[True, False],
    packing_format=["yourformat"],
    bits=[4, 8],
    group_size=[32, 64, 128],
    priority=2,
)

Step 6: Test

python
def test_yourformat_export(tiny_opt_model_path, dataloader):
    ar = AutoRound(
        tiny_opt_model_path,
        bits=4,
        group_size=128,
        dataset=dataloader,
        iters=2,
        nsamples=2,
    )
    compressed_model, _ = ar.quantize()
    ar.save_quantized(output_dir="./tmp_yourformat", format="yourformat")

    # Verify saved files exist
    assert os.path.exists("./tmp_yourformat/model.safetensors")

    # Verify model can be loaded back
    from transformers import AutoModelForCausalLM

    loaded = AutoModelForCausalLM.from_pretrained("./tmp_yourformat")

Reference: Existing Export Format Implementations

DirectoryFormat NameKey Patterns
export_to_autoround/auto_roundNative format, QuantLinear packing, safetensors
export_to_autogptq/auto_gptqGPTQ-compatible INT packing
export_to_awq/auto_awqAWQ-compatible format
export_to_gguf/ggufBinary GGUF format with super-block quantization, uses @register_qtype()
export_to_llmcompressor/llm_compressorCompressedTensors format for vLLM

Key Registration Points

WhatWhereMechanism
Format classauto_round/formats.py@OutputFormat.register("name")
Support matrixOutputFormat.support_schemesClass attribute list
Backend infoauto_round/inference/backend.pyBackendInfos["name"] dict
CLI format registryauto_round/utils/common.pySupportedFormats._support_format tuple

© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/add-export-format of intel/auto-round.

Open the folder on GitHubat commit 6afaecd

Compare with similar skills

Add Export Format next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Export Format compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Export Format this skillintel/auto-round1.6k—~1.9kAutomated safety check: PassApache-2.0
Jetson Inference Mem TuneNVIDIA/skills3.5k1 repos~2.9kAutomated safety check: PassApache-2.0
Agentsop LLM Engine Selectionagentsope/SkillAlchemy457—~6.1kAutomated safety check: PassMIT
Agentsop Vllmagentsope/SkillAlchemy457—~6.1kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT

Similar skills

  • Official

    Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

    3.5k GitHub starsUsed in 1 repo~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Agentsop LLM Engine Selection

    agentsope/SkillAlchemy

    Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

    457 GitHub stars~6.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agentsop Vllm

    agentsope/SkillAlchemy

    Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

    457 GitHub stars~6.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from intel/auto-round

  • Official

    Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

    1.6k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Adapt New LLM

    intel/auto-round

    Official

    Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

    1.6k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Add Inference Backend

    intel/auto-round

    Official

    Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

    1.6k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Official

    Add a new quantization data type to AutoRound (e.g., INT, FP8, MXFP, NVFP, GGUF variants).

    1.6k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Add Vlm Model

    intel/auto-round

    Official

    Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

    1.6k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Review PR

    intel/auto-round

    Official

    Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files…

    1.6k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed

Questions about Add Export Format

What does Add Export Format do?

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor). Add Export Format is an agent skill from intel/auto-round, published by the product's own GitHub organization., autoround, autogptq, autoawq, gguf, llmcompressor).

When should I use Add Export Format?

Add Export Format fits situations like: implementing a new quantized model serialization format; adding a new packing method; extending export compatibility for deployment frameworks like vLLM.

How do I install Add Export Format in Claude Code?

Run `npx skills add intel/auto-round --skill add-export-format -a claude-code`. Or copy the skill folder (.claude/skills/add-export-format in intel/auto-round) into .claude/skills/add-export-format in your project. Claude Code loads it when a task matches its description.

How do I install Add Export Format in Codex?

Run `npx skills add intel/auto-round --skill add-export-format -a codex`. Or copy the skill folder (.claude/skills/add-export-format in intel/auto-round) into .agents/skills/add-export-format in your project. Codex loads it when a task matches its description.

Can I use Add Export Format in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/auto-round --skill add-export-format -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-export-format, .gemini/skills/add-export-format, .github/skills/add-export-format and .opencode/skills/add-export-format in your project.

What does Add Export Format need to run?

SKILL.md names no scripts, command-line tools or credentials: Add Export Format is instructions for the agent only. Our summary lists: Python 3.

Does Add Export Format access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add Export Format safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Export Format use?

Add Export Format is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Export Format use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Export Format?

Skills that share tags, products or a category with Add Export Format: Jetson Inference Mem Tune (NVIDIA/skills, 3.5k stars), Agentsop LLM Engine Selection (agentsope/SkillAlchemy, 457 stars), Agentsop Vllm (agentsope/SkillAlchemy, 457 stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Export Format?

intel (a GitHub organization, an official publisher) maintains it in intel/auto-round, which has 1,628 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 5, 2026.

Source: intel/auto-round on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.