Official agent skill

Adapt New LLM

by intel in intel/auto-round

Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Adapt New LLM

skills CLI
$ npx skills add intel/auto-round --skill adapt-new-llm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intel/auto-round adapt-new-llm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/adapt-new-llm .claude/skills/adapt-new-llm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
adapt-new-llm
GitHub stars
1.6k
Token cost
~2.4k tokens
SKILL.md length
627 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

  • Works in 8 steps: Diagnose the Problem → Fix Block Detection → Handle MoE (Mixture-of-Experts) Models → …
  • Quantization fails for a new model type
  • SKILL.md covers Overview, Step 0: Diagnose the Problem, Step 1: Fix Block Detection and Step 2: Handle MoE…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Adapt New LLM is an agent skill from intel/auto-round, published by the product's own GitHub organization. Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box. Use when quantization fails for a new model type, block detection doesn't find layers, MoE models need unfusing, custom forward passes are needed, or non-standard linear layer types need handling.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. The repository describes itself as: A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包. The licence is Apache-2.0.

When your agent uses it

  • Quantization fails for a new model type
  • Block detection doesnt find layers
  • MoE models need unfusing
  • Custom forward passes are needed

Example prompts

  • “/adapt-new-llm”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Diagnose the Problem
  2. Fix Block Detection
  3. Handle MoE (Mixture-of-Experts) Models
  4. Add Custom Forward Pass
  5. Handle Non-Standard Linear Layers
  6. Handle Shared Cache Keys
  7. Test
  8. Update Documentation

What it can do on your machine

Read from SKILL.md and the folder at commit ae21ef9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Adapt New LLM loads about 2.4k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 627 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intel/auto-round at commit ae21ef9, republished under its Apache-2.0 licence (© intel). 627 words, ~2,437 tokens.

Download SKILL.mdSave it as .claude/skills/adapt-new-llm/SKILL.md (or your agent's skills folder).
name
adapt-new-llm
description
Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box. Use when quantization fails for a new model type, block detection doesn't find layers, MoE models need unfusing, custom forward passes are needed, or non-standard linear layer types need handling.

Adapting AutoRound for a New LLM Architecture

Overview

Most standard Transformers-based LLMs work with AutoRound out-of-the-box. This skill covers what to do when a new model architecture requires code changes. The need for adaptation typically arises from:

  • Non-standard layer hierarchy (block detection fails)
  • Fused Mixture-of-Experts (MoE) weights
  • Non-standard linear layer types (not nn.Linear or Conv1D)
  • Complex multi-component architectures (multimodal routing)
  • Shared cache keys or position embeddings

Step 0: Diagnose the Problem

Try quantizing the model first:

python
from auto_round import AutoRound

ar = AutoRound("your-org/your-model", scheme="W4A16", iters=2, nsamples=2)
ar.quantize_and_save(output_dir="./test_output", format="auto_round")

Common failure modes and their fixes:

Error / SymptomRoot CauseFix Section
"No quantizable layers found"Block detection failedStep 1
"Quantized 0/N layers"Layers not nn.Linear/Conv1DStep 4
Shape mismatch in MoE layersFused expert weightsStep 2
Wrong outputs / calibration divergesForward pass not exercised correctlyStep 3
Cache key errors (Gemma3-style)Shared position embeddingsStep 5

Step 1: Fix Block Detection

AutoRound discovers quantizable blocks via get_block_names() which searches recursively for nn.ModuleList instances. If your model has a non-standard layer hierarchy, block detection may fail.

Check current detection
python
from auto_round.utils import get_block_names

model = ...  # loaded model
print(get_block_names(model))
Option A: Use to_quant_block_names parameter

For simple cases, override block names without code changes:

python
ar = AutoRound(
    model,
    to_quant_block_names="model.decoder.layers",  # explicit path
)
Option B: Register in SPECIAL_MULTIMODAL_BLOCK

For multimodal or multi-component models, add a custom block handler in auto_round/special_model_handler.py:

python
def _get_your_model_multimodal_block(model, quant_vision=False):
    """Get block names for YourModel.

    YourModel structure:
    - encoder.layers: encoder blocks
    - decoder.layers: decoder blocks
    """
    block_names = []

    if quant_vision and hasattr(model, "encoder"):
        block_names.append([f"encoder.layers.{i}" for i in range(len(model.encoder.layers))])

    block_names.append([f"decoder.layers.{i}" for i in range(len(model.decoder.layers))])

    return block_names


# Register: key must match model.config.model_type
SPECIAL_MULTIMODAL_BLOCK["your_model_type"] = _get_your_model_multimodal_block

Also add to support lists if applicable:

python
# If text-only calibration works for this multimodal model:
SUPPORT_ONLY_TEXT_MODELS.append("your_model_type")

# If batch_size must be limited:
mllms_with_limited_bs = (..., "your_model_type")

Step 2: Handle MoE (Mixture-of-Experts) Models

MoE models often have fused 3D expert weights (shape [num_experts, hidden, intermediate]) that must be "unfused" into per-expert nn.Linear layers for quantization.

Check if auto-handled

Transformers >= 5.0 has a linear_loop experts interface that auto-unfuses most MoE models. Test first — it may just work.

Register custom unfusing

If auto-unfusing fails, create a custom module in auto_round/modeling/fused_moe/:

1. Create auto_round/modeling/fused_moe/your_moe.py:

python
"""Unfuse fused MoE weights for YourModel."""

import torch
import torch.nn as nn
from auto_round.modeling.fused_moe.replace_modules import register_replacement


@register_replacement("YourMoELayer")
def replace_your_moe_layer(module, name, model):
    """Replace FusedMoE with per-expert nn.Linear layers."""
    experts = nn.ModuleList()
    for i in range(module.num_experts):
        linear = nn.Linear(module.hidden_size, module.intermediate_size, bias=False)
        linear.weight.data = module.weight[i].clone()
        experts.append(linear)
    return experts

2. Register in BUILTIN_MODULES:

Edit auto_round/modeling/fused_moe/replace_modules.py:

python
BUILTIN_MODULES["your_model_type"] = LazyImport("auto_round.modeling.fused_moe.your_moe")
Existing MoE implementations
Model TypeFilePattern
llama4fused_moe/llama4.pyCustom replacement for no use_experts_implementation
deepseek_v2fused_moe/deepseek_v2.pyq_scale calibration for Gaudi
step3p5fused_moe/step3_5_moe.pySplits fused MoELinear
qwen3_omni_moefused_moe/qwen3_omni.pyThinker + talker MoE

Step 3: Add Custom Forward Pass

Some models have non-standard forward passes that don't get calibrated correctly with the default model.forward(). This is common for multi-component architectures.

Edit _handle_special_model() in auto_round/special_model_handler.py:

python
def _your_model_forward(model, **kwargs):
    """Custom forward that routes through all quantizable components."""
    # Example: route through both encoder and decoder
    encoder_output = model.encoder(**kwargs)
    decoder_output = model.decoder(encoder_output, **kwargs)
    return decoder_output


def _handle_special_model(model):
    ...
    if hasattr(model, "config") and model.config.model_type == "your_model_type":
        from functools import partial

        model.forward = partial(_your_model_forward, model)
    return model
When is this needed?
  • Model has multiple sub-models (thinker/talker, encoder/decoder)
  • Default forward doesn't exercise all quantizable layers
  • Model needs special input preprocessing during calibration
Existing examples
ModelCustom ForwardPurpose
deepseek_vl_v2_deepseek_vl2_forwardRoute through language component
qwen2_5_omni_qwen2_5_omni_forwardRoute through thinker → talker
qwen3_omni_moe_qwen3_omni_moe_forwardHandle MoE routing in omni model
Show full SKILL.md (241 more words)Show less

Step 4: Handle Non-Standard Linear Layers

AutoRound quantizes these layer types by default:

python
# auto_round/utils/common.py
SUPPORTED_LAYER_TYPES = (torch.nn.Linear, transformers.pytorch_utils.Conv1D)
INNER_SUPPORTED_LAYER_TYPES = ("FP8Linear",)  # matched by class name string

If your model uses a custom linear type (e.g., QuantizedLinear, FP8Linear), it won't be quantized unless registered.

Option A: String-based matching

INNER_SUPPORTED_LAYER_TYPES matches by class name string — useful for external classes that can't be imported directly:

python
INNER_SUPPORTED_LAYER_TYPES = ("FP8Linear", "YourCustomLinear")
Option B: Type-based registration

If you can import the class:

python
from your_library import YourLinear

SUPPORTED_LAYER_TYPES = SUPPORTED_LAYER_TYPES + (YourLinear,)

Step 5: Handle Shared Cache Keys

Some models share tensors across blocks during inference (e.g., Gemma3's rotary position embeddings). These must be declared so the calibration cache doesn't duplicate or corrupt them.

Edit SPECIAL_SHARED_CACHE_KEYS in auto_round/special_model_handler.py:

python
SPECIAL_SHARED_CACHE_KEYS["YourModelForCausalLM"] = ("shared_position_embeddings", "shared_rope")

The key is the class name of the model (not model_type).

Step 6: Test

python
def test_your_model_quantization():
    ar = AutoRound(
        "your-org/your-model",
        scheme="W4A16",
        iters=2,
        nsamples=2,
        batch_size=2,
    )
    compressed_model, layer_config = ar.quantize()
    # Verify layers were quantized
    assert len(layer_config) > 0, "No layers were quantized"

    ar.save_quantized(output_dir="./tmp_your_model", format="auto_round")

    # Verify inference works
    from auto_round.utils import model_infer

    output = model_infer(compressed_model, tokenizer, "Hello world")
    assert output is not None

Step 7: Update Documentation

  1. Add model to supported list in README.md
  2. Update README_CN.md with equivalent Chinese content
  3. Add example script if the model has notable differences

Checklist

  • get_block_names() finds all quantizable blocks
  • MoE layers (if any) are unfused correctly
  • calib() runs without shape errors
  • All target layers are quantized (check "Quantized X/Y layers" log)
  • Forward pass exercises all quantizable components
  • Quantized model produces valid outputs
  • Export to target format works
  • README.md + README_CN.md updated

Key Files

FilePurpose
auto_round/special_model_handler.pyBlock handlers, custom forwards, shared cache keys
auto_round/modeling/fused_moe/replace_modules.pyMoE unfusing registry (BUILTIN_MODULES)
auto_round/utils/common.pySUPPORTED_LAYER_TYPES, INNER_SUPPORTED_LAYER_TYPES
auto_round/utils/model.pyget_block_names(), is_mllm_model(), model loading
auto_round/compressors/data_driven.pyNew-architecture quantization loop and block scheduling
auto_round/algorithms/quantization/base.pyQuantizer block execution, sampling, and diffusion output configs
auto_round/calibration/llm.pyLLM calibration data collection and calib() flow
auto_round/autoround.pyAutoRound factory — model type routing logic

© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/adapt-new-llm of intel/auto-round.

Open the folder on GitHubat commit ae21ef9

Compare with similar skills

Adapt New LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Adapt New LLM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Adapt New LLM this skillintel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
CI Fails Buildkiteguqiong96/Lvllm4642 repos~349Automated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    464 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Subwave LLM Bench

    perminder-klair/subwave

    Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

    1.4k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from intel/auto-round

  • Official

    Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

    1.6k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Add Export Format

    intel/auto-round

    Official

    Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

    1.6k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Add Inference Backend

    intel/auto-round

    Official

    Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Official

    Add a new quantization data type to AutoRound (e.g., INT, FP8, MXFP, NVFP, GGUF variants).

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Add Vlm Model

    intel/auto-round

    Official

    Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

    1.6k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Review PR

    intel/auto-round

    Official

    Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files…

    1.6k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about Adapt New LLM

What does Adapt New LLM do?

Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box. Adapt New LLM is an agent skill from intel/auto-round, published by the product's own GitHub organization. Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

When should I use Adapt New LLM?

Adapt New LLM fits situations like: quantization fails for a new model type; block detection doesnt find layers; moE models need unfusing; custom forward passes are needed.

How do I install Adapt New LLM in Claude Code?

Run `npx skills add intel/auto-round --skill adapt-new-llm -a claude-code`. Or copy the skill folder (.claude/skills/adapt-new-llm in intel/auto-round) into .claude/skills/adapt-new-llm in your project. Claude Code loads it when a task matches its description.

How do I install Adapt New LLM in Codex?

Run `npx skills add intel/auto-round --skill adapt-new-llm -a codex`. Or copy the skill folder (.claude/skills/adapt-new-llm in intel/auto-round) into .agents/skills/adapt-new-llm in your project. Codex loads it when a task matches its description.

Can I use Adapt New LLM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/auto-round --skill adapt-new-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adapt-new-llm, .gemini/skills/adapt-new-llm, .github/skills/adapt-new-llm and .opencode/skills/adapt-new-llm in your project.

What does Adapt New LLM need to run?

SKILL.md names no scripts, command-line tools or credentials: Adapt New LLM is instructions for the agent only. Our summary lists: Python 3.

Does Adapt New LLM access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Adapt New LLM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Adapt New LLM use?

Adapt New LLM is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Adapt New LLM use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Adapt New LLM?

Skills that share tags, products or a category with Adapt New LLM: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Adapt New LLM?

intel (a GitHub organization, an official publisher) maintains it in intel/auto-round, which has 1,628 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.

Source: intel/auto-round on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.