SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.
$ npx skills add intel/auto-round --skill adapt-new-llm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install intel/auto-round adapt-new-llm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/adapt-new-llm .claude/skills/adapt-new-llm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "adapt-new-llm" agent skill from https://github.com/intel/auto-round/tree/main/.claude/skills/adapt-new-llm into .claude/skills/adapt-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adapt-new-llm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/intel/auto-round/tree/main/.claude/skills/adapt-new-llmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add intel/auto-round --skill adapt-new-llm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install intel/auto-round adapt-new-llm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/adapt-new-llm .agents/skills/adapt-new-llm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "adapt-new-llm" agent skill from https://github.com/intel/auto-round/tree/main/.claude/skills/adapt-new-llm into .agents/skills/adapt-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adapt-new-llm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/auto-round --skill adapt-new-llm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install intel/auto-round adapt-new-llm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/adapt-new-llm .cursor/skills/adapt-new-llm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "adapt-new-llm" agent skill from https://github.com/intel/auto-round/tree/main/.claude/skills/adapt-new-llm into .cursor/skills/adapt-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adapt-new-llm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/intel/auto-round.git --path .claude/skills/adapt-new-llm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add intel/auto-round --skill adapt-new-llm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install intel/auto-round adapt-new-llm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/adapt-new-llm .gemini/skills/adapt-new-llm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "adapt-new-llm" agent skill from https://github.com/intel/auto-round/tree/main/.claude/skills/adapt-new-llm into .gemini/skills/adapt-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adapt-new-llm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install intel/auto-round adapt-new-llmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add intel/auto-round --skill adapt-new-llm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/adapt-new-llm .github/skills/adapt-new-llm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "adapt-new-llm" agent skill from https://github.com/intel/auto-round/tree/main/.claude/skills/adapt-new-llm into .github/skills/adapt-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adapt-new-llm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add intel/auto-round --skill adapt-new-llm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install intel/auto-round adapt-new-llm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/intel/auto-round.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/adapt-new-llm .opencode/skills/adapt-new-llm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "adapt-new-llm" agent skill from https://github.com/intel/auto-round/tree/main/.claude/skills/adapt-new-llm into .opencode/skills/adapt-new-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "adapt-new-llm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
adapt-new-llmAdapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.
Adapt New LLM is an agent skill from intel/auto-round, published by the product's own GitHub organization. Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box. Use when quantization fails for a new model type, block detection doesn't find layers, MoE models need unfusing, custom forward passes are needed, or non-standard linear layer types need handling.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. The repository describes itself as: A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ae21ef9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Adapt New LLM loads about 2.4k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 627 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from intel/auto-round at commit ae21ef9, republished under its Apache-2.0 licence (© intel). 627 words, ~2,437 tokens.
.claude/skills/adapt-new-llm/SKILL.md (or your agent's skills folder).Most standard Transformers-based LLMs work with AutoRound out-of-the-box. This skill covers what to do when a new model architecture requires code changes. The need for adaptation typically arises from:
nn.Linear or Conv1D)Try quantizing the model first:
from auto_round import AutoRound
ar = AutoRound("your-org/your-model", scheme="W4A16", iters=2, nsamples=2)
ar.quantize_and_save(output_dir="./test_output", format="auto_round")Common failure modes and their fixes:
| Error / Symptom | Root Cause | Fix Section |
|---|---|---|
| "No quantizable layers found" | Block detection failed | Step 1 |
| "Quantized 0/N layers" | Layers not nn.Linear/Conv1D | Step 4 |
| Shape mismatch in MoE layers | Fused expert weights | Step 2 |
| Wrong outputs / calibration diverges | Forward pass not exercised correctly | Step 3 |
| Cache key errors (Gemma3-style) | Shared position embeddings | Step 5 |
AutoRound discovers quantizable blocks via get_block_names() which searches
recursively for nn.ModuleList instances. If your model has a non-standard
layer hierarchy, block detection may fail.
from auto_round.utils import get_block_names
model = ... # loaded model
print(get_block_names(model))to_quant_block_names parameterFor simple cases, override block names without code changes:
ar = AutoRound(
model,
to_quant_block_names="model.decoder.layers", # explicit path
)SPECIAL_MULTIMODAL_BLOCKFor multimodal or multi-component models, add a custom block handler in
auto_round/special_model_handler.py:
def _get_your_model_multimodal_block(model, quant_vision=False):
"""Get block names for YourModel.
YourModel structure:
- encoder.layers: encoder blocks
- decoder.layers: decoder blocks
"""
block_names = []
if quant_vision and hasattr(model, "encoder"):
block_names.append([f"encoder.layers.{i}" for i in range(len(model.encoder.layers))])
block_names.append([f"decoder.layers.{i}" for i in range(len(model.decoder.layers))])
return block_names
# Register: key must match model.config.model_type
SPECIAL_MULTIMODAL_BLOCK["your_model_type"] = _get_your_model_multimodal_blockAlso add to support lists if applicable:
# If text-only calibration works for this multimodal model:
SUPPORT_ONLY_TEXT_MODELS.append("your_model_type")
# If batch_size must be limited:
mllms_with_limited_bs = (..., "your_model_type")MoE models often have fused 3D expert weights (shape
[num_experts, hidden, intermediate]) that must be "unfused" into per-expert
nn.Linear layers for quantization.
Transformers >= 5.0 has a linear_loop experts interface that auto-unfuses
most MoE models. Test first — it may just work.
If auto-unfusing fails, create a custom module in
auto_round/modeling/fused_moe/:
1. Create auto_round/modeling/fused_moe/your_moe.py:
"""Unfuse fused MoE weights for YourModel."""
import torch
import torch.nn as nn
from auto_round.modeling.fused_moe.replace_modules import register_replacement
@register_replacement("YourMoELayer")
def replace_your_moe_layer(module, name, model):
"""Replace FusedMoE with per-expert nn.Linear layers."""
experts = nn.ModuleList()
for i in range(module.num_experts):
linear = nn.Linear(module.hidden_size, module.intermediate_size, bias=False)
linear.weight.data = module.weight[i].clone()
experts.append(linear)
return experts2. Register in BUILTIN_MODULES:
Edit auto_round/modeling/fused_moe/replace_modules.py:
BUILTIN_MODULES["your_model_type"] = LazyImport("auto_round.modeling.fused_moe.your_moe")| Model Type | File | Pattern |
|---|---|---|
llama4 | fused_moe/llama4.py | Custom replacement for no use_experts_implementation |
deepseek_v2 | fused_moe/deepseek_v2.py | q_scale calibration for Gaudi |
step3p5 | fused_moe/step3_5_moe.py | Splits fused MoELinear |
qwen3_omni_moe | fused_moe/qwen3_omni.py | Thinker + talker MoE |
Some models have non-standard forward passes that don't get calibrated correctly
with the default model.forward(). This is common for multi-component
architectures.
Edit _handle_special_model() in auto_round/special_model_handler.py:
def _your_model_forward(model, **kwargs):
"""Custom forward that routes through all quantizable components."""
# Example: route through both encoder and decoder
encoder_output = model.encoder(**kwargs)
decoder_output = model.decoder(encoder_output, **kwargs)
return decoder_output
def _handle_special_model(model):
...
if hasattr(model, "config") and model.config.model_type == "your_model_type":
from functools import partial
model.forward = partial(_your_model_forward, model)
return model| Model | Custom Forward | Purpose |
|---|---|---|
deepseek_vl_v2 | _deepseek_vl2_forward | Route through language component |
qwen2_5_omni | _qwen2_5_omni_forward | Route through thinker → talker |
qwen3_omni_moe | _qwen3_omni_moe_forward | Handle MoE routing in omni model |
AutoRound quantizes these layer types by default:
# auto_round/utils/common.py
SUPPORTED_LAYER_TYPES = (torch.nn.Linear, transformers.pytorch_utils.Conv1D)
INNER_SUPPORTED_LAYER_TYPES = ("FP8Linear",) # matched by class name stringIf your model uses a custom linear type (e.g., QuantizedLinear, FP8Linear),
it won't be quantized unless registered.
INNER_SUPPORTED_LAYER_TYPES matches by class name string — useful for
external classes that can't be imported directly:
INNER_SUPPORTED_LAYER_TYPES = ("FP8Linear", "YourCustomLinear")If you can import the class:
from your_library import YourLinear
SUPPORTED_LAYER_TYPES = SUPPORTED_LAYER_TYPES + (YourLinear,)Some models share tensors across blocks during inference (e.g., Gemma3's rotary position embeddings). These must be declared so the calibration cache doesn't duplicate or corrupt them.
Edit SPECIAL_SHARED_CACHE_KEYS in auto_round/special_model_handler.py:
SPECIAL_SHARED_CACHE_KEYS["YourModelForCausalLM"] = ("shared_position_embeddings", "shared_rope")The key is the class name of the model (not model_type).
def test_your_model_quantization():
ar = AutoRound(
"your-org/your-model",
scheme="W4A16",
iters=2,
nsamples=2,
batch_size=2,
)
compressed_model, layer_config = ar.quantize()
# Verify layers were quantized
assert len(layer_config) > 0, "No layers were quantized"
ar.save_quantized(output_dir="./tmp_your_model", format="auto_round")
# Verify inference works
from auto_round.utils import model_infer
output = model_infer(compressed_model, tokenizer, "Hello world")
assert output is not NoneREADME.mdREADME_CN.md with equivalent Chinese contentget_block_names() finds all quantizable blockscalib() runs without shape errors| File | Purpose |
|---|---|
auto_round/special_model_handler.py | Block handlers, custom forwards, shared cache keys |
auto_round/modeling/fused_moe/replace_modules.py | MoE unfusing registry (BUILTIN_MODULES) |
auto_round/utils/common.py | SUPPORTED_LAYER_TYPES, INNER_SUPPORTED_LAYER_TYPES |
auto_round/utils/model.py | get_block_names(), is_mllm_model(), model loading |
auto_round/compressors/data_driven.py | New-architecture quantization loop and block scheduling |
auto_round/algorithms/quantization/base.py | Quantizer block execution, sampling, and diffusion output configs |
auto_round/calibration/llm.py | LLM calibration data collection and calib() flow |
auto_round/autoround.py | AutoRound factory — model type routing logic |
© intel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/adapt-new-llm of intel/auto-round.
Open the folder on GitHubat commit ae21ef9
Adapt New LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Adapt New LLM this skillintel/auto-round | 1.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| CI Fails Buildkiteguqiong96/Lvllm | 464 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
guqiong96/Lvllm
Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.
perminder-klair/subwave
Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…
intel/auto-round
Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).
intel/auto-round
Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).
intel/auto-round
Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).
intel/auto-round
Add a new quantization data type to AutoRound (e.g., INT, FP8, MXFP, NVFP, GGUF variants).
intel/auto-round
Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.
intel/auto-round
Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files…
Categories
Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box. Adapt New LLM is an agent skill from intel/auto-round, published by the product's own GitHub organization. Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.
Adapt New LLM fits situations like: quantization fails for a new model type; block detection doesnt find layers; moE models need unfusing; custom forward passes are needed.
Run `npx skills add intel/auto-round --skill adapt-new-llm -a claude-code`. Or copy the skill folder (.claude/skills/adapt-new-llm in intel/auto-round) into .claude/skills/adapt-new-llm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add intel/auto-round --skill adapt-new-llm -a codex`. Or copy the skill folder (.claude/skills/adapt-new-llm in intel/auto-round) into .agents/skills/adapt-new-llm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intel/auto-round --skill adapt-new-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adapt-new-llm, .gemini/skills/adapt-new-llm, .github/skills/adapt-new-llm and .opencode/skills/adapt-new-llm in your project.
SKILL.md names no scripts, command-line tools or credentials: Adapt New LLM is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Adapt New LLM is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Adapt New LLM: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
intel (a GitHub organization, an official publisher) maintains it in intel/auto-round, which has 1,628 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.
Source: intel/auto-round on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.