Agent skill

Quark Torch Quant Plan

by amd in amd/Quark

Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent.

MITAuto-check passedAI & LLM Engineering

Install Quark Torch Quant Plan

skills CLI
$ npx skills add amd/Quark --skill quark-torch-quant-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-quant-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan .claude/skills/quark-torch-quant-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-quant-plan
GitHub stars
181
Token cost
~2.3k tokens
SKILL.md length
905 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent.

  • Works in 5 steps: Check prerequisites: Is… → Gather intent: What does the user care… → Present the decision table: Show… → …
  • The user needs quantization scheme recommendations
  • SKILL.md covers Purpose, Inputs, Outputs: quant_plan.json and Available Quantization Schemes…, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quark Torch Quant Plan is an agent skill from amd/Quark. Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent. Use when the user needs quantization scheme recommendations, exclusion lists, algorithm selection, KV cache decisions, per-layer overrides, or a draft quantplan. Trigger for "quantize with FP8", "what scheme should I use", "plan PTQ", "INT4 quantization", "choose quantization config", "quantization plan", or when the user has a model analysis and needs to decide how to quantize it.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. The licence is MIT.

When your agent uses it

  • The user needs quantization scheme recommendations
  • Exclusion lists
  • Algorithm selection
  • KV cache decisions

Example prompts

  • “quantize with FP8”
  • “what scheme should I use”
  • “plan PTQ”
  • “/quark-torch-quant-plan”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Check prerequisites: Is model_analysis.json available? If not, route to quark-torch-model-intake first.
  2. Gather intent: What does the user care about most — accuracy, size, speed? What hardware will run inference?
  3. Present the decision table: Show defaults, explain the tradeoffs, and let the user adjust.
  4. Confirm: Always required. Show the final plan summary before writing it.
  5. Emit: Write quant_plan.json.

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch Quant Plan loads about 2.3k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 905 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 905 words, ~2,284 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-quant-plan/SKILL.md (or your agent's skills folder).
name
quark-torch-quant-plan
description
Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent. Use when the user needs quantization scheme recommendations, exclusion lists, algorithm selection, KV cache decisions, per-layer overrides, or a draft quant_plan. Trigger for "quantize with FP8", "what scheme should I use", "plan PTQ", "INT4 quantization", "choose quantization config", "quantization plan", or when the user has a model analysis and needs to decide how to quantize it.
layer
l1-atomic
primary_artifact
quant_plan.json
source_knowledge
quark/torch/quantization/config/template.py, examples/torch/language_modeling/llm_ptq/README.md, examples/torch/language_modeling/llm_ptq/quantize_quark.py

quark-torch-quant-plan

Purpose

Convert a model analysis plus the user's intent into a confirmed quant_plan.json. This skill makes the quantization decisions — which scheme, which algorithm, what to exclude — without generating scripts or executing PTQ. The plan is the contract between the user's intent and the execution step.

Inputs

  • model_analysis.json from quark-torch-model-intake
  • env_context.json for accelerator-aware scheme recommendations
  • User preferences (target precision, accuracy goal)

Outputs: quant_plan.json

Records the chosen scheme, algorithm, layer overrides, calibration settings, and evaluation intent.

Schema: quant_plan.schema.json

json
{
  "model": {
    "model_type": "qwen3",
    "analysis_ref": "./model_analysis.json"
  },
  "global_scheme": "fp8",
  "kv_cache_scheme": "fp8",
  "exclude_layers": ["lm_head"],
  "layer_quant_config": {},
  "algorithm": null,
  "calibration": {
    "dataset": "pileval",
    "num_calib_data": 128,
    "seq_len": 512
  },
  "evaluation_intent": "smoke",
  "requires_confirmation": false
}

Available Quantization Schemes (21 total)

Weight-Only INT4 (best for deployment size reduction)
SchemeDescriptionUse Case
int4_wo_32INT4, group size 32Highest accuracy among INT4
int4_wo_64INT4, group size 64Good balance
int4_wo_128INT4, group size 128Smaller overhead
int4_wo_per_channelINT4, per-channelLeast overhead
uint4_wo_32/64/128/per_channelUnsigned INT4 variantsGGUF export compatibility
Weight+Activation INT8
SchemeDescriptionUse Case
int8INT8 per-tensor for both W and ACPU deployment, good accuracy
FP8 (best accuracy-size tradeoff for GPU inference)
SchemeDescriptionUse Case
fp8FP8 E4M3 per-tensorStandard GPU quantization
ptpc_fp8Per-Token-Per-Channel FP8Higher accuracy, dynamic activation quantization
OCP Microscaling Formats
SchemeDescriptionUse Case
mxfp4OCP MXFP4Aggressive compression
mxfp6_e3m2OCP MXFP6 (E3M2)Better range
mxfp6_e2m3OCP MXFP6 (E2M3)Better precision
mxfp4_mxfp6_e2m3MXFP4 weights + MXFP6 activationsMixed precision
mxfp4_fp8MXFP4 weights + FP8 activationsMixed precision
AMD-Specific
SchemeDescriptionUse Case
amdfp4amdfp4, group size 16AMD MI300X optimized
amdfp4_g32amdfp4, group size 32AMD MI300X, less overhead
Other
SchemeDescriptionUse Case
nvfp4NVFP4: FP4 group_size=16 with FP8 E4M3 scaleNVIDIA Blackwell/Hopper
mx6MX6 formatExperimental
bfp16Block Floating Point 16-bitExperimental
int4_wa_64INT4 weights + activations, group 64Research

Available Algorithms (7 primary)

AlgorithmCompatible SchemesDescription
awqINT4/UINT4 weight-onlyActivation-aware weight quantization — finds optimal per-channel scaling
gptqINT4/UINT4 weight-onlySecond-order weight optimization — often better than AWQ for small models
smoothquantINT8, FP8Migrates quantization difficulty from activations to weights
autosmoothquantINT8, FP8Automatic SmoothQuant with optimal alpha search
rotationVariousRotation-based optimization to equalize weight distribution
gptaqINT4/UINT4GPTAQ variant combining GPTQ with activation quantization
qronosVariousCustom algorithm for time-series-aware quantization

Algorithms can be combined: --quant_algo awq,smoothquant

KV Cache Quantization

  • Only fp8 is supported for KV cache (--kv_cache_dtype fp8)
  • Adds --min_kv_scale option (default 0.0) to prevent extreme scale values
  • --kv_cache_post_rope quantizes KV cache after RoPE (inside cache) instead of at k_proj/v_proj outputs — can improve accuracy for some models

Decision Guide

Help the user choose based on their priorities:

"I want the best accuracy" → fp8 or ptpc_fp8, optionally with smoothquant "I want the smallest model" → int4_wo_32 with awq or gptq "I need CPU deployment" → int8 (the only scheme that works well on CPU) "I need GGUF format" → uint4_wo_32 with awq, export as GGUF "I'm on AMD MI300X" → amdfp4 for best hardware utilization "I'm on NVIDIA H100/Blackwell" → fp8 or nvfp4 "I want to experiment" → mxfp4 for aggressive compression research

Decision Table (MUST show to user)

ALWAYS present this table to the user and WAIT for confirmation before finalizing. Do not skip this step.

Fill in the "Value" column based on the user's request and model analysis, then show:

DecisionValueReason
global_scheme(fill)(why this scheme)
kv_cache_scheme(fill: fp8 or null)(explain)
exclude_layers["lm_head"]Standard — lm_head stays full precision
layer_quant_config(fill: dict of pattern -> scheme, or {} if none)(explain which patterns and why)
algorithm(fill: algorithm or null)(explain)
calibration_datasetpilevalFast default
num_calib_data128Standard default
seq_len512Standard default
evaluation_intentsmokeQuick PPL check post-quantization

After showing the table, ask: "Confirm this plan? Any changes?"

Do NOT proceed until the user confirms.

Show full SKILL.md (326 more words)Show less

Layer-Specific Overrides via layer_quant_config

The layer_quant_config plan field is a dict of pattern -> scheme pairs. It is the single mechanism for any "quantize layer/module X with scheme Y" intent — including attention modules, MoE experts, lm_head, etc. Each entry emits one --layer_quant_scheme PATTERN SCHEME CLI argument.

json
"layer_quant_config": {
  "*self_attn*": "fp8",
  "lm_head": "int8",
  "*experts*": "fp8"
}

translates to:

bash
--quant_scheme <global_scheme> \
--layer_quant_scheme '*self_attn*' fp8 \
--layer_quant_scheme lm_head int8 \
--layer_quant_scheme '*experts*' fp8

When a user asks for attention-module quantization (e.g. "self_attn in fp8"), populate this field with the appropriate pattern (commonly *self_attn* for LLaMA-style models; adjust for models whose attention submodule has a different name). Do NOT introduce a dedicated attention field — keep all per-pattern overrides in layer_quant_config.

Rules

  • Keep scope to plan creation only. Do not generate scripts, do not run quantization, do not export. Those are separate skills.
  • Require a model analysis first. Without knowing the model architecture and layer count, you cannot make informed scheme recommendations. If model_analysis.json is missing, route back to quark-torch-model-intake.
  • Always present the decision table before finalizing. The user should explicitly confirm the scheme, algorithm, and exclusions.
  • If a risky scheme is chosen (e.g., mxfp4 on a model where accuracy loss may be significant), keep the user's choice but record the risk in the plan.

Interaction Flow

  1. Check prerequisites: Is model_analysis.json available? If not, route to quark-torch-model-intake first.
  2. Gather intent: What does the user care about most — accuracy, size, speed? What hardware will run inference?
  3. Present the decision table: Show defaults, explain the tradeoffs, and let the user adjust.
  4. Confirm: Always required. Show the final plan summary before writing it.
  5. Emit: Write quant_plan.json.

Recovery

  • If the model analysis is incomplete, produce a draft plan with requires_confirmation: true and note what facts are missing.
  • If the user picks an unusual combination (e.g., awq with fp8 — AWQ is designed for INT4), explain why it might not work well and suggest alternatives, but respect the user's choice if they insist.
  • If calibration dataset preferences are unclear, default to pileval with 128 samples — it is the fastest option and works for most models.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch Quant Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch Quant Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch Quant Plan this skillamd/Quark181—~2.3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
CI Fails Buildkiteguqiong96/Lvllm4642 repos~349Automated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    464 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Subwave LLM Bench

    perminder-klair/subwave

    Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

    1.4k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 10 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 10 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 10 days ago
    Auto-check passed

Questions about Quark Torch Quant Plan

What does Quark Torch Quant Plan do?

Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent. Quark Torch Quant Plan is an agent skill from amd/Quark. Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent.

When should I use Quark Torch Quant Plan?

Quark Torch Quant Plan fits situations like: the user needs quantization scheme recommendations; exclusion lists; algorithm selection; KV cache decisions.

How do I install Quark Torch Quant Plan in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-quant-plan -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan in amd/Quark) into .claude/skills/quark-torch-quant-plan in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch Quant Plan in Codex?

Run `npx skills add amd/Quark --skill quark-torch-quant-plan -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan in amd/Quark) into .agents/skills/quark-torch-quant-plan in your project. Codex loads it when a task matches its description.

Can I use Quark Torch Quant Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-quant-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-quant-plan, .gemini/skills/quark-torch-quant-plan, .github/skills/quark-torch-quant-plan and .opencode/skills/quark-torch-quant-plan in your project.

What does Quark Torch Quant Plan need to run?

SKILL.md names no scripts, command-line tools or credentials: Quark Torch Quant Plan is instructions for the agent only.

Does Quark Torch Quant Plan access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch Quant Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch Quant Plan use?

Quark Torch Quant Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch Quant Plan use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Torch Quant Plan?

Skills that share tags, products or a category with Quark Torch Quant Plan: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch Quant Plan?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.