Agent skill

Quark Onnx Quant Plan

by amd in amd/Quark

Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.

MITAuto-check passedAI & LLM Engineering

Install Quark Onnx Quant Plan

skills CLI
$ npx skills add amd/Quark --skill quark-onnx-quant-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-onnx-quant-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan .claude/skills/quark-onnx-quant-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-onnx-quant-plan
GitHub stars
181
Token cost
~4.8k tokens
SKILL.md length
1,721 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.

  • Works in 9 steps: Check prerequisites: Is… → Confirm deployment target: From… → Narrow presets: Start from… → …
  • The user needs preset selection (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX / MXFP …)
  • SKILL.md covers Purpose, Inputs, Outputs: quant_plan.json and Available Presets (from…, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quark Onnx Quant Plan is an agent skill from amd/Quark. Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. Use when the user needs preset selection (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX / MXFP …), calibration method choice (MinMax / Entropy / Percentile / Distribution / NonOverflow / MinMSE / LayerWisePercentile), algorithm selection (CLE / AdaRound / AdaQuant / BiasCorrection / AutoMixprecision), deployment-target gating (CPU / CUDA / ROCm / AMD NPU CNN / AMD NPU Transformer), op-type include/exclude lists, weights-only INT4…

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Performance reviews, LLM inference and serving and Deployment. It works with ONNX and CUDA. The licence is MIT.

When your agent uses it

  • The user needs preset selection (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX / MXFP …)
  • Calibration method choice (MinMax / Entropy / Percentile / Distribution / NonOverflow / MinMSE / LayerWisePercentile)
  • Algorithm selection (CLE / AdaRound / AdaQuant / BiasCorrection / AutoMixprecision)
  • Deployment-target gating (CPU / CUDA / ROCm / AMD NPU CNN / AMD NPU Transformer)

Example prompts

  • “what preset should I use for my ONNX model”
  • “choose XINT8 vs A8W8”
  • “plan ONNX PTQ”
  • “/quark-onnx-quant-plan”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Check prerequisites: Is model_analysis.json available? If not, route to
  2. Confirm deployment target: From session_context.json.constraints.deployment_target or
  3. Narrow presets: Start from model_analysis.json.quantization_targets.onnx_specific.preset_candidates,
  4. Pick calibration method: Default per target (MinMax for general; PowerOfTwo_MinMSE for
  5. Pick algorithms: Default per architecture and preset (CLE for INT8 CNN;
  6. Set extra options: From the table of common knobs, only set what's needed; document each.
  7. Present the decision table: Show defaults, explain the tradeoffs, let the user adjust.
  8. Confirm: Always required. Show the final plan summary before writing.
  9. Emit: Write quant_plan.json. Surface any new constraints back to quark-onnx-router so

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Onnx Quant Plan loads about 4.8k tokens when it runs. Until then it costs about 235 tokens; SKILL.md has 1,721 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~235
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,721 words, ~4,808 tokens.

Download SKILL.mdSave it as .claude/skills/quark-onnx-quant-plan/SKILL.md (or your agent's skills folder).
name
quark-onnx-quant-plan
description
Build a Quark ONNX PTQ quantization plan from `model_analysis.json` and user intent. Use when the user needs preset selection (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX* / MXFP* …), calibration method choice (MinMax / Entropy / Percentile / Distribution / NonOverflow / MinMSE / LayerWisePercentile), algorithm selection (CLE / AdaRound / AdaQuant / BiasCorrection / AutoMixprecision), deployment-target gating (CPU / CUDA / ROCm / AMD NPU CNN / AMD NPU Transformer), op-type include/exclude lists, weights-only INT4 (MatMulNBits) decisions, dynamic vs static quantization, or a draft `quant_plan.json`. Trigger for "what preset should I use for my ONNX model", "choose XINT8 vs A8W8", "plan ONNX PTQ", "INT4 "BFP16 / MXFP4 for my model", "calibration method for Ryzen AI", "should I use AdaRound or AdaQuant", "SmoothQuant alpha", or when the user has a model analysis and needs to decide how to quantize an ONNX model.
layer
l1-atomic
primary_artifact
quant_plan.json
source_knowledge
quark/onnx/quantization/config/custom_config.py, quark/onnx/quantization/config/algorithm.py, quark/onnx/quantization/config/config.py…

quark-onnx-quant-plan

Purpose

Convert an ONNX model_analysis.json plus the user's intent into a confirmed quant_plan.json. This skill makes the quantization decisions for the ONNX-to-ONNX flow — which preset, which calibration method, which algorithm, which op types to include/exclude, whether to enable an NPU target, whether to use external-data — without generating scripts or executing the quantization. The plan is the contract between the user's intent and the execution step.

Inputs

  • model_analysis.json from quark-onnx-model-intake (architecture, op histogram, opsets, external-data state, preset_candidates, risks)
  • env_context.json for accelerator-aware preset gating (CUDA major / ROCm major / NPU presence)
  • onnx_install_result.json (optional) — gates BFP/MX/Extended presets (require the custom-ops library to be compiled for the target EP)
  • User preferences: deployment target, accuracy goal, model-size goal, calibration data availability

Outputs: quant_plan.json

Records the chosen preset (or custom config), calibration method, algorithm list, layer/op overrides, NPU flag, external-data setting, and evaluation intent. Shares the schema with the Torch plan; ONNX-specific fields live under onnx_specific.

Schema: quant_plan.schema.json

json
{
  "model": {
    "model_type": "onnx",
    "architecture_guess": "cnn",
    "analysis_ref": "./model_analysis.json"
  },
  "backend": "onnx",
  "deployment_target": "npu_cnn",
  "preset": "XINT8",
  "calibration": {
    "method": "PowerOfTwo_MinMSE",
    "data_size": 200,
    "batch_size": 1,
    "use_external_data_format": false,
    "optimize_mem": false,
    "worker_num": 1
  },
  "algorithms": ["CLE"],
  "onnx_specific": {
    "enable_npu_cnn": true,
    "enable_npu_transformer": false,
    "include_cle": true,
    "include_fast_ft": false,
    "op_types_to_quantize": null,
    "nodes_to_quantize": null,
    "nodes_to_exclude": null,
    "subgraphs_to_exclude": [],
    "extra_options": {
      "OpTypesToExcludeOutputQuantization": []
    },
    "use_external_data_format": false,
    "execution_providers": ["CPUExecutionProvider"]
  },
  "evaluation_intent": "smoke",
  "requires_confirmation": false
}

Available Presets (from DefaultConfigMapping)

Quark ONNX ships 50+ named presets in quark/onnx/quantization/config/custom_config.py. Pick the smallest viable set for the user's architecture + deployment target; never list them all.

AMD NPU CNN — Ryzen AI / VAI (enable_npu_cnn=True, PoF2 scales, NHWC)
PresetDescriptionPicks
XINT8INT8 input + INT8 weight, optimized for NPUDefault for any CNN targeting NPU
XINT8_ADAROUNDXINT8 + AdaRound fast-finetuneWhen base XINT8 loses accuracy
XINT8_ADAQUANTXINT8 + AdaQuant fast-finetuneWhen AdaRound is not enough
VINT8INT8 optimized for VAIMLVAIML deployment
AMD NPU Transformer (enable_npu_transformer=True)
PresetDescriptionPicks
INT16_TRANSFORMER_DEFAULTINT16 activations + INT8 weights, fastOutlier-heavy activations
INT16_TRANSFORMER_ACCURATEINT16 + accuracy algorithmsLargest accuracy headroom
General CPU / CUDA / ROCm INT8 (deployment-agnostic)
PresetDescriptionPicks
A8W8INT8 sym activations + INT8 sym weightsStandard CPU/GPU INT8
A8W8_ADAROUND / A8W8_ADAQUANT+ fast-finetuneAccuracy-tight A8W8
A16W8INT16 sym activations + INT8 sym weightsWide-activation needs
A16W8_ADAROUND / A16W8_ADAQUANT+ fast-finetuneAccuracy-tight A16W8
S8S8_AAWS / U8S8_AAWS / U8U8_AAWA / S16S8_ASWS / U16S8_AAWS (+ ADAROUND / ADAQUANT variants)Various sym/asym INT8/INT16 combosFine-tune sym/asym choice per ORT target
INT8_CNN_DEFAULT / INT8_CNN_ACCURATE / INT16_CNN_DEFAULT / INT16_CNN_ACCURATECNN-tuned INT8/INT16CPU/GPU CNN deployment
Float Fallbacks
PresetDescriptionPicks
FP16 / FP16_ADAQUANTFP16 W+AAccuracy-first when INT is too lossy
BF16 / BF16_ADAQUANTBFloat16 W+ASame, with BF16 range
Block Formats (require Quark custom-ops library compiled for the target EP)
PresetDescriptionPicks
BFP16 / BFP16_ADAQUANTBlock Floating Point 16-bitAMD accelerator deployments
MX4 / MX6 / MX9 (+ ADAQUANT)Micro-Exponents BFP variantsBit-budget exploration
MXFP4E2M1 / MXFP6E2M3 / MXFP6E3M2 / MXFP8E4M3 / MXFP8E5M2 / MXINT8 (+ ADAQUANT)OCP MX formatsModern AMD/NVIDIA accelerators
Mixed-Precision
PresetDescriptionPicks
BF16_BFP16 / BF16_MIXED_BFP16 / BF16_MIXED_BFP16_ADAQUANTBF16 + BFP16 hybridHigh-accuracy + AMD HW
BF16_MXINT8 / BF16_MIXED_MXINT8 / BF16_MIXED_MXINT8_ADAQUANTBF16 + MXInt8 hybridSame, OCP MX flavor
MX9_INT8MX9 + INT8 hybridBit-budget exploration
S16S16_MIXED_S8S8INT16 + INT8 mixedOutlier-aware INT mix

Available Calibration Methods

From quark/onnx/calibration/methods.py (+ ORT built-ins):

MethodWhen to pick
MinMaxDefault for most CNN / weights-only; cheap, deterministic
EntropyKL-divergence-based; helps when activations have long tails
PercentileClip outliers at a chosen percentile (default 99.999)
DistributionDistribution-matching; useful for FP8 p3/same
LayerWisePercentileAuto-picks per-tensor optimal percentile (MAE/MSE) — AMD-specific
PowerOfTwo_NonOverflow (a.k.a. NonOverflow)Required for AMD NPU XINT8 — picks the smallest PoF2 scale that doesn't overflow
PowerOfTwo_MinMSE (a.k.a. MinMSE)Required for AMD NPU XINT8 — picks the PoF2 scale that minimizes MSE; usually better than NonOverflow
Int16Method.MinMaxFor INT16 configs

NPU CNN / XINT8 targets must use a PowerOfTwo* method. Non-PoF2 scales are rejected at NPU runtime — flag this as a hard constraint in the plan.

Available Algorithms

From quark/onnx/quantization/config/algorithm.py and examples/onnx/accuracy_improvement/:

AlgorithmKindCompatible presetsDescription
CLE (Cross-Layer Equalization)PreINT8 CNN configsFolds BN, equalizes per-channel scales across consecutive Conv/Linear layers (Nagel et al., 2019)
BiasCorrectionPostINT8 CNNPost-hoc bias adjustment (Nagel et al., 2019)
AdaRoundPost (fast-finetune)XINT8 / A8W8 / A16W8 / block formatsAdaptive rounding optimization; needs cal data + LR + iterations; GPU-accelerated
AdaQuantPost (fast-finetune)Same as AdaRoundLayer-wise calibration tuning; usually after AdaRound is not enough
AutoMixprecisionPostBlock formats / mixed-precisionAuto-selects sensitivity-based per-layer dtype; can do dual BFP16+MX hybrid

Algorithms compose: e.g. CLE (pre) + AdaRound (post) is a common XINT8 recipe. Combinations beyond two algorithms are usually a red flag — flag them in risks.

Decision Guide

Help the user choose based on their priorities and the architecture from model_analysis.json.model.onnx_specific.architecture_guess:

User intentArchitectureRecommended starting plan
"Best accuracy, AMD GPU"anyBF16 or BF16_MIXED_BFP16
"INT8 CNN, CPU/GPU deployment"cnnINT8_CNN_DEFAULT or A8W8; add CLE if accuracy drops
"INT8 CNN → Ryzen AI NPU"cnnXINT8 + CLE, calibration = PowerOfTwo_MinMSE; usually NHWC pre-conversion via quark.onnx.tools.convert_nchw_to_nhwc
"Ryzen AI NPU, accuracy-tight CNN"cnnXINT8_ADAROUND (then XINT8_ADAQUANT if still short)
"Block format MXFP4 / BFP16 experimentation"anyBFP16 or MXFP4E2M1; require quark.onnx.operators.custom_ops to be compiled
"Hybrid mixed-precision for best size/accuracy"anyBF16_MIXED_BFP16 or S16S16_MIXED_S8S8 + AutoMixprecision

Deployment-Target Gating (HARD constraints)

TargetPreset must satisfyCalibration must beNotes
npu_cnn (Ryzen AI CNN)enable_npu_cnn=True, PoF2 symmetric INT8 per-tensor, NCHW→NHWC donePowerOfTwo_MinMSE or PowerOfTwo_NonOverflowReject A8W8 / BFP16 / MX* / FP16 if user requests npu_cnn
npu_transformer (Ryzen AI Transformer)enable_npu_transformer=True, INT8/INT16 QDQ on MatMul/GemmMinMax / Percentile typicallyReject CNN-only presets
cudaonnxruntime-gpu present, CUDAExecutionProvider availableAnyBlock-format presets require custom-ops
rocmonnxruntime_rocm (ROCm 6.x) or CPU ORT on ROCm 7.x (tools/ci/install_onnxruntime.sh)AnyCustom-ops library must compile for ROCm
cpuAnyAnyBlock-format presets work via CPU custom-ops; expect speed cost

If the deployment target conflicts with a requested preset, the plan must either (a) downgrade to a viable preset and explain, or (b) leave it unset with a high-severity risk in quant_plan.json.

Op-Type / Node Include/Exclude Levers

Three commonly-used knobs the plan should expose:

  • op_types_to_quantize — restrict QDQ insertion to a subset, e.g. ["Conv"] for CNN-only quantization.
  • nodes_to_quantize / nodes_to_exclude — surgical per-node control by graph node name. Node names change after pre-processing (NCHW→NHWC, BN folding, etc.), so resolve them after any pre-processing pass.
  • extra_options["OpTypesToExcludeOutputQuantization"] — keep certain op outputs in float while still quantizing their inputs/weights.

Extra Options Commonly Set in quant_plan.onnx_specific.extra_options

From real examples in examples/onnx/:

OptionTypical valueSource example
SimplifyModelTrue / Falsetoggle OnnxSlim pre-pass
QuantizeFP16TrueFP16-input models
OpTypesToExcludeOutputQuantization["Add", "Mul"] etc.keep selected op outputs in float
FastFinetune{"DataSize": 200, "BatchSize": 2, "NumIterations": 1000, "LearningRate": …, "OptimAlgorithm": "adaround"/"adaquant", "OptimDevice": "cuda:0"/"cpu", "InferDevice": "cuda:0"/"cpu", "EarlyStop": True}AdaRound / AdaQuant tutorials and Auto-Search tutorials
Show full SKILL.md (740 more words)Show less

Decision Table (MUST show to user)

ALWAYS present this table and WAIT for confirmation before finalizing. Fill the "Value" column from the user's request, the model analysis, and the deployment-target gates above.

DecisionValueReason
backendonnxFixed for this skill
deployment_target(fill: cpu / cuda / rocm / npu_cnn / npu_transformer)(from env or user)
preset(fill: name from DefaultConfigMapping or "custom")(why)
calibration.method(fill: MinMax / Entropy / Percentile / Distribution / LayerWisePercentile / PowerOfTwo_MinMSE / PowerOfTwo_NonOverflow)(why)
calibration.data_size200 (default)(why)
calibration.batch_size1–4(why)
algorithms(fill: list e.g. ["CLE"], ["AdaRound"], ["AdaQuant"], ["BiasCorrection"])(why)
onnx_specific.enable_npu_cnn(fill: bool)Hard-tied to deployment_target == "npu_cnn"
onnx_specific.enable_npu_transformer(fill: bool)Hard-tied to deployment_target == "npu_transformer"
onnx_specific.include_cle(fill: bool)CNN INT8 default true
onnx_specific.op_types_to_quantizenull or ["MatMul"] etc.(why)
onnx_specific.use_external_data_formattrue if model > 2 GB (from model_analysis.json)Required for large models
onnx_specific.extra_options(fill: dict)(why — list each key)
evaluation_intentsmoke (default), mlperf, mAP(why)

After showing the table, ask: "Confirm this plan? Any changes?"

Do NOT proceed until the user confirms.

Per-Layer Overrides

For fine-grained control, individual layers can override the global config via QLayerConfig (see examples/onnx/yolo_quantization/quantize_yolo.py):

python
from quark.onnx import QConfig, QLayerConfig, XInt8Spec, CLEConfig

config = QConfig(
    global_config=QLayerConfig(activation=XInt8Spec(), weight=XInt8Spec()),
    algo_config=[CLEConfig()],
    EnableNPUCnn=True,
    exclude=[
        # YOLOX-style: keep a specific subgraph in float
        (["/_head/_modules_list.14/Transpose"], ["/_head/_modules_list.14/Concat_9"]),
    ],
)

Record any per-layer overrides under onnx_specific.subgraphs_to_exclude (list of (start_nodes, end_nodes) tuples) or onnx_specific.nodes_to_exclude (flat name list).

Rules

  • Keep scope to plan creation only. Do not generate scripts, do not run quantization, do not export. Those are separate skills.
  • Require a model analysis first. Without architecture, op histogram, opsets, and the already-QDQ flags, you cannot make informed preset / op-type recommendations. If model_analysis.json is missing, route back to quark-onnx-model-intake.
  • Honor hard deployment-target constraints. NPU CNN ⇒ XINT8 family + PowerOfTwo* calibration + NHWC. NPU Transformer ⇒ INT8_TRANSFORMER_* / INT16_TRANSFORMER_*. Reject conflicting preset requests with a clear explanation.
  • Custom-op preconditions for block formats. Any preset in {BFP16*, MX*, MXFP*, MXINT8*, BF16_*BFP*, BF16_*MXINT8*} requires the Quark custom-ops library to have compiled. If onnx_install_result.json shows the compile failed, suppress these presets from recommendations and emit a risk pointing back to quark-onnx-install.
  • External-data flag tracks model size. If model_analysis.json.model.estimated_size_gb > 2, set onnx_specific.use_external_data_format = true automatically. Otherwise default false.
  • Always present the decision table before finalizing. The user must explicitly confirm.
  • Record risks. If the user picks a risky combination (e.g. XINT8 without PoF2 calibration, MXFP4 on a model that hasn't been validated with AutoMixprecision), keep the user's choice but record the risk in the plan.
  • Do not duplicate fields between model_analysis.json and quant_plan.json. The plan references the analysis via analysis_ref.

Interaction Flow

  1. Check prerequisites: Is model_analysis.json available? If not, route to quark-onnx-model-intake first. Is onnx_install_result.json available? If a block-format preset is on the table, require it.
  2. Confirm deployment target: From session_context.json.constraints.deployment_target or ask. Apply the hard gates from the table above.
  3. Narrow presets: Start from model_analysis.json.quantization_targets.onnx_specific.preset_candidates, filter by deployment target and custom-op availability, present the top 1–3 with rationale.
  4. Pick calibration method: Default per target (MinMax for general; PowerOfTwo_MinMSE for NPU CNN; Percentile for outlier-heavy transformers; consider LayerWisePercentile when accuracy is critical). Decide data_size and batch_size.
  5. Pick algorithms: Default per architecture and preset (CLE for INT8 CNN; AdaRound→AdaQuant for accuracy-tight CNN; BiasCorrection as a cheap post-hoc fixup). At most two algorithms unless justified.
  6. Set extra options: From the table of common knobs, only set what's needed; document each.
  7. Present the decision table: Show defaults, explain the tradeoffs, let the user adjust.
  8. Confirm: Always required. Show the final plan summary before writing.
  9. Emit: Write quant_plan.json. Surface any new constraints back to quark-onnx-router so they land in session_context.json.

Recovery

  • If model_analysis.json shows has_qdq_already=true or has_quark_custom_ops=true — do not produce a plan. Route back to quark-onnx-model-intake with a recommendation to remove existing QDQ first (quark.onnx.tools.remove_qdq).
  • If the user picks a preset that conflicts with their deployment target — explain why it won't work (e.g. "A8W8 uses non-PoF2 scales which the AMD NPU CNN runtime rejects"), suggest the closest viable alternative (e.g. XINT8), and respect the user's choice if they insist — recording the risk.
  • If model_analysis.json shows architecture unknown — produce a draft plan with requires_confirmation: true, list the op histogram, and ask the user which of CNN / transformer / hybrid path to follow.
  • If calibration data is unavailable — explain that PTQ requires representative inputs for accurate scale estimation, and ask the user to supply a small calibration set (≥ a few dozen samples) before the plan can be finalized.
  • If a block-format preset is requested but custom-ops compile is unverified — hand off to quark-onnx-install to verify, then resume.
  • If the user picks an unusual combination (e.g. CLE + transformer, three or more algorithms) — explain why it's atypical and suggest the standard recipe, but respect the user's choice with a recorded risk.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Onnx Quant Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Onnx Quant Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Onnx Quant Plan this skillamd/Quark181—~4.8kAutomated safety check: PassMIT
Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassCustom licence
Embedded AI Deploymentmatlab/agent-skills-playground1811 repos~3.4kAutomated safety check: PassCustom licence
Astreawarpfront/hipfire653—~2.6kAutomated safety check: PassCustom licence
Vss Setup Behavior AnalyticsNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.7kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Matlab Use Visual Inspection

    matlab/matlab-agentic-toolkit

    Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

    1.1k GitHub stars~3.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Embedded AI Deployment

    matlab/agent-skills-playground

    Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).

    181 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Astrea

    warpfront/hipfire

    A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…

    653 GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Vss Setup Behavior Analytics

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses to deploy the vss-behavior-analytics service standalone (entrypoint, config-source, optional calibration).

    1.9k GitHub stars~2.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Benchmark Tune

    Mesh-LLM/mesh-llm

    A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

    3.5k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 11 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 11 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 11 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 11 days ago
    Auto-check passed

Works with

Questions about Quark Onnx Quant Plan

What does Quark Onnx Quant Plan do?

Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. Quark Onnx Quant Plan is an agent skill from amd/Quark.json and user intent.

When should I use Quark Onnx Quant Plan?

Quark Onnx Quant Plan fits situations like: the user needs preset selection (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX / MXFP …); calibration method choice (MinMax / Entropy / Percentile / Distribution / NonOverflow / MinMSE / LayerWisePercentile); algorithm selection (CLE / AdaRound / AdaQuant / BiasCorrection / AutoMixprecision); deployment-target gating (CPU / CUDA / ROCm / AMD NPU CNN / AMD NPU Transformer).

How do I install Quark Onnx Quant Plan in Claude Code?

Run `npx skills add amd/Quark --skill quark-onnx-quant-plan -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan in amd/Quark) into .claude/skills/quark-onnx-quant-plan in your project. Claude Code loads it when a task matches its description.

How do I install Quark Onnx Quant Plan in Codex?

Run `npx skills add amd/Quark --skill quark-onnx-quant-plan -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan in amd/Quark) into .agents/skills/quark-onnx-quant-plan in your project. Codex loads it when a task matches its description.

Can I use Quark Onnx Quant Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-onnx-quant-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-onnx-quant-plan, .gemini/skills/quark-onnx-quant-plan, .github/skills/quark-onnx-quant-plan and .opencode/skills/quark-onnx-quant-plan in your project.

What does Quark Onnx Quant Plan need to run?

SKILL.md names no scripts, command-line tools or credentials: Quark Onnx Quant Plan is instructions for the agent only.

Does Quark Onnx Quant Plan access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Onnx Quant Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Onnx Quant Plan use?

Quark Onnx Quant Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Onnx Quant Plan use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Onnx Quant Plan?

Skills that share tags, products or a category with Quark Onnx Quant Plan: Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Embedded AI Deployment (matlab/agent-skills-playground, 181 stars), Astrea (warpfront/hipfire, 653 stars) and Vss Setup Behavior Analytics (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Onnx Quant Plan?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.