Matlab Use Visual Inspection
matlab/matlab-agentic-toolkit
Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.
Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.
$ npx skills add amd/Quark --skill quark-onnx-quant-plan -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/Quark quark-onnx-quant-plan --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan .claude/skills/quark-onnx-quant-plan && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "quark-onnx-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan into .claude/skills/quark-onnx-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-onnx-quant-plan", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-planType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/Quark --skill quark-onnx-quant-plan -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/Quark quark-onnx-quant-plan --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan .agents/skills/quark-onnx-quant-plan && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "quark-onnx-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan into .agents/skills/quark-onnx-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-onnx-quant-plan", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-onnx-quant-plan -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/Quark quark-onnx-quant-plan --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan .cursor/skills/quark-onnx-quant-plan && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "quark-onnx-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan into .cursor/skills/quark-onnx-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-onnx-quant-plan", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/Quark.git --path .claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/Quark --skill quark-onnx-quant-plan -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/Quark quark-onnx-quant-plan --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan .gemini/skills/quark-onnx-quant-plan && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "quark-onnx-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan into .gemini/skills/quark-onnx-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-onnx-quant-plan", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/Quark quark-onnx-quant-planInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/Quark --skill quark-onnx-quant-plan -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan .github/skills/quark-onnx-quant-plan && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "quark-onnx-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan into .github/skills/quark-onnx-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-onnx-quant-plan", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-onnx-quant-plan -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/Quark quark-onnx-quant-plan --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan .opencode/skills/quark-onnx-quant-plan && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "quark-onnx-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan into .opencode/skills/quark-onnx-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-onnx-quant-plan", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
quark-onnx-quant-planBuild a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.
Quark Onnx Quant Plan is an agent skill from amd/Quark. Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. Use when the user needs preset selection (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX / MXFP …), calibration method choice (MinMax / Entropy / Percentile / Distribution / NonOverflow / MinMSE / LayerWisePercentile), algorithm selection (CLE / AdaRound / AdaQuant / BiasCorrection / AutoMixprecision), deployment-target gating (CPU / CUDA / ROCm / AMD NPU CNN / AMD NPU Transformer), op-type include/exclude lists, weights-only INT4…
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Performance reviews, LLM inference and serving and Deployment. It works with ONNX and CUDA. The licence is MIT.
9 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json and python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Quark Onnx Quant Plan loads about 4.8k tokens when it runs. Until then it costs about 235 tokens; SKILL.md has 1,721 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,721 words, ~4,808 tokens.
.claude/skills/quark-onnx-quant-plan/SKILL.md (or your agent's skills folder).Convert an ONNX model_analysis.json plus the user's intent into a confirmed quant_plan.json.
This skill makes the quantization decisions for the ONNX-to-ONNX flow — which preset, which
calibration method, which algorithm, which op types to include/exclude, whether to enable an NPU
target, whether to use external-data — without generating scripts or executing the
quantization. The plan is the contract between the user's intent and the execution step.
model_analysis.json from quark-onnx-model-intake (architecture, op histogram, opsets,
external-data state, preset_candidates, risks)env_context.json for accelerator-aware preset gating (CUDA major / ROCm major / NPU presence)onnx_install_result.json (optional) — gates BFP/MX/Extended presets (require the custom-ops
library to be compiled for the target EP)Records the chosen preset (or custom config), calibration method, algorithm list, layer/op
overrides, NPU flag, external-data setting, and evaluation intent. Shares the schema with the
Torch plan; ONNX-specific fields live under onnx_specific.
Schema: quant_plan.schema.json
{
"model": {
"model_type": "onnx",
"architecture_guess": "cnn",
"analysis_ref": "./model_analysis.json"
},
"backend": "onnx",
"deployment_target": "npu_cnn",
"preset": "XINT8",
"calibration": {
"method": "PowerOfTwo_MinMSE",
"data_size": 200,
"batch_size": 1,
"use_external_data_format": false,
"optimize_mem": false,
"worker_num": 1
},
"algorithms": ["CLE"],
"onnx_specific": {
"enable_npu_cnn": true,
"enable_npu_transformer": false,
"include_cle": true,
"include_fast_ft": false,
"op_types_to_quantize": null,
"nodes_to_quantize": null,
"nodes_to_exclude": null,
"subgraphs_to_exclude": [],
"extra_options": {
"OpTypesToExcludeOutputQuantization": []
},
"use_external_data_format": false,
"execution_providers": ["CPUExecutionProvider"]
},
"evaluation_intent": "smoke",
"requires_confirmation": false
}DefaultConfigMapping)Quark ONNX ships 50+ named presets in quark/onnx/quantization/config/custom_config.py. Pick the
smallest viable set for the user's architecture + deployment target; never list them all.
enable_npu_cnn=True, PoF2 scales, NHWC)| Preset | Description | Picks |
|---|---|---|
XINT8 | INT8 input + INT8 weight, optimized for NPU | Default for any CNN targeting NPU |
XINT8_ADAROUND | XINT8 + AdaRound fast-finetune | When base XINT8 loses accuracy |
XINT8_ADAQUANT | XINT8 + AdaQuant fast-finetune | When AdaRound is not enough |
VINT8 | INT8 optimized for VAIML | VAIML deployment |
enable_npu_transformer=True)| Preset | Description | Picks |
|---|---|---|
INT16_TRANSFORMER_DEFAULT | INT16 activations + INT8 weights, fast | Outlier-heavy activations |
INT16_TRANSFORMER_ACCURATE | INT16 + accuracy algorithms | Largest accuracy headroom |
| Preset | Description | Picks |
|---|---|---|
A8W8 | INT8 sym activations + INT8 sym weights | Standard CPU/GPU INT8 |
A8W8_ADAROUND / A8W8_ADAQUANT | + fast-finetune | Accuracy-tight A8W8 |
A16W8 | INT16 sym activations + INT8 sym weights | Wide-activation needs |
A16W8_ADAROUND / A16W8_ADAQUANT | + fast-finetune | Accuracy-tight A16W8 |
S8S8_AAWS / U8S8_AAWS / U8U8_AAWA / S16S8_ASWS / U16S8_AAWS (+ ADAROUND / ADAQUANT variants) | Various sym/asym INT8/INT16 combos | Fine-tune sym/asym choice per ORT target |
INT8_CNN_DEFAULT / INT8_CNN_ACCURATE / INT16_CNN_DEFAULT / INT16_CNN_ACCURATE | CNN-tuned INT8/INT16 | CPU/GPU CNN deployment |
| Preset | Description | Picks |
|---|---|---|
FP16 / FP16_ADAQUANT | FP16 W+A | Accuracy-first when INT is too lossy |
BF16 / BF16_ADAQUANT | BFloat16 W+A | Same, with BF16 range |
| Preset | Description | Picks |
|---|---|---|
BFP16 / BFP16_ADAQUANT | Block Floating Point 16-bit | AMD accelerator deployments |
MX4 / MX6 / MX9 (+ ADAQUANT) | Micro-Exponents BFP variants | Bit-budget exploration |
MXFP4E2M1 / MXFP6E2M3 / MXFP6E3M2 / MXFP8E4M3 / MXFP8E5M2 / MXINT8 (+ ADAQUANT) | OCP MX formats | Modern AMD/NVIDIA accelerators |
| Preset | Description | Picks |
|---|---|---|
BF16_BFP16 / BF16_MIXED_BFP16 / BF16_MIXED_BFP16_ADAQUANT | BF16 + BFP16 hybrid | High-accuracy + AMD HW |
BF16_MXINT8 / BF16_MIXED_MXINT8 / BF16_MIXED_MXINT8_ADAQUANT | BF16 + MXInt8 hybrid | Same, OCP MX flavor |
MX9_INT8 | MX9 + INT8 hybrid | Bit-budget exploration |
S16S16_MIXED_S8S8 | INT16 + INT8 mixed | Outlier-aware INT mix |
From quark/onnx/calibration/methods.py (+ ORT built-ins):
| Method | When to pick |
|---|---|
MinMax | Default for most CNN / weights-only; cheap, deterministic |
Entropy | KL-divergence-based; helps when activations have long tails |
Percentile | Clip outliers at a chosen percentile (default 99.999) |
Distribution | Distribution-matching; useful for FP8 p3/same |
LayerWisePercentile | Auto-picks per-tensor optimal percentile (MAE/MSE) — AMD-specific |
PowerOfTwo_NonOverflow (a.k.a. NonOverflow) | Required for AMD NPU XINT8 — picks the smallest PoF2 scale that doesn't overflow |
PowerOfTwo_MinMSE (a.k.a. MinMSE) | Required for AMD NPU XINT8 — picks the PoF2 scale that minimizes MSE; usually better than NonOverflow |
Int16Method.MinMax | For INT16 configs |
NPU CNN / XINT8 targets must use a PowerOfTwo* method. Non-PoF2 scales are rejected at NPU
runtime — flag this as a hard constraint in the plan.
From quark/onnx/quantization/config/algorithm.py and examples/onnx/accuracy_improvement/:
| Algorithm | Kind | Compatible presets | Description |
|---|---|---|---|
CLE (Cross-Layer Equalization) | Pre | INT8 CNN configs | Folds BN, equalizes per-channel scales across consecutive Conv/Linear layers (Nagel et al., 2019) |
BiasCorrection | Post | INT8 CNN | Post-hoc bias adjustment (Nagel et al., 2019) |
AdaRound | Post (fast-finetune) | XINT8 / A8W8 / A16W8 / block formats | Adaptive rounding optimization; needs cal data + LR + iterations; GPU-accelerated |
AdaQuant | Post (fast-finetune) | Same as AdaRound | Layer-wise calibration tuning; usually after AdaRound is not enough |
AutoMixprecision | Post | Block formats / mixed-precision | Auto-selects sensitivity-based per-layer dtype; can do dual BFP16+MX hybrid |
Algorithms compose: e.g. CLE (pre) + AdaRound (post) is a common XINT8 recipe. Combinations
beyond two algorithms are usually a red flag — flag them in risks.
Help the user choose based on their priorities and the architecture from
model_analysis.json.model.onnx_specific.architecture_guess:
| User intent | Architecture | Recommended starting plan |
|---|---|---|
| "Best accuracy, AMD GPU" | any | BF16 or BF16_MIXED_BFP16 |
| "INT8 CNN, CPU/GPU deployment" | cnn | INT8_CNN_DEFAULT or A8W8; add CLE if accuracy drops |
| "INT8 CNN → Ryzen AI NPU" | cnn | XINT8 + CLE, calibration = PowerOfTwo_MinMSE; usually NHWC pre-conversion via quark.onnx.tools.convert_nchw_to_nhwc |
| "Ryzen AI NPU, accuracy-tight CNN" | cnn | XINT8_ADAROUND (then XINT8_ADAQUANT if still short) |
| "Block format MXFP4 / BFP16 experimentation" | any | BFP16 or MXFP4E2M1; require quark.onnx.operators.custom_ops to be compiled |
| "Hybrid mixed-precision for best size/accuracy" | any | BF16_MIXED_BFP16 or S16S16_MIXED_S8S8 + AutoMixprecision |
| Target | Preset must satisfy | Calibration must be | Notes |
|---|---|---|---|
npu_cnn (Ryzen AI CNN) | enable_npu_cnn=True, PoF2 symmetric INT8 per-tensor, NCHW→NHWC done | PowerOfTwo_MinMSE or PowerOfTwo_NonOverflow | Reject A8W8 / BFP16 / MX* / FP16 if user requests npu_cnn |
npu_transformer (Ryzen AI Transformer) | enable_npu_transformer=True, INT8/INT16 QDQ on MatMul/Gemm | MinMax / Percentile typically | Reject CNN-only presets |
cuda | onnxruntime-gpu present, CUDAExecutionProvider available | Any | Block-format presets require custom-ops |
rocm | onnxruntime_rocm (ROCm 6.x) or CPU ORT on ROCm 7.x (tools/ci/install_onnxruntime.sh) | Any | Custom-ops library must compile for ROCm |
cpu | Any | Any | Block-format presets work via CPU custom-ops; expect speed cost |
If the deployment target conflicts with a requested preset, the plan must either (a) downgrade
to a viable preset and explain, or (b) leave it unset with a high-severity risk in quant_plan.json.
Three commonly-used knobs the plan should expose:
op_types_to_quantize — restrict QDQ insertion to a subset, e.g. ["Conv"] for
CNN-only quantization.nodes_to_quantize / nodes_to_exclude — surgical per-node control by graph node name.
Node names change after pre-processing (NCHW→NHWC, BN folding, etc.), so resolve them after
any pre-processing pass.extra_options["OpTypesToExcludeOutputQuantization"] — keep certain op outputs in float
while still quantizing their inputs/weights.quant_plan.onnx_specific.extra_optionsFrom real examples in examples/onnx/:
| Option | Typical value | Source example |
|---|---|---|
SimplifyModel | True / False | toggle OnnxSlim pre-pass |
QuantizeFP16 | True | FP16-input models |
OpTypesToExcludeOutputQuantization | ["Add", "Mul"] etc. | keep selected op outputs in float |
FastFinetune | {"DataSize": 200, "BatchSize": 2, "NumIterations": 1000, "LearningRate": …, "OptimAlgorithm": "adaround"/"adaquant", "OptimDevice": "cuda:0"/"cpu", "InferDevice": "cuda:0"/"cpu", "EarlyStop": True} | AdaRound / AdaQuant tutorials and Auto-Search tutorials |
ALWAYS present this table and WAIT for confirmation before finalizing. Fill the "Value" column from the user's request, the model analysis, and the deployment-target gates above.
| Decision | Value | Reason |
|---|---|---|
backend | onnx | Fixed for this skill |
deployment_target | (fill: cpu / cuda / rocm / npu_cnn / npu_transformer) | (from env or user) |
preset | (fill: name from DefaultConfigMapping or "custom") | (why) |
calibration.method | (fill: MinMax / Entropy / Percentile / Distribution / LayerWisePercentile / PowerOfTwo_MinMSE / PowerOfTwo_NonOverflow) | (why) |
calibration.data_size | 200 (default) | (why) |
calibration.batch_size | 1–4 | (why) |
algorithms | (fill: list e.g. ["CLE"], ["AdaRound"], ["AdaQuant"], ["BiasCorrection"]) | (why) |
onnx_specific.enable_npu_cnn | (fill: bool) | Hard-tied to deployment_target == "npu_cnn" |
onnx_specific.enable_npu_transformer | (fill: bool) | Hard-tied to deployment_target == "npu_transformer" |
onnx_specific.include_cle | (fill: bool) | CNN INT8 default true |
onnx_specific.op_types_to_quantize | null or ["MatMul"] etc. | (why) |
onnx_specific.use_external_data_format | true if model > 2 GB (from model_analysis.json) | Required for large models |
onnx_specific.extra_options | (fill: dict) | (why — list each key) |
evaluation_intent | smoke (default), mlperf, mAP | (why) |
After showing the table, ask: "Confirm this plan? Any changes?"
Do NOT proceed until the user confirms.
For fine-grained control, individual layers can override the global config via QLayerConfig
(see examples/onnx/yolo_quantization/quantize_yolo.py):
from quark.onnx import QConfig, QLayerConfig, XInt8Spec, CLEConfig
config = QConfig(
global_config=QLayerConfig(activation=XInt8Spec(), weight=XInt8Spec()),
algo_config=[CLEConfig()],
EnableNPUCnn=True,
exclude=[
# YOLOX-style: keep a specific subgraph in float
(["/_head/_modules_list.14/Transpose"], ["/_head/_modules_list.14/Concat_9"]),
],
)Record any per-layer overrides under onnx_specific.subgraphs_to_exclude (list of
(start_nodes, end_nodes) tuples) or onnx_specific.nodes_to_exclude (flat name list).
model_analysis.json is missing, route back to quark-onnx-model-intake.XINT8 family + PowerOfTwo*
calibration + NHWC. NPU Transformer ⇒ INT8_TRANSFORMER_* / INT16_TRANSFORMER_*. Reject
conflicting preset requests with a clear explanation.{BFP16*, MX*, MXFP*, MXINT8*, BF16_*BFP*, BF16_*MXINT8*} requires the Quark custom-ops library to have compiled. If
onnx_install_result.json shows the compile failed, suppress these presets from
recommendations and emit a risk pointing back to quark-onnx-install.model_analysis.json.model.estimated_size_gb > 2,
set onnx_specific.use_external_data_format = true automatically. Otherwise default false.XINT8 without PoF2 calibration,
MXFP4 on a model that hasn't been validated with AutoMixprecision), keep the user's choice
but record the risk in the plan.model_analysis.json and quant_plan.json. The plan
references the analysis via analysis_ref.model_analysis.json available? If not, route to
quark-onnx-model-intake first. Is onnx_install_result.json available? If a block-format
preset is on the table, require it.session_context.json.constraints.deployment_target or
ask. Apply the hard gates from the table above.model_analysis.json.quantization_targets.onnx_specific.preset_candidates,
filter by deployment target and custom-op availability, present the top 1–3 with rationale.PowerOfTwo_MinMSE for
NPU CNN; Percentile for outlier-heavy transformers; consider LayerWisePercentile when
accuracy is critical). Decide data_size and batch_size.quant_plan.json. Surface any new constraints back to quark-onnx-router so
they land in session_context.json.model_analysis.json shows has_qdq_already=true or has_quark_custom_ops=true — do
not produce a plan. Route back to quark-onnx-model-intake with a recommendation to
remove existing QDQ first (quark.onnx.tools.remove_qdq).A8W8 uses non-PoF2 scales which the AMD NPU CNN runtime rejects"), suggest the
closest viable alternative (e.g. XINT8), and respect the user's choice if they insist —
recording the risk.model_analysis.json shows architecture unknown — produce a draft plan with
requires_confirmation: true, list the op histogram, and ask the user which of CNN /
transformer / hybrid path to follow.quark-onnx-install to verify, then resume.CLE + transformer, three or more algorithms)
— explain why it's atypical and suggest the standard recipe, but respect the user's choice
with a recorded risk.© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan of amd/Quark.
Open the folder on GitHubat commit 313cb0b
Quark Onnx Quant Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Quark Onnx Quant Plan this skillamd/Quark | 181 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit | 1.1k | — | ~3.1k | Automated safety check: Pass | Custom licence | |
| Embedded AI Deploymentmatlab/agent-skills-playground | 181 | 1 repos | ~3.4k | Automated safety check: Pass | Custom licence | |
| Astreawarpfront/hipfire | 653 | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| Vss Setup Behavior AnalyticsNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 |
matlab/matlab-agentic-toolkit
Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.
matlab/agent-skills-playground
Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).
warpfront/hipfire
A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses to deploy the vss-behavior-analytics service standalone (entrypoint, config-source, optional calibration).
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Mesh-LLM/mesh-llm
A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…
amd/Quark
Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.
amd/Quark
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.
amd/Quark
Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/Quark
Install or verify the AMD Quark package and its dependencies.
amd/Quark
L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…
Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. Quark Onnx Quant Plan is an agent skill from amd/Quark.json and user intent.
Quark Onnx Quant Plan fits situations like: the user needs preset selection (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX / MXFP …); calibration method choice (MinMax / Entropy / Percentile / Distribution / NonOverflow / MinMSE / LayerWisePercentile); algorithm selection (CLE / AdaRound / AdaQuant / BiasCorrection / AutoMixprecision); deployment-target gating (CPU / CUDA / ROCm / AMD NPU CNN / AMD NPU Transformer).
Run `npx skills add amd/Quark --skill quark-onnx-quant-plan -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan in amd/Quark) into .claude/skills/quark-onnx-quant-plan in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/Quark --skill quark-onnx-quant-plan -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-quant-plan in amd/Quark) into .agents/skills/quark-onnx-quant-plan in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-onnx-quant-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-onnx-quant-plan, .gemini/skills/quark-onnx-quant-plan, .github/skills/quark-onnx-quant-plan and .opencode/skills/quark-onnx-quant-plan in your project.
SKILL.md names no scripts, command-line tools or credentials: Quark Onnx Quant Plan is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Quark Onnx Quant Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Quark Onnx Quant Plan: Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Embedded AI Deployment (matlab/agent-skills-playground, 181 stars), Astrea (warpfront/hipfire, 653 stars) and Vss Setup Behavior Analytics (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.
Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.