Agent skill

Quark Onnx Ptq Workflow

by amd in amd/Quark

End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output.

MITAuto-check passedAI & LLM Engineering

Install Quark Onnx Ptq Workflow

skills CLI
$ npx skills add amd/Quark --skill quark-onnx-ptq-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-onnx-ptq-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/onnx/quark-onnx-ptq-workflow .claude/skills/quark-onnx-ptq-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-onnx-ptq-workflow
GitHub stars
181
Token cost
~4.5k tokens
SKILL.md length
1,751 words
Files
2
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output.

  • Works in 4 steps: Model Intake → Quantization Plan → Manifest Generation → …
  • The user wants a complete ONNX-to-ONNX PTQ pipeline: model intake
  • SKILL.md covers Purpose, Inputs, Outputs: run_manifest.yaml and Interaction Flow, plus 9 more sections
  • Calls python3

What it does

Quark Onnx Ptq Workflow is an agent skill from amd/Quark. End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output. Use when the user wants a complete ONNX-to-ONNX PTQ pipeline: model intake, quantization planning, calibration-script generation, manifest, and confirmed execution. Trigger for "quantize my .onnx", "run ONNX PTQ end to end", "full ONNX quantization pipeline", "quantize yolov8/resnet50/yolonas with XINT8/A8W8/BFP16/MXFP", "weights-only INT4 for my .onnx LLM", or any request that spans more than one…

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `example-xint8-yolov8n.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving, Performance reviews and Computer vision. It works with ONNX and Python. The licence is MIT.

When your agent uses it

  • The user wants a complete ONNX-to-ONNX PTQ pipeline: model intake
  • Quantization planning
  • Calibration-script generation
  • Confirmed execution

Example prompts

  • “quantize my .onnx”
  • “run ONNX PTQ end to end”
  • “full ONNX quantization pipeline”
  • “/quark-onnx-ptq-workflow”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Model Intake
  2. Quantization Plan
  3. Manifest Generation
  4. Execute PTQ

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Onnx Ptq Workflow loads about 4.5k tokens when it runs. Until then it costs about 184 tokens; SKILL.md has 1,751 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~184
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,751 words, ~4,497 tokens.

Download SKILL.mdSave it as .claude/skills/quark-onnx-ptq-workflow/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
quark-onnx-ptq-workflow
description
End-to-end ONNX PTQ workflow for AMD Quark — from a `.onnx` file (and calibration data) to a quantized `.onnx` output. Use when the user wants a complete ONNX-to-ONNX PTQ pipeline: model intake, quantization planning, calibration-script generation, manifest, and confirmed execution. Trigger for "quantize my .onnx", "run ONNX PTQ end to end", "full ONNX quantization pipeline", "quantize yolov8/resnet50/yolo_nas with XINT8/A8W8/BFP16/MXFP*", "weights-only INT4 for my .onnx LLM", or any request that spans more than one ONNX PTQ step. When in doubt between routing to an atomic skill vs. the workflow, prefer this workflow if the user's request implies they want to go from a `.onnx` file to a quantized output.
layer
l2-workflows
primary_artifact
run_manifest.yaml
source_knowledge
examples/onnx/yolo_quantization/quantize_yolo.py, tutorials/onnx/image_classification/onnx_image_classification_tutorial.ipynb…

quark-onnx-ptq-workflow

📘 Quick-start example — read this first. A fully worked end-to-end walkthrough is available at example-xint8-yolov8n.md (XINT8 quantization of YOLOv8n for AMD NPU CNN deployment, calibrated on COCO val2017). It is the fastest way to see exactly what this workflow produces — open it alongside this SKILL.md before running anything.

Purpose

Chain the ONNX PTQ path — model intake → quantization planning → script + manifest generation → confirmed execution — while keeping the user informed at each checkpoint. This workflow orchestrates the atomic ONNX skills so the user does not have to manually chain them. Unlike the Torch flow there is no single shipped quantize_quark.py for ONNX; the workflow generates a small standalone Python script in the user's working directory that imports from quark.onnx, then runs that script after the user confirms.

📂 Worked example: see example-xint8-yolov8n.md for the full YOLOv8n + XINT8 + NPU-CNN walkthrough referenced throughout the steps below.

Inputs

  • Input .onnx model path (with optional sibling .onnx_data external-weights file)
  • Calibration data: folder of representative samples (or CalibrationDataReader Python class)
  • Output .onnx path (and optional .onnx_data if use_external_data_format=True)
  • User goal: scheme (XINT8 / A8W8 / A16W8 / BF16 / BFP16 / MX* / MXFP*), deployment target (CPU / CUDA / ROCm / AMD NPU CNN / AMD NPU Transformer), accuracy target
  • session_context.json with constraints.backend = "onnx" for user goal and constraints
  • env_context.json for hardware / execution-provider facts
  • workspace_context.json for validated paths
  • onnx_install_result.json and quark_install_result.json to confirm runtime is ready

Outputs: run_manifest.yaml

Records the generated calibration/quantization script path, the exact python3 invocation, and the resolved QConfig. Side artifacts: the generated script in the user's working directory, the quantized .onnx (and .onnx_data if external) in the user's output directory. The manifest is built in Step 3.

Schema: run_manifest.schema.json

Interaction Flow

  1. Intake — call quark-onnx-model-intake to produce model_analysis.json
  2. Plan — call quark-onnx-quant-plan to produce quant_plan.json
  3. Manifest — generate a standalone quantization script in the user's working directory + run_manifest.yaml, then stop for user approval
  4. Execute — run the confirmed script and report the artifact paths

CRITICAL RULES

  1. NEVER call ModelQuantizer.quantize_model(...) directly from this workflow. Always generate a standalone script the user reviews first.
  2. NEVER skip a step. Even if the user provides all details upfront, execute each step in order.
  3. STOP at every checkpoint and wait for user confirmation before continuing.
  4. Show concrete output at each step — tables, JSON, the full script body, the exact command — not just prose descriptions.
  5. NEVER modify Quark's own source code, examples, or tutorials. The Quark repo (quark/, examples/, tutorials/, tools/, docs/, tests/) is read-only from this workflow's perspective. See Upstream Quark Code is Read-Only below.
  6. NEVER silently fall back to CPU execution. If the user requested CUDA/ROCm and the provider is unavailable, stop and surface the gap to quark-onnx-install / quark-onnx-debug — do not quietly degrade.

Upstream Quark Code is Read-Only

The Quark repository is the upstream source of truth. This workflow may read it freely (config sources, example scripts, tutorial notebooks, custom-op headers) but must not write into it. That includes:

  • No edits to quark/ package source.
  • No edits to anything under examples/onnx/ — including examples/onnx/yolo_quantization/quantize_yolo.py. Even small "just to make the script accept my preprocessing" patches are forbidden, because they make the run irreproducible against a clean Quark install.
  • No edits to anything under tutorials/onnx/ — the Ryzen AI tutorials for YOLOv8 / ResNet-50 / image classification are reference material only.
  • No edits to tools/, docs/, tests/, or pyproject.toml / requirements.txt.

If a shipped example or tutorial as-is does not cover what the user needs (different model, different preprocessing, different exclude list, different calibration source), write a fresh standalone script in the user's working directory (or /tmp/) that imports from quark.onnx. Pattern:

python
# user_workspace/my_onnx_ptq.py — NOT inside the Quark repo
from quark.onnx import ModelQuantizer, QConfig, QLayerConfig, XInt8Spec, CLEConfig
from onnxruntime.quantization.calibrate import CalibrationDataReader
# ... user-specific data reader + QConfig the shipped example doesn't cover ...

Then reference that script in the run_manifest.yaml instead of a modified example. The shipped scripts stay untouched, the user's customization is local to their workspace, and the run remains reproducible against any Quark version.

If the user actually needs an upstream Quark change (a new preset, a new algorithm config), surface it as a Quark contribution — do not silently patch their local checkout.

Required Artifact Flow

text
Step 1 (Intake)   ──► model_analysis.json
Step 2 (Plan)     ──► quant_plan.json
Step 3 (Manifest) ──► run_manifest.yaml  +  user_workspace/<name>_ptq.py
Step 4 (Execute)  ──► quantized .onnx (+ .onnx_data if external)  (only after user says yes)

Step 1: Model Intake

Goal: Analyze the .onnx graph and produce model_analysis.json. Hand off to quark-onnx-model-intake.

Actions
  1. Confirm the input path. It must end in .onnx. If a sibling .onnx_data exists, both must move together.
  2. Run quark-onnx-model-intake. That skill produces:
    • opset / IR version
    • input / output names, shapes, dtypes
    • op-type histogram and quantizable-op count
    • whether QDQ nodes are already present (if yes — model is already quantized; warn and stop)
    • external-data status (model.onnx_data present? total bytes? >2 GB?)
    • deployment-target compatibility (CPU / CUDA / ROCm / NPU CNN / NPU Transformer)
  3. Surface risks:
    • Already quantized → stop, do not re-quantize.
    • opset < 13 → recommend upgrading to opset 17+ for best support.
    • Dynamic input shapes → calibration data reader must produce matching shapes; warn if shapes vary.
    • 2 GB model → external data must be handled; set use_external_data_format=True later.

    • NPU CNN target → only XINT8 / A8W8 are gated as supported; warn early.
Output to Show User

Present a summary table:

text
Model Analysis:
  Model path:       ./models/yolov8n.onnx
  Opset / IR:       17 / 8
  Input:            images, [1, 3, 640, 640], float32
  Output:           output0, [1, 84, 8400], float32
  Total ops:        ~226 nodes (Conv, MatMul, Add, Mul, Sigmoid, Concat, …)
  Quantizable ops:  ~118 (Conv + MatMul + …)
  QDQ already?      No
  External data:    No (6.2 MB inline)
  Risks:            None
  NPU CNN target:   Compatible (Conv-heavy, no unsupported ops)
>>> CHECKPOINT 1: Confirm model analysis is correct before continuing

Step 2: Quantization Plan

Goal: Build quant_plan.json from the model analysis and user's stated preferences. Hand off to quark-onnx-quant-plan.

Actions
  1. Determine the preset. If the user stated one (e.g., "XINT8"), use it. Otherwise, recommend based on deployment target + priority:

    DeploymentPriorityRecommended PresetAlgorithm
    AMD NPU CNN (Ryzen AI)Best accuracyXINT8 + EnableNPUCnn=TrueCLE (optional AdaRound)
    AMD NPU TransformerBest accuracyBFP16none
    CPU generalSmallest modelA8W8CLE (optional)
    CPU generalBest accuracyA16W8MinMax
    CUDA / ROCmBest accuracyBF16none
    CUDA / ROCmHigh throughputBFP16 / MXFP4none
    AnyRecover lost accuracy+ AdaRound / AdaQuantas additional algorithm
  2. Fill the decision table. Show ALL decisions with defaults:

    DecisionValueReason
    presetXINT8User requested XINT8 (NPU CNN)
    activation_specXInt8Spec()Matches preset
    weight_specXInt8Spec()Matches preset
    calibration_methodMinMax (default)Standard for XINT8
    algo_config[CLEConfig()]Improves XINT8 accuracy on Conv networks
    EnableNPUCnnTrueXINT8 + NPU CNN deployment
    use_external_data_formatFalseModel < 2 GB
    exclude_nodes / exclude_subgraphs[]None identified
    calibration_data_path./calib_data/User-provided folder of representative samples
    num_calib_data100Standard default for vision
    batch_size1Safe default
    evaluation_intentsmokeQuick mAP / Prec@1 check after quantization
  3. Map deployment-target gates explicitly. Note any preset+target combinations that are unsupported (e.g., BFP16 on NPU CNN). Surface them, do not silently change the user's choice.

  4. Ask the user if they want to change anything.

>>> CHECKPOINT 2: User MUST confirm or adjust the plan before continuing

Wait for the user to say "ok", "confirm", "looks good", "continue", or similar. If they request changes (e.g., "add AdaRound", "raise calib data to 500"), update the table and re-present.


Step 3: Manifest Generation

Goal: Translate the confirmed plan into a runnable standalone script in the user's working directory and a run_manifest.yaml. Do not write into the Quark repo.

Show full SKILL.md (708 more words)Show less
Actions
  1. Decide the script path. Place it in the user's working directory, e.g. ./<model_name>_ptq.py. Never write under examples/onnx/ or tutorials/onnx/.

  2. Build the script body. Three pieces:

    a. CalibrationDataReader — pick the right reader for the input modality. Vision models follow the pattern from tutorials/onnx/ryzen_ai/yolov8/ and tutorials/onnx/image_classification/:

    python
    from onnxruntime.quantization.calibrate import CalibrationDataReader
    import onnxruntime as ort
    import numpy as np, os, cv2
    
    class ImageDataReader(CalibrationDataReader):
        def __init__(self, calib_folder, model_path, hw=(640, 640)):
            sess = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
            self.input_name = sess.get_inputs()[0].name
            self.data = self._load(calib_folder, *hw)
            self.iter = None
        def _load(self, folder, h, w):
            out = []
            for f in sorted(os.listdir(folder)):
                if not f.lower().endswith((".jpg", ".jpeg", ".png")): continue
                img = cv2.imread(os.path.join(folder, f))
                img = cv2.resize(img, (w, h))
                arr = img.transpose(2, 0, 1).astype(np.float32) / 255.0
                out.append(np.expand_dims(arr, 0))
            return out
        def get_next(self):
            if self.iter is None:
                self.iter = iter([{self.input_name: d} for d in self.data])
            return next(self.iter, None)
        def rewind(self): self.iter = None

    b. QConfig built from the plan:

    python
    from quark.onnx import (
        ModelQuantizer, QConfig, QLayerConfig,
        XInt8Spec, Int8Spec, Int16Spec, BFloat16Spec, BFP16Spec,
        CLEConfig, AdaRoundConfig, AdaQuantConfig, CalibMethod,
    )
    activation_spec = XInt8Spec()
    weight_spec     = XInt8Spec()
    algo_config     = [CLEConfig()]
    config = QConfig(
        global_config=QLayerConfig(activation=activation_spec, weight=weight_spec),
        algo_config=algo_config,
        EnableNPUCnn=True,
        use_external_data_format=False,
        exclude=[],
    )

    c. Driver that wires the data reader + config + I/O paths:

    python
    dr = ImageDataReader("./calib_data", "./models/yolov8n.onnx")
    ModelQuantizer(config).quantize_model(
        "./models/yolov8n.onnx",
        "./models/yolov8n_xint8.onnx",
        dr,
    )
  3. Write the script to disk at the chosen path. Print the full script body back to the user so they can see exactly what will run — never just say "generated a script".

  4. Build the exact command:

    bash
    python3 ./yolov8n_ptq.py
  5. Show expected output layout:

    text
    ./models/
      ├── yolov8n.onnx                  (input, unchanged)
      └── yolov8n_xint8.onnx            (quantized output)
    ./yolov8n_ptq.py                    (generated by this workflow)
    ./run_manifest.yaml                 (this workflow's manifest)

    If use_external_data_format=True, also expect yolov8n_xint8.onnx_data next to the .onnx.

>>> CHECKPOINT 3: Show the script body and the command, then ask "shall I run this?"

Do NOT proceed to execution unless the user explicitly confirms. Acceptable confirmations: "yes", "run it", "go", "execute", or similar.

If the user says "no" or wants changes, go back to the relevant step (most commonly back to Step 2 to change preset / algorithm / exclude lists).


Step 4: Execute PTQ

Goal: Run the generated script and report results.

Precondition

This step runs ONLY after the user explicitly confirms in Step 3.

Actions
  1. Create the output directory if needed:

    bash
    mkdir -p ./models
  2. Run the generated script. Monitor for the common ONNX-side failures:

    • ONNX Runtime EP not available → hand off to quark-onnx-install / quark-onnx-debug.
    • Custom-op library load failure (BFPQuantizeDequantize, MXQuantizeDequantize, Extended*) → hand off to quark-onnx-debug.
    • OOM during calibration → suggest reducing num_calib_data, reducing batch size, or moving calibration to CPU (OptimDevice="cpu").
    • External-data not found → confirm .onnx_data sits next to .onnx.
    • AdaRound / AdaQuant diverged → suggest EarlyStop=True, lower learning rate, more iterations.
    • Calibration shapes mismatch → confirm the data reader produces the model's declared input shape.
  3. After completion, verify outputs exist:

    bash
    ls -lh ./models/yolov8n_xint8.onnx*
  4. Report results:

    text
    Quantization complete:
      Output:          ./models/yolov8n_xint8.onnx
      Model size:      ~3.5 MB (input was 12.3 MB → ~3.5× smaller)
      Format:          ONNX (QDQ inserted, com.amd.quark custom ops where applicable)
      External data:   No
Error Recovery
  • If quantization fails, do NOT retry blindly. Report the error verbatim and hand off to quark-onnx-debug.
  • For OOM, suggest: reduce num_calib_data first, then drop batch_size to 1, then move calibration to CPU.
  • For EP issues, do not "fix" by silently switching providers — surface the gap and let quark-onnx-install resolve it.

Complete Example

⭐ Recommended starting point for new users.

The full end-to-end walkthrough — XINT8 quantization of YOLOv8n for AMD NPU CNN deployment, calibrated on COCO val2017 — is in example-xint8-yolov8n.md sitting next to this SKILL.md. It shows every checkpoint (intake → plan → manifest → execute) with real numbers, the generated script, and the resulting quantized model layout. Mirror it for your own CNN.

Recovery

  • If any upstream artifact is missing, stop and name the missing producer skill. Do not improvise a partial artifact.
  • If the workflow reaches a blocked state (e.g., model already quantized, preset incompatible with deployment target, calibration data not in the model's input shape), report the blocker and suggest the specific fix.
  • If execution fails, report the error with diagnostic context and hand off to quark-onnx-debug rather than attempting ad-hoc patches.
  • If the user wants to change a decision mid-workflow (e.g., switch from XINT8 to BFP16 after seeing model analysis), go back to the relevant step — do not restart from scratch.
  • If the user pasted a Torch traceback or a HuggingFace path, do not try to repair-route — hand off to quark-torch-router and stop.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills-impl/l2-workflows/onnx/quark-onnx-ptq-workflow of amd/Quark.

  • SKILL.md
  • example-xint8-yolov8n.md

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Onnx Ptq Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Onnx Ptq Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Onnx Ptq Workflow this skillamd/Quark181—~4.5kAutomated safety check: PassMIT
Onboard Jetpack5 Inference BackendsEGalahad/sim2real145—~1.1kAutomated safety check: PassNone
Running Openmed Ondevicemaziyarpanahi/openmed5.5k—~2kAutomated safety check: PassApache-2.0
Tao Finetune ClipNVIDIA/skills3.5k—~4kAutomated safety check: NotesApache-2.0
Tao Port Huggingface ModelNVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

    145 GitHub stars~1.1k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Running Openmed Ondevice

    maziyarpanahi/openmed

    Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows.

    5.5k GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Tao Finetune Clip

    NVIDIA/skills

    Official

    CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

    3.5k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Official

    Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

    3.5k GitHub stars~4.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Vlm Model

    intel/auto-round

    Official

    Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

    1.6k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 12 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 12 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 12 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 12 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Quark Onnx Ptq Workflow

What does Quark Onnx Ptq Workflow do?

End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output. Quark Onnx Ptq Workflow is an agent skill from amd/Quark.onnx output.

When should I use Quark Onnx Ptq Workflow?

Quark Onnx Ptq Workflow fits situations like: the user wants a complete ONNX-to-ONNX PTQ pipeline: model intake; quantization planning; calibration-script generation; confirmed execution.

How do I install Quark Onnx Ptq Workflow in Claude Code?

Run `npx skills add amd/Quark --skill quark-onnx-ptq-workflow -a claude-code`. Or copy the skill folder (.claude/skills-impl/l2-workflows/onnx/quark-onnx-ptq-workflow in amd/Quark) into .claude/skills/quark-onnx-ptq-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Quark Onnx Ptq Workflow in Codex?

Run `npx skills add amd/Quark --skill quark-onnx-ptq-workflow -a codex`. Or copy the skill folder (.claude/skills-impl/l2-workflows/onnx/quark-onnx-ptq-workflow in amd/Quark) into .agents/skills/quark-onnx-ptq-workflow in your project. Codex loads it when a task matches its description.

Can I use Quark Onnx Ptq Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-onnx-ptq-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-onnx-ptq-workflow, .gemini/skills/quark-onnx-ptq-workflow, .github/skills/quark-onnx-ptq-workflow and .opencode/skills/quark-onnx-ptq-workflow in your project.

What does Quark Onnx Ptq Workflow need to run?

Going by SKILL.md and its folder, Quark Onnx Ptq Workflow needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Quark Onnx Ptq Workflow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Onnx Ptq Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Onnx Ptq Workflow use?

Quark Onnx Ptq Workflow is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Onnx Ptq Workflow use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Onnx Ptq Workflow?

Skills that share tags, products or a category with Quark Onnx Ptq Workflow: Onboard Jetpack5 Inference Backends (EGalahad/sim2real, 145 stars), Running Openmed Ondevice (maziyarpanahi/openmed, 5.5k stars), Tao Finetune Clip (NVIDIA/skills, 3.5k stars) and Tao Port Huggingface Model (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Onnx Ptq Workflow?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.