Agent skill

Quark Onnx Router

by amd in amd/Quark

Route Quark ONNX user goals to the correct atomic skill. An agent skill from amd/Quark.

MITAuto-check passedAI & LLM Engineering

Install Quark Onnx Router

skills CLI
$ npx skills add amd/Quark --skill quark-onnx-router -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-onnx-router --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-router .claude/skills/quark-onnx-router && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-onnx-router
GitHub stars
181
Token cost
~2.7k tokens
SKILL.md length
1,189 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Route Quark ONNX user goals to the correct atomic skill. An agent skill from amd/Quark.

  • Works in 5 steps: Confirm the backend. If the user… → Extract the core intent. Strip filler… → Check for L0 prerequisites. Before… → …
  • A user describes an ONNX quantization task in plain language — such as install onnxruntime
  • SKILL.md covers Purpose, Inputs, Outputs: session_context.json and CRITICAL ROUTING RULE, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quark Onnx Router is an agent skill from amd/Quark. Route Quark ONNX user goals to the correct atomic skill. Use when a user describes an ONNX quantization task in plain language — such as "install onnxruntime", "analyze my .onnx model", "choose a preset for my YOLO model", "plan ONNX PTQ", "quantize this .onnx with XINT8/BFP16/MXFP4", "validate my quantized .onnx", "debug a failed ONNX quantization", or any request that involves Quark's ONNX-to-ONNX flow. This is the ONNX-side entry point — trigger it whenever the user's intent involves Quark ONNX and the correct…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving, Computer vision and Plain language and style rules. It works with ONNX. The licence is MIT.

When your agent uses it

  • A user describes an ONNX quantization task in plain language — such as install onnxruntime
  • Analyze my .onnx model
  • Choose a preset for my YOLO model
  • Quantize this .onnx with XINT8/BFP16/MXFP4

Example prompts

  • “install onnxruntime”
  • “analyze my .onnx model”
  • “choose a preset for my YOLO model”
  • “/quark-onnx-router”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the backend. If the user mentions .onnx, quantize_static, ModelQuantizer, an ONNX Runtime execution provider, opset, or QDQ, treat…
  2. Extract the core intent. Strip filler and figure out what the user actually needs. "I want to quantize yolov8n.onnx with XINT8 for CPU" →…
  3. Check for L0 prerequisites. Before routing to an L1 skill, check if downstream skills need facts that are missing
  4. Pick the smallest fit. If the user only wants to inspect a .onnx graph, route to quark-onnx-model-intake — do not start the full workflow…
  5. Handle ambiguity honestly. If the goal is unclear, record the likely routes in open_questions and ask. Example: "I want to use Quark with…

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Onnx Router loads about 2.7k tokens when it runs. Until then it costs about 145 tokens; SKILL.md has 1,189 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,189 words, ~2,658 tokens.

Download SKILL.mdSave it as .claude/skills/quark-onnx-router/SKILL.md (or your agent's skills folder).
name
quark-onnx-router
description
Route Quark ONNX user goals to the correct atomic skill. Use when a user describes an ONNX quantization task in plain language — such as "install onnxruntime", "analyze my .onnx model", "choose a preset for my YOLO model", "plan ONNX PTQ", "quantize this .onnx with XINT8/BFP16/MXFP4", "validate my quantized .onnx", "debug a failed ONNX quantization", or any request that involves Quark's ONNX-to-ONNX flow. This is the ONNX-side entry point — trigger it whenever the user's intent involves Quark ONNX and the correct downstream skill is not immediately obvious.
layer
l1-atomic
primary_artifact
session_context.json
source_knowledge
docs/source/install.rst, docs/source/onnx/basic_usage_onnx.rst, docs/source/onnx/onnx_examples.rst, examples/onnx/model_support.md

quark-onnx-router

Purpose

Translate a user's natural-language ONNX quantization goal into the smallest correct skill boundary. The router exists because Quark's ONNX flow spans install, intake, planning, execution, debug, and validation — picking the wrong skill wastes time and produces wrong artifacts. The router is the sole producer of session_context.json for the ONNX backend; every other artifact (env, workspace, install results, model analysis, quant plan) is produced by its own owning skill, and the router only carries forward references to those files.

Inputs

  • User goal stated in natural language
  • env_context.json (optional, when routing depends on hardware / execution-provider facts)
  • workspace_context.json (optional, when routing depends on validated .onnx / .onnx_data paths)

Outputs: session_context.json

Carries the user goal, selected workflow (or atomic skill if no workflow applies), constraints with backend = "onnx", *_ref pointers to other artifacts, and unresolved questions.

Schema: session_context.schema.json

json
{
  "user_goal": "Quantize ./models/yolov8n.onnx with XINT8 and validate against the FP32 baseline",
  "workflow": "quark-onnx-ptq-workflow",
  "constraints": {
    "backend": "onnx",
    "offline": false,
    "execution_mode": "interactive_execute"
  },
  "env_context_ref": null,
  "workspace_context_ref": null,
  "onnx_install_result_ref": null,
  "quark_install_result_ref": null,
  "model_analysis_ref": null,
  "quant_plan_ref": null,
  "open_questions": [
    "Deployment target (CPU / CUDA / ROCm / NPU CNN / NPU Transformer) not yet confirmed — affects preset gating in quark-onnx-quant-plan"
  ]
}

Set each *_ref field to the path of the corresponding artifact once its producer skill has run. The router never embeds hardware, workspace, install, model, or plan facts inline — those belong in their owning artifacts. Always set constraints.backend = "onnx" so downstream skills and any future cross-backend orchestrator can disambiguate from the Torch flow.

CRITICAL ROUTING RULE

Vision / CNN .onnx PTQ requests MUST route to quark-onnx-ptq-workflow.

Signals: model name contains yolo, resnet, mobilenet, efficientnet, etc.; inputs are image tensors [N, C, H, W]; preset is XINT8 / A8W8 / A16W8 / BF16 / BFP16; user mentions mAP / Prec@1 / NPU CNN / Ryzen AI.

Examples:

  • "quantize yolov8n.onnx with XINT8 for Ryzen AI" → quark-onnx-ptq-workflow
  • "quantize resnet50 .onnx with A8W8" → quark-onnx-ptq-workflow
  • "run ONNX PTQ on my image-classification model" → quark-onnx-ptq-workflow

NEVER invoke quark.onnx.ModelQuantizer.quantize_model(...) directly from within the router. Execution belongs inside the workflow, after a confirmed quant_plan.json and a generated script the user has reviewed (the workflow's 4-step flow: intake → plan → manifest+script → confirmed execution).

Skill Map

User IntentTarget SkillWhy
Quantize a .onnx vision/CNN model (YOLO / ResNet / MobileNet / …)quark-onnx-ptq-workflowVision-specific: image data reader, CNN presets (XINT8 / A8W8 / BFP16), EnableNPUCnn, mAP / Prec@1 eval
Full end-to-end ONNX PTQ: from .onnx to quantized .onnxquark-onnx-ptq-workflowMulti-step workflow orchestration
Install ONNX Runtime, fix CPU/GPU variant, missing CUDA/ROCm EPquark-onnx-installONNX Runtime install is separate from Quark package install
Install Quark, set up Quark dependenciesquark-installBackend-neutral Quark package install (assumes ORT already set up)
Inspect a .onnx model, check opset / IR / op-type histogram, NPU compatibilityquark-onnx-model-intakeModel facts are prerequisites for planning
Choose preset (XINT8 / A8W8 / BFP16 / MX* / MatMulNBits), calibration method, algorithmquark-onnx-quant-planPlanning is separate from execution
Debug a failed ONNX quantization, ORT EP error, custom-op load failure, calibration crashquark-onnx-debugError recovery has its own diagnostic flow
Validate quantized .onnx, verify QDQ insertion, check non-quantized initializersquark-onnx-result-validatorPost-quantization byte-level + structural validation
Check environment, detect GPU, verify setupquark-env-preflightL0 fact collection (shared backend-neutral)
Validate .onnx / .onnx_data paths, output directoryquark-workspace-validateL0 path validation (shared backend-neutral)

Routing Logic

  1. Confirm the backend. If the user mentions .onnx, quantize_static, ModelQuantizer, an ONNX Runtime execution provider, opset, or QDQ, treat the request as ONNX. If they mention HuggingFace, safetensors, transformers, torch._dynamo, or quantize_quark.py, hand off to quark-torch-router instead — never silently mix backends.

  2. Extract the core intent. Strip filler and figure out what the user actually needs. "I want to quantize yolov8n.onnx with XINT8 for CPU" → vision PTQ → quark-onnx-ptq-workflow.

  3. Check for L0 prerequisites. Before routing to an L1 skill, check if downstream skills need facts that are missing:

    • If hardware / accelerator / execution-provider facts are needed but unknown → run quark-env-preflight first.
    • If the .onnx (and any sibling .onnx_data) path or output directory needs validation → run quark-workspace-validate first.
  4. Pick the smallest fit. If the user only wants to inspect a .onnx graph, route to quark-onnx-model-intake — do not start the full workflow. If they only want to install ORT, route to quark-onnx-install. But if they say "quantize my .onnx", use quark-onnx-ptq-workflow.

  5. Handle ambiguity honestly. If the goal is unclear, record the likely routes in open_questions and ask. Example: "I want to use Quark with my ONNX model" — does that mean inspect, plan, quantize end-to-end, or validate an already-quantized file?

Show full SKILL.md (529 more words)Show less

Rules

  • Do not guess deployment targets, EP availability, or paths. If the deployment target (CPU / CUDA / ROCm / NPU CNN / NPU Transformer) is unknown, do not assume CPU — gate it as an open question. Preset gating in quark-onnx-quant-plan depends on this.
  • Do not load .onnx files in the router. The router only reads paths and string intent; loading the model is quark-onnx-model-intake's job (especially relevant for >2 GB models with external data).
  • Explain the routing. Tell the user why you chose a particular skill: "Since you want to quantize a .onnx, I'll start with model intake, then we'll pick a preset, then I'll generate a script you confirm before running."
  • Reuse existing context. If a session_context.json already exists from a previous step, read it and carry forward — do not start from scratch.
  • Refuse cross-backend mixing. Never feed a .onnx into quark-torch-*; never feed safetensors into quark-onnx-*. If the user's input doesn't match this router's backend, hand off to quark-torch-router.

Interaction Flow

  1. Listen: Restate the user's goal in one sentence to confirm understanding and confirm the backend is ONNX.
  2. Assess: Check what facts are already known vs. missing. Do L0 skills need to run first? Is the deployment target known?
  3. Route: Name the target skill (or first link in the chain) and explain why it is the right fit. For PTQ requests, state the chain explicitly and that execution will only happen after the user confirms the generated script.
  4. Hand off: Produce the initial session_context.json (with constraints.backend = "onnx") and pass control to the chosen skill.

Recovery

  • If the goal is ambiguous, present the 2-3 most likely interpretations and ask the user to pick.
  • If required inputs are missing, list them explicitly — do not route to a downstream skill with known gaps that will immediately fail (e.g. starting quark-onnx-quant-plan without a model_analysis.json).
  • If L0 facts are unresolved, hand off a partial session_context.json with the gaps documented rather than forcing a premature routing decision.
  • If the user pasted a Torch traceback or a HuggingFace path, do not try to repair-route — hand off to quark-torch-router and stop.

Examples

Example 1: Clear vision PTQ request

User: "Quantize ./models/yolov8n.onnx with XINT8 for Ryzen AI deployment, output to ./output/yolov8n-xint8.onnx" Router: → quark-onnx-ptq-workflow (vision task family: YOLO + XINT8 + NPU CNN) Explain: "I'll quantize your YOLOv8 ONNX model with the vision PTQ workflow: 4 steps (intake → plan → generated script + manifest → confirmed execution)."

Example 2: ONNX Runtime install request

User: "I need to install onnxruntime-rocm 7.1" Router: → quark-onnx-install (clear ORT install intent, EP stated)

Example 3: Debug request

User: "quantize_static is failing with 'CUDAExecutionProvider not available'" Router: → quark-onnx-debug (ORT EP error, established diagnostic flow)

Example 4: Validation-only request

User: "Here's the quantized output — did the QDQ insertion actually happen?" Router: → quark-onnx-result-validator

Example 5: Ambiguous request

User: "I want to use Quark with my ONNX model" Router: Ask — "Do you want to: (a) inspect the graph's opset and op-types, (b) plan and run a full PTQ, (c) install ONNX Runtime / Quark first, or (d) validate an already-quantized output?"

Example 6: Cross-backend mistake

User: "Quantize Qwen/Qwen3-8B with FP8" (HuggingFace repo id, no .onnx) Router: Hand off → quark-torch-router and stop. Do not attempt to route this through the ONNX flow.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/onnx/quark-onnx-router of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Onnx Router next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Onnx Router compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Onnx Router this skillamd/Quark181—~2.7kAutomated safety check: PassMIT
Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassCustom licence
Contextpilot SavingsEfficientContext/ContextPilot140—~1.4kAutomated safety check: PassMIT
Tao Finetune ClipNVIDIA/skills3.5k—~4kAutomated safety check: NotesApache-2.0
Tao Port Huggingface ModelNVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Matlab Use Visual Inspection

    matlab/matlab-agentic-toolkit

    Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

    1.1k GitHub stars~3.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Contextpilot Savings

    EfficientContext/ContextPilot

    A skill your agent uses when a user asks how many tokens (or how much context/cost) ContextPilot has saved, or wants a ContextPilot savings status/summary inside Hermes Agent — e.g.

    140 GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Tao Finetune Clip

    NVIDIA/skills

    Official

    CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

    3.5k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Official

    Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

    3.5k GitHub stars~4.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vision

    gridaco/grida

    Query images with a local Ollama vision model without loading the image into the main agent context.

    2.7k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 12 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 12 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 12 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 12 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Quark Onnx Router

What does Quark Onnx Router do?

Route Quark ONNX user goals to the correct atomic skill. An agent skill from amd/Quark. Quark Onnx Router is an agent skill from amd/Quark. Route Quark ONNX user goals to the correct atomic skill.

When should I use Quark Onnx Router?

Quark Onnx Router fits situations like: A user describes an ONNX quantization task in plain language — such as install onnxruntime; analyze my .onnx model; choose a preset for my YOLO model; quantize this .onnx with XINT8/BFP16/MXFP4.

How do I install Quark Onnx Router in Claude Code?

Run `npx skills add amd/Quark --skill quark-onnx-router -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-router in amd/Quark) into .claude/skills/quark-onnx-router in your project. Claude Code loads it when a task matches its description.

How do I install Quark Onnx Router in Codex?

Run `npx skills add amd/Quark --skill quark-onnx-router -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-router in amd/Quark) into .agents/skills/quark-onnx-router in your project. Codex loads it when a task matches its description.

Can I use Quark Onnx Router in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-onnx-router -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-onnx-router, .gemini/skills/quark-onnx-router, .github/skills/quark-onnx-router and .opencode/skills/quark-onnx-router in your project.

What does Quark Onnx Router need to run?

SKILL.md names no scripts, command-line tools or credentials: Quark Onnx Router is instructions for the agent only.

Does Quark Onnx Router access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Onnx Router safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Onnx Router use?

Quark Onnx Router is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Onnx Router use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Onnx Router?

Skills that share tags, products or a category with Quark Onnx Router: Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Contextpilot Savings (EfficientContext/ContextPilot, 140 stars), Tao Finetune Clip (NVIDIA/skills, 3.5k stars) and Tao Port Huggingface Model (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Onnx Router?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.