Agent skill

Quark Onnx Debug

by amd in amd/Quark

Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts.

MITAuto-check passedAI & LLM Engineering

Install Quark Onnx Debug

skills CLI
$ npx skills add amd/Quark --skill quark-onnx-debug -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-onnx-debug --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/onnx/quark-onnx-debug .claude/skills/quark-onnx-debug && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-onnx-debug
GitHub stars
181
Token cost
~4.8k tokens
SKILL.md length
333 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts.

  • Works in 6 steps: Gather evidence: error message, stack… → Classify: install / custom-op / ORT… → Diagnose: match against known patterns… → …
  • The user reports an error
  • SKILL.md covers Purpose, Inputs, Outputs: validation_report.md and Interaction Flow, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quark Onnx Debug is an agent skill from amd/Quark. Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts. Use when the user reports an error, stack trace, invalid artifact, missing dependency, ORT execution-provider mismatch, silent CPU fallback, OOM during calibration, custom-op load failure (BFPQuantizeDequantize / MXQuantizeDequantize / Extended), or unexpected quantization results from the ONNX flow. Trigger for "Quark ONNX error", "onnxruntime error", "quantizestatic failed", "calibration crashed"…

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Debugging, Performance reviews and LLM inference and serving. It works with ONNX and Python. The licence is MIT.

When your agent uses it

  • The user reports an error
  • Invalid artifact
  • Missing dependency
  • ORT execution-provider mismatch

Example prompts

  • “Quark ONNX error”
  • “onnxruntime error”
  • “quantizestatic failed”
  • “/quark-onnx-debug”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Gather evidence: error message, stack trace, exact call, model size, config, EPs.
  2. Classify: install / custom-op / ORT runtime / model / calibration / config / algorithm /
  3. Diagnose: match against known patterns and run diagnostic commands if needed.
  4. Propose fix: present the smallest change that resolves the issue.
  5. Confirm: get user approval before reinstalls, custom-ops rebuilds, or long recalibrations.
  6. Verify: after the fix, re-run the failing step (or a smaller smoke variant) to confirm

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Onnx Debug loads about 4.8k tokens when it runs. Until then it costs about 230 tokens; SKILL.md has 333 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~230
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 333 words, ~4,766 tokens.

Download SKILL.mdSave it as .claude/skills/quark-onnx-debug/SKILL.md (or your agent's skills folder).
name
quark-onnx-debug
description
Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts. Use when the user reports an error, stack trace, invalid artifact, missing dependency, ORT execution-provider mismatch, silent CPU fallback, OOM during calibration, custom-op load failure (BFPQuantizeDequantize / MXQuantizeDequantize / Extended*), or unexpected quantization results from the ONNX flow. Trigger for "Quark ONNX error", "onnxruntime error", "quantize_static failed", "calibration crashed", "CUDAExecutionProvider not available", "ROCMExecutionProvider not available", "custom op library load failed", "model.onnx larger than 2GB", "external data not found", "AdaRound diverged", "GPTQ ONNX failed", "QuaRot failed", "NPU power-of-2 scale", any Python traceback mentioning quark.onnx / onnxruntime / onnx, or when the user pastes an error message related to Quark ONNX workflows.
layer
l1-atomic
primary_artifact
validation_report.md
source_knowledge
quark/onnx/quantization/quantize.py, quark/onnx/quantization/api.py, quark/onnx/quantization/input_check.py, quark/onnx/calibration/calibrators.py…

quark-onnx-debug

Purpose

Convert ONNX-flow failures into a structured diagnostic report with the smallest safe recovery path. Debugging Quark ONNX issues is tricky because errors can originate from many layers — Python environment, the onnx package version, the installed onnxruntime* variant, custom-op compilation, CUDA / ROCm / NPU EPs, calibration data plumbing, graph optimization, or the chosen quantization config itself. This skill systematically narrows down the root cause.

Inputs

  • Error message and stack trace from a failing run
  • The exact command or ModelQuantizer / quantize_static invocation that triggered it
  • env_context.json, onnx_install_result.json, quark_install_result.json (optional, for environment and install state)
  • The model path / size and (if relevant) the QuantizationConfig used

Outputs: validation_report.md

Diagnostic report with root cause, evidence, and the smallest safe fix.

Schema: validation_report.schema.json

markdown
# Debug Report

## Common Error Patterns

### Installation Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ModuleNotFoundError: No module named 'quark'` | Quark not installed or wrong Python env | `pip install amd-quark` or activate correct conda env |
| `ModuleNotFoundError: No module named 'onnxruntime'` | ONNX Runtime not installed | Hand off to `quark-onnx-install` |
| `ModuleNotFoundError: No module named 'onnx'` | `onnx` package not installed | `pip install "onnx>=1.21.0,<=1.22.0"` |
| `ImportError: quark.onnx.operators.custom_ops` / compile failure | Missing C++ compiler (Linux `g++`, Windows VS 2022) or `ROCM_PATH`/`CUDA_HOME` unset for GPU build | `apt install g++` (Linux), install VS 2022 (Windows), `export ROCM_PATH=/opt/rocm` or `export CUDA_HOME=/usr/local/cuda` |
| Both `onnxruntime` and `onnxruntime-gpu` (or `_rocm`) installed | Variant collision — wrong EP loads | `pip uninstall -y onnxruntime onnxruntime-gpu onnxruntime_rocm`, then reinstall the single correct variant via `quark-onnx-install` |
| `onnx` schema / `opset` errors at import | `onnx` outside `>=1.16.0,<=1.19.0` | Pin within the supported range |
| `onnxruntime` ABI / symbol errors when loading custom ops | ORT version outside `>=1.22.2,<=1.24.2` (custom-ops built against a different ABI) | Reinstall ORT within the supported range, then re-trigger custom-ops compile |

### ONNX Runtime / Execution Provider Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `get_available_providers()` lacks `CUDAExecutionProvider` | CPU `onnxruntime` installed on a CUDA box | Reinstall `onnxruntime-gpu` via `quark-onnx-install` |
| `get_available_providers()` lacks `ROCMExecutionProvider` on ROCm 6.x | Wrong variant (need `onnxruntime_rocm` from AMD Artifactory) | Reinstall `onnxruntime_rocm` via `quark-onnx-install` |
| `ROCMExecutionProvider` missing on ROCm 7.x | **By design** — ROCm 7.x falls back to CPU `onnxruntime` per `tools/ci/install_onnxruntime.sh` (build incompatibility) | Document the CPU-only fallback; do not attempt a ROCm wheel install |
| Silent CPU fallback (no GPU utilization during calibration) | EP not requested or unavailable | Pass `execution_providers=['CUDAExecutionProvider']` / `['ROCMExecutionProvider']` explicitly; verify with `get_available_providers()` |
| `Failed to load library libonnxruntime_providers_cuda.so` | CUDA toolkit version mismatch with `onnxruntime-gpu` build | Match `onnxruntime-gpu` to the system CUDA major (CUDA 11 → Azure DevOps index; CUDA 12/13 → pypi default) |
| `RuntimeError: ... onnxruntime::rocm::...` | ROCm driver / library mismatch | Verify `rocm-smi` matches the wheel's ROCm version |

### Custom Op Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `RuntimeError: Failed to load custom op library` | Custom-ops library was not compiled, or compiled against a different ORT version | Re-run `python -c "import quark.onnx.operators.custom_ops"` and inspect the compile output |
| `Op (BFPQuantizeDequantize) ... is not a registered function/op` | The session was created without registering Quark's custom-op library | Use `quark.onnx.ModelQuantizer` / `quantize_static`, or pass the custom-op `.so`/`.dll` via `SessionOptions.register_custom_ops_library()` |
| Custom-op symbol-not-found on GPU but works on CPU | GPU kernel of the custom op was not built (missing `ROCM_PATH` / `CUDA_HOME` at compile time) | Set the env var, delete the cached `.so`/`.dll`, re-import to recompile |

### Model Loading / Shape Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `FileNotFoundError: Input model file ... does not exist.` (`api.py:103`) | Wrong `model_input` path | Check absolute path to `model.onnx` |
| `onnx.onnx_cpp2py_export.checker.ValidationError` | Model fails ONNX schema check | Run `onnx.checker.check_model()` to localize; rebuild the model with a compatible opset |
| `Message ... exceeds 2GB` / ProtoBuf size error | Model > 2 GB without external data | Save with `save_as_external_data=True`; pass `use_external_data_format=True` to the quantizer |
| `external data file not found` | `.onnx` moved but `.onnx_data` / weight blobs left behind | Move the model directory as a whole, or re-export with external data adjacent |
| Shape-inference failure | Incomplete shapes in the model | Run `quark.onnx.tools.fix_shapes` / `onnx.shape_inference.infer_shapes` before quantization |

### Calibration Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ValueError: The data reader should implement the '__len__' method to provide the data size.` (`calibrators.py:239`) | Custom `CalibrationDataReader` missing `__len__` | Implement `__len__` returning sample count |
| `ValueError: No data is collected.` (`calibrators.py:253`) | Data reader yielded zero batches before exhaustion | Check reader's `get_next()` returns at least one batch |
| `TypeError: compute_data must return a TensorsData not <...>` (`calibrators.py:259`) | Custom calibrator subclass returned wrong type | Return `onnxruntime.quantization.calibrate.TensorsData` |
| `ValueError: No collector created and can't generate calibration data.` | `collect_data()` never called before `compute_data()` | Call `collect_data()` first, or use `ModelQuantizer.quantize_model()` which handles the order |
| `ValueError: Invalid averaging constant, which should not be < 0 or > 1.` (`calibrators.py:563`) | `moving_average=True` with out-of-range constant | Use a value in `[0, 1]`, typically `0.01` |
| `ValueError: Unsupported calibration method` (`calibrators.py:1237`) | Typo or unsupported `CalibrationMethod` | Check `CalibrationMethod` enum in `quark.onnx.quantization.config.config` |
| OOM during calibration | Activation cache too large | Set `optimize_mem=True` (disk cache) and/or `optimize_disk=True`; reduce `worker_num` if memory bound; reduce calibration sample count |

### Quantization Config Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ValueError: Only ExtendedQuantFormat.QDQ supports wide bits quantization types.` (`input_check.py:67`) | INT16/UINT16/INT32 with `QuantFormat.QOperator` | Switch to `ExtendedQuantFormat.QDQ` |
| `ValueError: Fast finetune does not support int4 or uint4.` (`input_check.py:118`) | AdaRound/AdaQuant + INT4 weights | Disable fast-finetune for INT4, or use GPTQ instead |
| `ValueError: Invalid quant overrides ... for tensor ...` (`input_check.py:86`) | Tensor name in `extra_options['QuantOverrides']` not in the graph | Verify tensor name with `onnx.load(...).graph` |
| `ValueError: The per-channel quant override ... can not be applied on <quant_type>` (`input_check.py:94`) | Per-channel override on a quant type that doesn't support it (e.g., FP types, block formats) | Remove per-channel from that override |
| `ValueError: For the crypto mode, the input model should be in onnx.ModelProto format.` (`input_check.py:137`) | `crypto_mode=True` with a file path | Load model first: `model = onnx.load(path)` |
| `ValueError: quantization config must be one of Config and QConfig.` (`api.py:73`) | Passed a dict or wrong type | Build a `quark.onnx.quantization.config.Config` or `QConfig` |
| `ValueError: Unexpected config name: <X>` (`custom_config.py:993`) | Typo in `get_default_config(<X>)` preset name | Pick from `DefaultConfigMapping` (XINT8, S8S8_AAWS, A8W8, A16W8, BF16, BFP16, MX4/6/9, MXFP8/6/4, …) |
| `nodes_to_quantize` ignored | Node names don't match graph after pre-processing (NCHW→NHWC, BN folding) | Run pre-processing first, dump the post-preprocess graph, then re-select node names |

### Algorithm Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| AdaRound / AdaQuant device mismatch | `optim_device='cuda:0'` but no `CUDAExecutionProvider` available | Either set `optim_device='cpu'` or fix the ORT install via `quark-onnx-install` |
| AdaRound / AdaQuant divergence (loss → NaN/Inf) | LR too high for the model, or activation outliers | Lower `learning_rate`, enable `CLE` pre-processing first |
| GPTQ shape mismatch | Group-size doesn't divide the weight dimension | Pick a `group_size` that divides the weight's reduction dim (commonly 32, 64, 128) |
| SmoothQuant alpha range | `SmoothAlpha` outside `[0, 1]` | Use a value in `[0.5, 0.85]` for LLMs |
| QuaRot rotation config invalid | Missing rotation pair definitions | Provide the rotation pairs in the algorithm config; see `quark/onnx/algorithm/quarot/` |
| `AutoSearchPro: 'model_input' can not be None.` (`auto_search_pro.py:139`) | Search invoked without a model | Pass the model path or `ModelProto` |
| `AutoSearchPro: 'calib_data_reader' can not be None.` (`auto_search_pro.py:142`) | Search invoked without calibration data | Provide a `CalibrationDataReader` |
| `AutoSearchPro: Unsupported search_algo` (`auto_search_pro.py:192`) | Sampler typo | Use `'TPE'` or `'Grid'` |

### NPU-Specific Errors (Ryzen AI / VAI)

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| Power-of-2 scale validation failure | Non-PoF2 scale produced for an NPU target | Use a PoF2 calibrator (`PowOfTwoCalibrater` MinMSE / NonOverflow); confirm `enable_npu_cnn=True` |
| `enable_npu_cnn=True` but model has unsupported ops | NPU CNN backend doesn't support the op | Fold / replace via `quark.onnx.tools.*`; see `optimizations/optimize.py` |
| NCHW vs NHWC mismatch | NPU expects NHWC but model is NCHW | Run `quark.onnx.tools.convert_nchw_to_nhwc` before quantization |
| `enable_dpu` deprecated warning | Old API | Switch to `enable_npu_cnn=True` |

### Export / Post-quantization Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `PermissionError` on output dir | No write permission | Check output directory permissions |
| Quantized model > 2 GB and external data not written | `use_external_data_format=False` on a large model | Set `use_external_data_format=True` |
| Inference using the quantized model errors on a custom op | Deployment env missing Quark's custom-ops library | Ship the compiled `.so`/`.dll` and register via `SessionOptions.register_custom_ops_library()` |

## Diagnostic Process

1. **Read the error** — capture the exact error message, full stack trace, and the command or
   `quantize_static` / `ModelQuantizer` call that triggered it.
2. **Identify the layer** — installation / ORT runtime / custom-op / model loading / calibration /
   config validation / algorithm / NPU-specific / export.
3. **Check the environment** — Python version, `onnx` version, installed `onnxruntime*` variants
   and versions, available EPs, Quark version, accelerator (CUDA major / ROCm major), whether the
   custom-ops library compiled.
4. **Match against known patterns** — use the tables above.
5. **Propose the fix** — smallest change that resolves the issue without side effects.

## Diagnostic Commands

```bash
# Environment snapshot
python -c "
import sys; print('Python:', sys.version.split()[0])
try:
    import onnx; print('onnx:', onnx.__version__)
except Exception as e: print('onnx: not installed', e)
try:
    import onnxruntime as ort
    print('onnxruntime:', ort.__version__)
    print('EPs:', ort.get_available_providers())
except Exception as e: print('onnxruntime: not installed', e)
try:
    import quark; print('Quark:', quark.__version__)
except Exception as e: print('Quark: not installed', e)
try:
    import quark.onnx; print('quark.onnx loaded from', quark.onnx.__file__)
except Exception as e: print('quark.onnx: load failed', e)
"

# Confirm only ONE onnxruntime variant is installed
pip list 2>/dev/null | grep -iE '^(onnx|onnxruntime|onnxslim|onnxscript|onnxruntime-genai|onnxruntime_rocm|onnxruntime-extensions)\b'

# Force custom-ops compile (first import) and capture failures
python -c "import quark.onnx.operators.custom_ops" 2>&1 | tail -30

# Model sanity check
python -c "
import onnx, sys
m = onnx.load(sys.argv[1])
print('opset:', [(o.domain or 'ai.onnx', o.version) for o in m.opset_import])
print('inputs:', [(i.name, [d.dim_value or d.dim_param for d in i.type.tensor_type.shape.dim]) for i in m.graph.input])
onnx.checker.check_model(m, full_check=False)
print('checker: OK')
" path/to/model.onnx

# GPU memory status
nvidia-smi --query-gpu=memory.used,memory.total --format=csv,noheader 2>/dev/null
rocm-smi --showmemuse 2>/dev/null
```

## Rules

- **Always ask for the full error message and the call that was run.** Partial errors lead to wrong
  diagnoses. For ONNX, also ask for: the `QuantizationConfig` / preset name, the model path and
  size, the chosen `execution_providers`, and the calibration data reader class.
- **Do not guess the fix.** Narrow down the root cause first, then propose a specific solution.
- **Distinguish "EP not available" from "EP available but not requested".** Both cause silent CPU
  fallback. Check `ort.get_available_providers()` *and* the providers actually passed to the
  session.
- **Confirm before mutating.** If the fix involves reinstalling `onnxruntime*`, rebuilding the
  custom-ops library, or re-running an hour-long calibration, present the plan and get
  confirmation.
- **Consider cascading effects.** Bumping `onnxruntime` may force a custom-ops rebuild and may break
  ABI with an older `onnx` pin. Bumping `onnx` outside `<=1.19.0` may break Quark's QDQ insertion.
  Always check version constraints (`requirements.txt`, `docs/source/install.rst`).
- **For ROCm 7.x with no ROCM EP, this is by design.** Do not propose a "fix" that downgrades to
  ROCm 6.x unless the user asks; instead document the CPU-fallback trade-off (see
  `tools/ci/install_onnxruntime.sh`).

## Example: CUDA EP missing during calibration

### Error Summary

`onnxruntime.capi.onnxruntime_pybind11_state.RuntimeException: ... CUDAExecutionProvider is not in
the list of available providers` raised from `ModelQuantizer.quantize_model()` on a CUDA 12 box.

### Root Cause

CPU `onnxruntime` was installed instead of `onnxruntime-gpu`; `get_available_providers()` returns
`['CPUExecutionProvider']` only.

### Evidence

- `pip list | grep onnxruntime` shows `onnxruntime 1.23.2` only (no `-gpu` variant).
- `python -c "import onnxruntime as ort; print(ort.get_available_providers())"` →
  `['CPUExecutionProvider']`.
- `nvidia-smi` reports a healthy CUDA 12.4 driver and an idle GPU.

### Fix

Reinstall the correct variant via `quark-onnx-install` (CUDA 12 path):

```bash
pip uninstall -y onnxruntime onnxruntime-gpu onnxruntime_rocm
pip install --no-cache-dir onnxruntime-gpu
python -c "import onnxruntime as ort; print(ort.get_available_providers())"   # expect CUDAExecutionProvider present
python -c "import quark.onnx.operators.custom_ops"                            # recompile against new ORT
```

### Prevention

Run `quark-onnx-install` (which routes through `quark-env-preflight`) before the first
quantization on a fresh environment, so the ORT variant is picked from the accelerator instead of
defaulting to CPU.

Interaction Flow

  1. Gather evidence: error message, stack trace, exact call, model size, config, EPs.
  2. Classify: install / custom-op / ORT runtime / model / calibration / config / algorithm / NPU / export.
  3. Diagnose: match against known patterns and run diagnostic commands if needed.
  4. Propose fix: present the smallest change that resolves the issue.
  5. Confirm: get user approval before reinstalls, custom-ops rebuilds, or long recalibrations.
  6. Verify: after the fix, re-run the failing step (or a smaller smoke variant) to confirm resolution.

Recovery

  • If the fix requires an onnxruntime* or onnx package change (wrong variant, version mismatch, missing variant), hand off to quark-onnx-install with the specific requirement noted.
  • If the fix requires an amd-quark package or shared-dependency change, hand off to quark-install with the specific requirement noted.
  • If the fix requires a C++ compiler install or ROCM_PATH/CUDA_HOME setup, surface the environment gap and let the user resolve it before re-running custom-ops compile.
  • If the root cause is upstream (ORT API change, onnx schema change), hand off to quark-doc-drift-check or quark-skill-sync.
  • If the error is in a custom config or unsupported op for an NPU target, suggest the relevant pre-processing tool from quark.onnx.tools (e.g. convert_nchw_to_nhwc, convert_qdq_to_qop, convert_a8w8_npu_to_a8w8_cpu).
  • If the issue cannot be reproduced from the supplied evidence, ask for the minimum repro (model, calibration data sample, config) before guessing.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/onnx/quark-onnx-debug of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Onnx Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Onnx Debug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Onnx Debug this skillamd/Quark181—~4.8kAutomated safety check: PassMIT
Onboard Jetpack5 Inference BackendsEGalahad/sim2real145—~1.1kAutomated safety check: PassNone
Running Openmed Ondevicemaziyarpanahi/openmed5.5k—~2kAutomated safety check: PassApache-2.0
Onnxtxtonnx/onnx22k—~1.3kAutomated safety check: PassApache-2.0
Engine Performancescragnog/HOT-Step-CPP170—~4.9kAutomated safety check: PassMIT
CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

    145 GitHub stars~1.1k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Running Openmed Ondevice

    maziyarpanahi/openmed

    Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows.

    5.5k GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Onnxtxt

    onnx/onnx

    Read or write ONNX text format ("onnxtxt"). An agent skill from onnx/onnx.

    22k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Engine Performance

    scragnog/HOT-Step-CPP

    Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

    170 GitHub stars~4.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 9 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 9 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 9 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 9 days ago
    Auto-check passed

Works with

Questions about Quark Onnx Debug

What does Quark Onnx Debug do?

Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts. Quark Onnx Debug is an agent skill from amd/Quark. Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts.

When should I use Quark Onnx Debug?

Quark Onnx Debug fits situations like: the user reports an error; invalid artifact; missing dependency; ORT execution-provider mismatch.

How do I install Quark Onnx Debug in Claude Code?

Run `npx skills add amd/Quark --skill quark-onnx-debug -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-debug in amd/Quark) into .claude/skills/quark-onnx-debug in your project. Claude Code loads it when a task matches its description.

How do I install Quark Onnx Debug in Codex?

Run `npx skills add amd/Quark --skill quark-onnx-debug -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/onnx/quark-onnx-debug in amd/Quark) into .agents/skills/quark-onnx-debug in your project. Codex loads it when a task matches its description.

Can I use Quark Onnx Debug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-onnx-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-onnx-debug, .gemini/skills/quark-onnx-debug, .github/skills/quark-onnx-debug and .opencode/skills/quark-onnx-debug in your project.

What does Quark Onnx Debug need to run?

SKILL.md names no scripts, command-line tools or credentials: Quark Onnx Debug is instructions for the agent only. Our summary lists: Python 3.

Does Quark Onnx Debug access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Onnx Debug safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Onnx Debug use?

Quark Onnx Debug is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Onnx Debug use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Onnx Debug?

Skills that share tags, products or a category with Quark Onnx Debug: Onboard Jetpack5 Inference Backends (EGalahad/sim2real, 145 stars), Running Openmed Ondevice (maziyarpanahi/openmed, 5.5k stars), Onnxtxt (onnx/onnx, 22k stars) and Engine Performance (scragnog/HOT-Step-CPP, 170 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Onnx Debug?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.