Magpie Kernel Evaluator
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.
$ npx skills add amd/Quark --skill quark-torch-debug -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/Quark quark-torch-debug --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-debug .claude/skills/quark-torch-debug && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "quark-torch-debug" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-debug into .claude/skills/quark-torch-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-debug", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-debugType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/Quark --skill quark-torch-debug -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/Quark quark-torch-debug --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-debug .agents/skills/quark-torch-debug && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "quark-torch-debug" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-debug into .agents/skills/quark-torch-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-debug", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-debug -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/Quark quark-torch-debug --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-debug .cursor/skills/quark-torch-debug && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "quark-torch-debug" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-debug into .cursor/skills/quark-torch-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-debug", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/Quark.git --path .claude/skills-impl/l1-atomic/torch/quark-torch-debug--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/Quark --skill quark-torch-debug -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/Quark quark-torch-debug --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-debug .gemini/skills/quark-torch-debug && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "quark-torch-debug" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-debug into .gemini/skills/quark-torch-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-debug", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/Quark quark-torch-debugInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/Quark --skill quark-torch-debug -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-debug .github/skills/quark-torch-debug && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-debug" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-debug into .github/skills/quark-torch-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-debug", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-debug -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/Quark quark-torch-debug --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-debug .opencode/skills/quark-torch-debug && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-debug" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-debug into .opencode/skills/quark-torch-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-debug", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
quark-torch-debugDiagnose failed Quark installation, PTQ execution, script generation, or export attempts.
Quark Torch Debug is an agent skill from amd/Quark. Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. Use when the user reports an error, stack trace, invalid artifact, missing dependency, CUDA OOM, version mismatch, or unexpected PTQ results. Trigger for "Quark error", "PTQ failed", "quantization crashed", "CUDA out of memory", "import error", "model loading failed", "wrong results", any Python traceback mentioning quark/torch/transformers, or when the user pastes an error message related to Quark workflows.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Debugging, LLM inference and serving and Deep learning. It works with CUDA, Python and PyTorch. The licence is MIT.
Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Quark Torch Debug loads about 1.9k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 226 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
.torch.kernel` | Missing C++ compiler | `sudo apt install build-essential` |Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 226 words, ~1,944 tokens.
.claude/skills/quark-torch-debug/SKILL.md (or your agent's skills folder).Convert failures into a structured diagnostic report with the smallest safe recovery path. Debugging Quark issues is tricky because errors can originate from many layers — Python environment, PyTorch, transformers, CUDA/ROCm drivers, model architecture, or Quark itself. This skill systematically narrows down the root cause.
env_context.json, pytorch_install_result.json, quark_install_result.json (optional, for environment and install state)Diagnostic report with root cause, evidence, and the smallest safe fix.
Schema: validation_report.schema.json
# Debug Report
## Common Error Patterns
### Installation Errors
| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ModuleNotFoundError: No module named 'quark'` | Quark not installed or wrong Python env | `pip install amd-quark` or activate correct conda env |
| `ImportError: quark.torch.kernel` | Missing C++ compiler | `sudo apt install build-essential` |
| `torch.cuda.is_available() == False` | CPU-only PyTorch installed | Reinstall PyTorch with correct `--index-url` |
| `RuntimeError: CUDA error: no kernel image` | PyTorch CUDA version ≠ system CUDA | Match PyTorch build to system CUDA version |
### Model Loading Errors
| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ValueError: Unrecognized model in config` | `trust_remote_code` needed | Add `--trust_remote_code` flag |
| `OSError: Can't load tokenizer` | Missing tokenizer files or sentencepiece | `pip install sentencepiece` and check model path |
| `ImportError: ... requires transformers>=X.Y.Z` | Transformers version too old | `pip install transformers==X.Y.Z` |
| `OutOfMemoryError` during loading | Model too large for single GPU | Use `--multi_gpu auto` or `--multi_device` |
### Quantization Errors
| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `CUDA out of memory` during quantization | Insufficient GPU memory | Reduce `--num_calib_data` or `--batch_size`, or use `--multi_gpu` |
| `RuntimeError: expected scalar type Half` | Data type mismatch | Set `--data_type float16` or `bfloat16` explicitly |
| `KeyError: 'model.layers.0...'` | Model architecture not matching template | Check `model_type` in config.json matches Quark template |
| `ValueError: ... is not a valid quantization scheme` | Typo in scheme name | Check against the 21 supported schemes |
| `AssertionError` in AWQ/GPTQ | Algorithm config mismatch | Verify `--quant_algo_config_file` matches model architecture |
### Export Errors
| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ModuleNotFoundError: gguf` | GGUF package missing | `pip install gguf>=0.10.0` |
| `ONNX export failed` | Model has unsupported ops | Try `--model_export hf_format` instead |
| `PermissionError` on output dir | No write permission | Check output directory permissions |
### Transformers Compatibility
Quark checks compatibility before quantization. Known issues:
- `seen_tokens` removed in transformers > 4.53.3
- `get_max_length` removed in transformers > 4.48.3
- `get_usable_length` removed in transformers > 4.53.3
- General constraint: `transformers < 5.3`
## Diagnostic Process
1. **Read the error** — get the exact error message, full stack trace, and the command that was run.
2. **Identify the layer** — is this an install problem, model loading, quantization runtime, or export issue?
3. **Check the environment** — Python version, PyTorch version, CUDA/ROCm version, Quark version, transformers version.
4. **Match against known patterns** — use the tables above.
5. **Propose the fix** — smallest change that resolves the issue without side effects.
## Diagnostic Commands
```bash
# Environment snapshot
python -c "
import sys; print('Python:', sys.version)
try:
import torch; print('PyTorch:', torch.__version__, 'CUDA:', torch.version.cuda, 'HIP:', torch.version.hip)
except: print('PyTorch: not installed')
try:
import quark; print('Quark:', quark.__version__)
except: print('Quark: not installed')
try:
import transformers; print('Transformers:', transformers.__version__)
except: print('Transformers: not installed')
"
# GPU memory status
nvidia-smi --query-gpu=memory.used,memory.total --format=csv,noheader 2>/dev/null
rocm-smi --showmemuse 2>/dev/null
# Check Quark compatibility
python -c "
from quark.torch.utils.llm.compatibility import check_compatibility_before_quantization
print('Compatibility check available')
"CUDA out of memory during quantization of Qwen/Qwen3-8B with FP8
Single GPU (24GB) insufficient for FP8 quantization with 512 calibration samples
Reduce calibration data or use multi-GPU:
# Option A: Reduce calibration samples
python quantize_quark.py ... --num_calib_data 64 --batch_size 1
# Option B: Use multi-GPU (if available)
python quantize_quark.py ... --multi_gpu autoFor models > 7B parameters on GPUs with < 48GB VRAM, start with --num_calib_data 64 and increase if memory allows.
## Interaction Flow
1. **Gather evidence**: Get the error message, stack trace, command, and environment info.
2. **Classify**: Determine which layer failed (install / load / quantize / export).
3. **Diagnose**: Match against known patterns and run diagnostic commands if needed.
4. **Propose fix**: Present the smallest change that resolves the issue.
5. **Confirm**: Get user approval before any environment changes.
6. **Verify**: After the fix, re-run the failing step to confirm resolution.
## Recovery
- If the root cause is upstream (transformers API change, PyTorch bug), hand off to `quark-torch-doc-drift-check` or `quark-torch-skill-sync`.
- If the fix requires a PyTorch reinstall (wrong build, version mismatch, CPU instead of GPU), hand off to `quark-torch-install` with the specific requirement noted.
- If the fix requires a Quark package or dependency change, hand off to `quark-install` with the specific requirement noted.
- If the error is in a custom model or unsupported architecture, suggest registering a custom template via `LLMTemplate.register_template()`.© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-debug of amd/Quark.
Open the folder on GitHubat commit 313cb0b
Quark Torch Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Quark Torch Debug this skillamd/Quark | 181 | — | ~1.9k | Automated safety check: Notes | MIT | |
| Magpie Kernel Evaluatoramd/skills | 395 | — | ~2.3k | Automated safety check: Pass | MIT | |
| The Art of Debuggingstas00/the-art-of-debugging | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Ako4allTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT | |
| Paddle Op DevPaddlePaddle/Paddle | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
evo-design/proto-tools
Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…
amd/Quark
Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.
amd/Quark
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.
amd/Quark
Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/Quark
Install or verify the AMD Quark package and its dependencies.
amd/Quark
L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…
Categories
Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. Quark Torch Debug is an agent skill from amd/Quark. Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.
Quark Torch Debug fits situations like: the user reports an error; invalid artifact; missing dependency; version mismatch.
Run `npx skills add amd/Quark --skill quark-torch-debug -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-debug in amd/Quark) into .claude/skills/quark-torch-debug in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/Quark --skill quark-torch-debug -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-debug in amd/Quark) into .agents/skills/quark-torch-debug in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-debug, .gemini/skills/quark-torch-debug, .github/skills/quark-torch-debug and .opencode/skills/quark-torch-debug in your project.
Going by SKILL.md and its folder, Quark Torch Debug needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Quark Torch Debug is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Quark Torch Debug: Magpie Kernel Evaluator (amd/skills, 395 stars), The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and Ako4all (TongmingLAIC/AKO4ALL, 369 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.
Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.