Agent skill

Quark Torch Debug

by amd in amd/Quark

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

MITAuto-check: notesAI & LLM Engineering

Install Quark Torch Debug

skills CLI
$ npx skills add amd/Quark --skill quark-torch-debug -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-debug --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-debug .claude/skills/quark-torch-debug && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-debug
GitHub stars
181
Token cost
~1.9k tokens
SKILL.md length
226 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

  • The user reports an error
  • SKILL.md covers Purpose, Inputs, Outputs: validation_report.md and Rules, plus 5 more sections
  • Calls python
  • Invalid artifact

What it does

Quark Torch Debug is an agent skill from amd/Quark. Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. Use when the user reports an error, stack trace, invalid artifact, missing dependency, CUDA OOM, version mismatch, or unexpected PTQ results. Trigger for "Quark error", "PTQ failed", "quantization crashed", "CUDA out of memory", "import error", "model loading failed", "wrong results", any Python traceback mentioning quark/torch/transformers, or when the user pastes an error message related to Quark workflows.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Debugging, LLM inference and serving and Deep learning. It works with CUDA, Python and PyTorch. The licence is MIT.

When your agent uses it

  • The user reports an error
  • Invalid artifact
  • Missing dependency
  • Version mismatch

Example prompts

  • “Quark error”
  • “PTQ failed”
  • “quantization crashed”
  • “/quark-torch-debug”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch Debug loads about 1.9k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 226 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:45
    .torch.kernel` | Missing C++ compiler | `sudo apt install build-essential` |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 226 words, ~1,944 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-debug/SKILL.md (or your agent's skills folder).
name
quark-torch-debug
description
Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. Use when the user reports an error, stack trace, invalid artifact, missing dependency, CUDA OOM, version mismatch, or unexpected PTQ results. Trigger for "Quark error", "PTQ failed", "quantization crashed", "CUDA out of memory", "import error", "model loading failed", "wrong results", any Python traceback mentioning quark/torch/transformers, or when the user pastes an error message related to Quark workflows.
layer
l1-atomic
primary_artifact
validation_report.md
source_knowledge
examples/torch/language_modeling/llm_ptq/quantize_quark.py, quark/torch/utils/llm/model_preparation.py, quark/torch/utils/llm/compatibility.py…

quark-torch-debug

Purpose

Convert failures into a structured diagnostic report with the smallest safe recovery path. Debugging Quark issues is tricky because errors can originate from many layers — Python environment, PyTorch, transformers, CUDA/ROCm drivers, model architecture, or Quark itself. This skill systematically narrows down the root cause.

Inputs

  • Error message and stack trace from a failing run
  • env_context.json, pytorch_install_result.json, quark_install_result.json (optional, for environment and install state)

Outputs: validation_report.md

Diagnostic report with root cause, evidence, and the smallest safe fix.

Schema: validation_report.schema.json

markdown
# Debug Report

## Common Error Patterns

### Installation Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ModuleNotFoundError: No module named 'quark'` | Quark not installed or wrong Python env | `pip install amd-quark` or activate correct conda env |
| `ImportError: quark.torch.kernel` | Missing C++ compiler | `sudo apt install build-essential` |
| `torch.cuda.is_available() == False` | CPU-only PyTorch installed | Reinstall PyTorch with correct `--index-url` |
| `RuntimeError: CUDA error: no kernel image` | PyTorch CUDA version ≠ system CUDA | Match PyTorch build to system CUDA version |

### Model Loading Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ValueError: Unrecognized model in config` | `trust_remote_code` needed | Add `--trust_remote_code` flag |
| `OSError: Can't load tokenizer` | Missing tokenizer files or sentencepiece | `pip install sentencepiece` and check model path |
| `ImportError: ... requires transformers>=X.Y.Z` | Transformers version too old | `pip install transformers==X.Y.Z` |
| `OutOfMemoryError` during loading | Model too large for single GPU | Use `--multi_gpu auto` or `--multi_device` |

### Quantization Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `CUDA out of memory` during quantization | Insufficient GPU memory | Reduce `--num_calib_data` or `--batch_size`, or use `--multi_gpu` |
| `RuntimeError: expected scalar type Half` | Data type mismatch | Set `--data_type float16` or `bfloat16` explicitly |
| `KeyError: 'model.layers.0...'` | Model architecture not matching template | Check `model_type` in config.json matches Quark template |
| `ValueError: ... is not a valid quantization scheme` | Typo in scheme name | Check against the 21 supported schemes |
| `AssertionError` in AWQ/GPTQ | Algorithm config mismatch | Verify `--quant_algo_config_file` matches model architecture |

### Export Errors

| Error | Likely Cause | Fix |
|-------|-------------|-----|
| `ModuleNotFoundError: gguf` | GGUF package missing | `pip install gguf>=0.10.0` |
| `ONNX export failed` | Model has unsupported ops | Try `--model_export hf_format` instead |
| `PermissionError` on output dir | No write permission | Check output directory permissions |

### Transformers Compatibility

Quark checks compatibility before quantization. Known issues:
- `seen_tokens` removed in transformers > 4.53.3
- `get_max_length` removed in transformers > 4.48.3
- `get_usable_length` removed in transformers > 4.53.3
- General constraint: `transformers < 5.3`

## Diagnostic Process

1. **Read the error** — get the exact error message, full stack trace, and the command that was run.
2. **Identify the layer** — is this an install problem, model loading, quantization runtime, or export issue?
3. **Check the environment** — Python version, PyTorch version, CUDA/ROCm version, Quark version, transformers version.
4. **Match against known patterns** — use the tables above.
5. **Propose the fix** — smallest change that resolves the issue without side effects.

## Diagnostic Commands

```bash
# Environment snapshot
python -c "
import sys; print('Python:', sys.version)
try:
    import torch; print('PyTorch:', torch.__version__, 'CUDA:', torch.version.cuda, 'HIP:', torch.version.hip)
except: print('PyTorch: not installed')
try:
    import quark; print('Quark:', quark.__version__)
except: print('Quark: not installed')
try:
    import transformers; print('Transformers:', transformers.__version__)
except: print('Transformers: not installed')
"

# GPU memory status
nvidia-smi --query-gpu=memory.used,memory.total --format=csv,noheader 2>/dev/null
rocm-smi --showmemuse 2>/dev/null

# Check Quark compatibility
python -c "
from quark.torch.utils.llm.compatibility import check_compatibility_before_quantization
print('Compatibility check available')
"

Rules

  • Always ask for the full error message and the command that was run. Partial errors lead to wrong diagnoses.
  • Do not guess the fix. Narrow down the root cause first, then propose a specific solution.
  • Confirm before mutating. If the fix involves reinstalling packages, changing environment, or modifying user files, present the plan and get confirmation.
  • Consider cascading effects. Upgrading transformers might fix one issue but break compatibility with Quark. Always check version constraints.

Error Summary

CUDA out of memory during quantization of Qwen/Qwen3-8B with FP8

Root Cause

Single GPU (24GB) insufficient for FP8 quantization with 512 calibration samples

Evidence

  • GPU memory: 22GB / 24GB used at crash point
  • Model size: ~16GB in FP16
  • Calibration data: 512 samples at seq_len=512

Fix

Reduce calibration data or use multi-GPU:

bash
# Option A: Reduce calibration samples
python quantize_quark.py ... --num_calib_data 64 --batch_size 1

# Option B: Use multi-GPU (if available)
python quantize_quark.py ... --multi_gpu auto

Prevention

For models > 7B parameters on GPUs with < 48GB VRAM, start with --num_calib_data 64 and increase if memory allows.

text

## Interaction Flow

1. **Gather evidence**: Get the error message, stack trace, command, and environment info.
2. **Classify**: Determine which layer failed (install / load / quantize / export).
3. **Diagnose**: Match against known patterns and run diagnostic commands if needed.
4. **Propose fix**: Present the smallest change that resolves the issue.
5. **Confirm**: Get user approval before any environment changes.
6. **Verify**: After the fix, re-run the failing step to confirm resolution.

## Recovery

- If the root cause is upstream (transformers API change, PyTorch bug), hand off to `quark-torch-doc-drift-check` or `quark-torch-skill-sync`.
- If the fix requires a PyTorch reinstall (wrong build, version mismatch, CPU instead of GPU), hand off to `quark-torch-install` with the specific requirement noted.
- If the fix requires a Quark package or dependency change, hand off to `quark-install` with the specific requirement noted.
- If the error is in a custom model or unsupported architecture, suggest registering a custom template via `LLMTemplate.register_template()`.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-debug of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch Debug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch Debug this skillamd/Quark181—~1.9kAutomated safety check: NotesMIT
Magpie Kernel Evaluatoramd/skills395—~2.3kAutomated safety check: PassMIT
The Art of Debuggingstas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Paddle Op DevPaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    395 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Fix Env

    evo-design/proto-tools

    Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…

    134 GitHub stars~2.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 9 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 9 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 9 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 9 days ago
    Auto-check passed

Questions about Quark Torch Debug

What does Quark Torch Debug do?

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. Quark Torch Debug is an agent skill from amd/Quark. Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

When should I use Quark Torch Debug?

Quark Torch Debug fits situations like: the user reports an error; invalid artifact; missing dependency; version mismatch.

How do I install Quark Torch Debug in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-debug -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-debug in amd/Quark) into .claude/skills/quark-torch-debug in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch Debug in Codex?

Run `npx skills add amd/Quark --skill quark-torch-debug -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-debug in amd/Quark) into .agents/skills/quark-torch-debug in your project. Codex loads it when a task matches its description.

Can I use Quark Torch Debug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-debug, .gemini/skills/quark-torch-debug, .github/skills/quark-torch-debug and .opencode/skills/quark-torch-debug in your project.

What does Quark Torch Debug need to run?

Going by SKILL.md and its folder, Quark Torch Debug needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Quark Torch Debug access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch Debug safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch Debug use?

Quark Torch Debug is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch Debug use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Torch Debug?

Skills that share tags, products or a category with Quark Torch Debug: Magpie Kernel Evaluator (amd/skills, 395 stars), The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and Ako4all (TongmingLAIC/AKO4ALL, 369 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch Debug?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.