Agent skill

Quark Torch Ptq

by amd in amd/Quark

Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request…

MITAuto-check passedAI & LLM Engineering

Install Quark Torch Ptq

skills CLI
$ npx skills add amd/Quark --skill quark-torch-ptq -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-ptq --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/quark-torch-ptq .claude/skills/quark-torch-ptq && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-ptq
GitHub stars
181
Token cost
~2.3k tokens
SKILL.md length
1,074 words
Files
11 (incl. references)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request…

  • Works in 4 steps: Model intake → Quantization plan → Manifest and execution confirmation → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Purpose, Prerequisites, Inputs and Outputs, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quark Torch Ptq is an agent skill from amd/Quark. Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request execution approval, and produce a quantized model. Applies to Llama, Qwen, Mistral, and similar transformer LLM requests involving FP8, INT4, or another Quark scheme. Does not handle .onnx model inputs or ONNX PTQ.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `evals/evals.json`, `references/contracts/model_analysis.schema.json` and `references/contracts/quant_plan.schema.json`).

It sits in AI & LLM Engineering, covering LLM inference and serving, Deep learning and Model hubs and datasets. It works with PyTorch, ONNX, Hugging Face and Mistral AI. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Deep learning
  • Tasks that involve Model hubs and datasets

Example prompts

  • “/quark-torch-ptq”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Model intake
  2. Quantization plan
  3. Manifest and execution confirmation
  4. Execute and verify

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch Ptq loads about 2.3k tokens when it runs, and up to ~9k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 1,074 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,074 words, ~2,337 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-ptq/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
quark-torch-ptq
description
Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request execution approval, and produce a quantized model. Applies to Llama, Qwen, Mistral, and similar transformer LLM requests involving FP8, INT4, or another Quark scheme. Does not handle .onnx model inputs or ONNX PTQ.

Quark Torch PTQ

Purpose

Take a PyTorch / Hugging Face LLM from model identification through confirmed AMD Quark PTQ. Perform intake, planning, manifest generation, execution, and output verification as one self-contained workflow.

This skill stops at the quantized model. It does not accept .onnx model input, train or fine-tune a model, or modify Quark package/source files.

Prerequisites

  • Python 3.11 to 3.13, accelerator-matched PyTorch 2.2 or newer, amd-quark[cli], and datasets.
  • The required ROCm version and GPU architecture depend on the PyTorch build and selected quantization scheme. Record the actual runtime, GPU architecture (gfx... from gcnArchName on AMD), host kernel, and driver, and verify support for the confirmed plan.
  • No container image is required. If running in a container, record its image name or digest when available and verify GPU device access.
  • Inspect and preserve HIP_VISIBLE_DEVICES, CUDA_VISIBLE_DEVICES, HSA_OVERRIDE_GFX_VERSION, PYTORCH_ROCM_ARCH, and PYTORCH_HIP_ALLOC_CONF. Include any required changes to device visibility, architecture, or memory allocation in the confirmed plan.

Inputs

  • Model source: a Hugging Face repository ID or local model directory.
  • Output directory.
  • Quantization intent: requested precision/scheme, target hardware, accuracy priority, and optional calibration settings.
  • Optional environment facts: Python, PyTorch, accelerator, available memory, installed amd-quark, and Transformers versions.

Do not require pre-existing workflow artifacts. Create all artifacts in the user's working directory by following this skill's local references:

Outputs

Produce these three artifacts before or during execution:

  1. model_analysis.json, validated against references/contracts/model_analysis.schema.json.
  2. quant_plan.json, validated against references/contracts/quant_plan.schema.json.
  3. run_manifest.yaml, validated against references/contracts/run_manifest.schema.json.

The quantized model and its configuration/tokenizer files are written under the confirmed output directory. Record actual files and the final status in the manifest.

Interaction Flow

Always complete the following four steps in order. Show concrete facts, artifacts, and commands. Stop at every checkpoint and wait for the user.

Step 1 — Model intake
  1. Confirm whether the model source is local or remote. For a local source, resolve it to an absolute path and verify that the directory and config.json exist. For a remote source, preserve the repository ID.
  2. Follow references/model-intake.md. Read configuration only; do not load model weights during intake.
  3. Determine model_type, architecture/loading hints, hidden-layer facts, multimodal or MoE signals, default exclusions, compatibility risks, and a defensible estimate of quantizable linear layers.
  4. Write schema-valid model_analysis.json and show its summary.
Checkpoint 1 — Confirm the model analysis

Ask the user to confirm or correct the model analysis.

Do not plan quantization until the user confirms.

Step 2 — Quantization plan
  1. Follow references/quant-plan.md and use the confirmed analysis plus the user's priorities.
  2. Select and explain the global scheme, optional KV-cache scheme, per-pattern overrides, exclusions, algorithms, and calibration data.
  3. Treat scheme, algorithm, and model-template support as version-dependent. Use the current installed Quark API or quark-cli torch-llm-ptq --help when a choice needs verification; do not rely on historical list sizes.
  4. Show a decision table, write schema-valid quant_plan.json, and set requires_confirmation: true until approved.
Checkpoint 2 — Confirm the quantization plan

Ask the user to confirm or adjust the complete plan.

After confirmation, update requires_confirmation to false. If the user changes a decision, rewrite and revalidate the plan before continuing.

Step 3 — Manifest and execution confirmation

Use the public quark-cli torch-llm-ptq command installed by amd-quark[cli]. Do not import its implementation module directly or locate, copy, generate, or patch another PTQ runner.

Build an argument-array-safe command equivalent to:

bash
quark-cli torch-llm-ptq \
  --model_dir "<MODEL_OR_ABSOLUTE_LOCAL_PATH>" \
  --output_dir "<ABSOLUTE_OUTPUT_PATH>" \
  --quant_scheme "<SCHEME>" \
  --num_calib_data "<N>" \
  --seq_len "<LENGTH>" \
  --device cuda \
  --no_trust_remote_code

Add only confirmed options:

  • --dataset <NAME> and --batch_size <N> when the plan changed them from the CLI defaults.
  • --kv_cache_dtype <SCHEME> for confirmed KV-cache quantization.
  • One --layer_quant_scheme <PATTERN> <SCHEME> per override.
  • --quant_algo <comma-separated-list> when required.
  • --exclude_layers <patterns...> only when overriding template defaults.
  • --multi_device when its constraints are understood.
  • --no_trust_remote_code unless the user explicitly accepts executing remote model code. The CLI trusts remote code when this flag is omitted.
  • --skip_evaluation when the requested scope ends strictly at model output.
  • --evaluation_dataset <NAME> when evaluation is requested with a non-default CLI-supported dataset.

Resolve local model and output directories to absolute paths, then quote all user-controlled paths and values. On ROCm, --device cuda is still the PyTorch device spelling; use HIP_VISIBLE_DEVICES to pin a GPU when needed. Resolve quark-cli from the same Python environment that provides amd-quark.

Write run_manifest.yaml with:

  • workflow: quark-torch-ptq
  • input and output paths;
  • all four checkpoint reasons;
  • an analyze step producing model_analysis.json;
  • a plan step producing quant_plan.json;
  • a generate step producing run_manifest.yaml;
  • a run step containing the exact command and expected output directory.

Confirm the environment before showing the command, using references/environment.md. A missing package or an accelerator-mismatched PyTorch build should surface here, not after the execution gate.

Show the full command, destination, estimated resource needs, remote-code choice, and expected outputs.

Show full SKILL.md (318 more words)Show less
Checkpoint 3 — Approve execution

Ask: “Shall I run this exact command?”

This is the execution gate. Do not create the output directory, download model weights, change the environment, or run PTQ without explicit approval such as “yes”, “run it”, or “execute”. A prior plan confirmation is not execution approval.

Step 4 — Execute and verify

Only after Checkpoint 3 approval:

  1. Create the output directory if needed.
  2. Run the exact confirmed quark-cli command.
  3. Monitor output. On failure, stop; collect the command, exit status, full error, versions, and resource state, then use references/troubleshooting.md. Never retry blindly.
  4. On success, inspect the output directory and report actual model shards, configuration/tokenizer files, size, format, and any requested metrics.
  5. Update run_manifest.yaml with the observed status and outputs without changing the recorded command.
Checkpoint 4 — Accept the verified result

Present the verified result and ask the user to accept it or request a bounded follow-up.

Do not claim success from exit status alone. If expected files are absent, report a partial/failed result and preserve diagnostics.

Recovery

  • Missing or invalid artifact: regenerate it from the corresponding local reference and validate it against the local schema; do not invent fields around a validation error.
  • Intake uncertainty: mark analysis_status as partial, record a risk, and ask for the missing fact. Do not load weights merely to fill metadata.
  • Unsupported model type or scheme: report the installed Quark evidence and ask whether to change the plan. Do not patch the installed CLI.
  • OOM or device failure: retain the failed manifest and propose the smallest plan change, such as fewer calibration samples, a shorter sequence length, or multi-device execution. Return to Checkpoint 2.
  • Dependency or compatibility failure: show exact installed and required versions. Get confirmation before package changes, then return to Checkpoint 3 with a newly recorded command.
  • User changes intent: return to the earliest affected checkpoint and preserve still-valid artifacts. Never bypass the execution confirmation.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/quark-torch-ptq of amd/Quark.

  • SKILL.md
  • evals/evals.json
  • references/contracts/model_analysis.schema.json
  • references/contracts/quant_plan.schema.json
  • references/contracts/run_manifest.schema.json
  • references/environment.md
  • references/example-fp8-qwen3-8b.md
  • references/model-intake.md
  • references/quant-plan.md
  • references/troubleshooting.md
  • skill-card.md

Open the folder on GitHubat commit 313cb0b

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in amd/Quark, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Quark Torch Ptq next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch Ptq compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch Ptq this skillamd/Quark181—~2.3kAutomated safety check: PassMIT
Tao Port Huggingface ModelNVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0
Model Builderqualcomm/qai-appbuilder247—~4.1kAutomated safety check: PassBSD-3-Clause
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT

Similar skills

  • Official

    Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

    3.5k GitHub stars~4.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Model Builder

    qualcomm/qai-appbuilder

    QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

    247 GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    334 GitHub stars~3.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 11 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 11 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 11 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 11 days ago
    Auto-check passed

Questions about Quark Torch Ptq

What does Quark Torch Ptq do?

Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request…. Quark Torch Ptq is an agent skill from amd/Quark. Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request execution approval, and produce a quantized model.

When should I use Quark Torch Ptq?

Quark Torch Ptq fits situations like: tasks that involve LLM inference and serving; tasks that involve Deep learning; tasks that involve Model hubs and datasets.

How do I install Quark Torch Ptq in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-ptq -a claude-code`. Or copy the skill folder (skills/quark-torch-ptq in amd/Quark) into .claude/skills/quark-torch-ptq in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch Ptq in Codex?

Run `npx skills add amd/Quark --skill quark-torch-ptq -a codex`. Or copy the skill folder (skills/quark-torch-ptq in amd/Quark) into .agents/skills/quark-torch-ptq in your project. Codex loads it when a task matches its description.

Can I use Quark Torch Ptq in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-ptq -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-ptq, .gemini/skills/quark-torch-ptq, .github/skills/quark-torch-ptq and .opencode/skills/quark-torch-ptq in your project.

What does Quark Torch Ptq need to run?

SKILL.md names no scripts, command-line tools or credentials: Quark Torch Ptq is instructions for the agent only. Our summary lists: Python 3.

Does Quark Torch Ptq access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch Ptq safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch Ptq use?

Quark Torch Ptq is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch Ptq use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.6k tokens, read only when the agent opens those files.

What are the alternatives to Quark Torch Ptq?

Skills that share tags, products or a category with Quark Torch Ptq: Tao Port Huggingface Model (NVIDIA/skills, 3.5k stars), Model Builder (qualcomm/qai-appbuilder, 247 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Add Model (guoqingbao/xinfer, 334 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch Ptq?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.