Agent skill

Quark Torch LLM Ptq Workflow

by amd in amd/Quark

Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output.

MITAuto-check passedAI & LLM Engineering

Install Quark Torch LLM Ptq Workflow

skills CLI
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-llm-ptq-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow .claude/skills/quark-torch-llm-ptq-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-llm-ptq-workflow
GitHub stars
181
Token cost
~2.8k tokens
SKILL.md length
1,077 words
Files
2
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output.

  • Works in 4 steps: Model Intake → Quantization Plan → Manifest Generation → …
  • The user wants a complete PTQ pipeline: model inspection
  • SKILL.md covers Purpose, Inputs, Outputs: run_manifest.yaml and Interaction Flow, plus 9 more sections
  • Calls python3

What it does

Quark Torch LLM Ptq Workflow is an agent skill from amd/Quark. Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output. Use when the user wants a complete PTQ pipeline: model inspection, quantization planning, script generation, and optional execution. Stops at the quantized output. Trigger for "quantize my model", "run PTQ", "run model quantization", "full quantization pipeline", "quantize Llama/Qwen/Mistral with FP8/INT4", or any request that spans more than one PTQ step. When in doubt between routing to an atomic skill vs. the workflow, prefer this…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `example-fp8-qwen3-8b.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Mistral AI and Qwen. The licence is MIT.

When your agent uses it

  • The user wants a complete PTQ pipeline: model inspection
  • Quantization planning
  • Script generation
  • Optional execution

Example prompts

  • “quantize my model”
  • “run PTQ”
  • “run model quantization”
  • “/quark-torch-llm-ptq-workflow”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Model Intake
  2. Quantization Plan
  3. Manifest Generation
  4. Execute PTQ

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch LLM Ptq Workflow loads about 2.8k tokens when it runs. Until then it costs about 159 tokens; SKILL.md has 1,077 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,077 words, ~2,794 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-llm-ptq-workflow/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
quark-torch-llm-ptq-workflow
description
Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output. Use when the user wants a complete PTQ pipeline: model inspection, quantization planning, script generation, and optional execution. Stops at the quantized output. Trigger for "quantize my model", "run PTQ", "run model quantization", "full quantization pipeline", "quantize Llama/Qwen/Mistral with FP8/INT4", or any request that spans more than one PTQ step. When in doubt between routing to an atomic skill vs. the workflow, prefer this workflow if the user's request implies they want to go from model to quantized output.
layer
l2-workflows
primary_artifact
run_manifest.yaml
source_knowledge
examples/torch/language_modeling/llm_ptq/quantize_quark.py, examples/torch/language_modeling/llm_ptq/example_quark_torch_llm_ptq.rst…

quark-torch-llm-ptq-workflow

Purpose

Chain the full PTQ path — model intake → quantization planning → manifest generation → confirmed execution — while keeping the user informed at each checkpoint. This workflow orchestrates the atomic skills so the user does not have to manually chain them. It stops at the quantized output; for a run that also validates and evaluates the result, use the quark-torch-llm-ptq-eval recipe.

Inputs

  • Model path (HuggingFace ID or local directory)
  • User goal (target precision, hardware, accuracy)
  • Output directory for the quantized model
  • session_context.json for user goal and constraints
  • env_context.json for hardware facts
  • workspace_context.json for validated paths
  • pytorch_install_result.json and quark_install_result.json to confirm runtime is ready

Outputs: run_manifest.yaml

Records the executed command and config. Side artifacts: quantized model files written to the user's output directory. The manifest is built in Step 3 above.

Schema: run_manifest.schema.json

Interaction Flow

  1. Intake → quark-torch-model-intake → model_analysis.json
  2. Plan → quark-torch-quant-plan → quant_plan.json
  3. Confirm → run_manifest.yaml (stop for user approval)
  4. Execute → run the confirmed command, write quantized model

CRITICAL RULES

  1. NEVER run quantize_quark.py directly. Always go through the 4 steps below.
  2. NEVER skip a step. Even if the user provides all details upfront, execute each step in order.
  3. STOP at every checkpoint and wait for user confirmation before continuing.
  4. Show concrete output at each step — tables, JSON, commands — not just prose descriptions.
  5. NEVER modify Quark's own source code or examples. The Quark repo (quark/, examples/, tools/, docs/, tests/) is read-only from this workflow's perspective. See Upstream Quark Code is Read-Only below.

Upstream Quark Code is Read-Only

The Quark repository (quark/, examples/, tools/, docs/, tests/, pyproject.toml, requirements.txt) may be read freely but never modified. This includes "just to make the script accept my flag" patches to examples/torch/language_modeling/llm_ptq/quantize_quark.py — they break reproducibility against a clean Quark install.

When the shipped example does not cover the user's needs, write a fresh standalone script in the user's working directory (or /tmp/) that imports from quark:

python
# user_workspace/my_custom_ptq.py — NOT inside the Quark repo
from quark.torch import ModelQuantizer
from quark.torch.quantization.config.config import Config
# ... user-specific logic ...

Reference that script in run_manifest.yaml. The shipped quantize_quark.py stays untouched; the run remains reproducible against any Quark version. If the user actually needs an upstream Quark change, surface it as a contribution — do not silently patch their local checkout.

Required Artifact Flow

text
Step 1 (Intake)   ──► model_analysis.json
Step 2 (Plan)     ──► quant_plan.json
Step 3 (Manifest) ──► run_manifest.yaml      (contains the exact command)
Step 4 (Execute)  ──► quantized model output (only after user says yes)

Step 1: Model Intake

Goal: Produce model_analysis.json.

Actions
  1. Locate the Quark PTQ script (quantize_quark.py) under <Quark repo>/examples/torch/language_modeling/llm_ptq/ — find / -name quantize_quark.py -path "*/llm_ptq/*" 2>/dev/null | head -1 if the path is unknown. Record it for Step 3.

  2. Call quark-torch-model-intake with the model path. It handles config parsing (no weight load), supported-template matching, and risk identification (MoE, >70B, transformers version constraints) and emits model_analysis.json.

Output to Show User

Render the summary from model_analysis.json:

text
Model Analysis:
  Model path:       Qwen/Qwen3-8B
  Model type:       qwen3
  Hidden layers:    36
  Linear layers:    ~224 (quantization targets)
  MoE:              No
  Exclude defaults: [lm_head]
  Risks:            <list>
  Compatibility:    OK
>>> CHECKPOINT 1: Confirm model analysis is correct before continuing

Step 2: Quantization Plan

Goal: Build quant_plan.json from the model analysis and user's stated preferences.

Actions
  1. Determine the scheme. If the user stated a scheme (e.g., "FP8"), use it. Otherwise, recommend based on their priority:

    PriorityRecommended SchemeAlgorithm
    Best accuracyfp8 or ptpc_fp8smoothquant (optional)
    Smallest modelint4_wo_32awq or gptq
    CPU deploymentint8none
    AMD MI300X/MI355Xfp8 or amdfp4none
    NVIDIA H100fp8none
    GGUF exportuint4_wo_32awq
  2. Fill the decision table. Show ALL decisions with defaults:

    DecisionValueReason
    global_schemefp8User requested FP8
    kv_cache_schemefp8Recommended for FP8 inference
    exclude_layers["lm_head"]Standard — lm_head stays full precision
    layer_quant_config{}No per-pattern overrides (or e.g. {"*self_attn*": "fp8"} if the user asked to quantize attention with a non-global scheme)
    algorithmnullRTN baseline (fastest)
    calibration_datasetpilevalFast default
    num_calib_data128Standard default
    seq_len512Standard default
    evaluation_intentsmokeQuick PPL check after quantization
  3. Ask the user if they want to change anything.

>>> CHECKPOINT 2: User MUST confirm or adjust the plan before continuing

Wait for the user to say "ok", "confirm", "looks good", "continue", or similar. If they request changes (e.g., "use smoothquant", "increase calibration to 256"), update the table and re-present.


Step 3: Manifest Generation

Goal: Translate the confirmed plan into the exact quantize_quark.py command and produce run_manifest.yaml.

Show full SKILL.md (473 more words)Show less
Actions
  1. Build the command. Map plan fields to CLI arguments:

    Plan FieldCLI Argument
    model path--model_dir
    output dir--output_dir
    global_scheme--quant_scheme
    kv_cache_scheme--kv_cache_dtype (only if non-null)
    layer_quant_configone --layer_quant_scheme PATTERN SCHEME per dict entry (only if non-empty)
    exclude_layers--exclude_layers
    algorithm--quant_algo (only if non-null)
    num_calib_data--num_calib_data
    seq_len--seq_len
    export format--model_export (default: hf_format)
    data type--data_type auto
    device--device cuda
  2. Present the exact command:

    bash
    python3 <path_to_quantize_quark.py> \
      --model_dir <MODEL> \
      --output_dir <OUTPUT> \
      --quant_scheme <SCHEME> \
      --kv_cache_dtype <KV_SCHEME> \
      --num_calib_data <N> \
      --seq_len <LEN> \
      --model_export hf_format \
      --data_type auto \
      --device cuda

    Note on layer_quant_config patterns. Patterns are wildcard module-name matches against the model's named_modules(). Common LLaMA-style picks: '*self_attn*' (attention block), '*experts*' (MoE experts — covers all expert FFN submodules in one entry), 'lm_head' (output head). For models with different naming (e.g. attention, attn, self_attention, DeepSeek MLA), the pattern matches nothing silently and no override is applied — inspect named_modules() and adjust the pattern before running.

  3. Show expected output — config.json + model.safetensors (possibly sharded) + tokenizer files + quark_profile.yaml under <OUTPUT>/.

>>> CHECKPOINT 3: Show the command and ask "shall I run this?"

Do NOT proceed to execution unless the user explicitly confirms. Acceptable confirmations: "yes", "run it", "go", "execute", or similar.

If the user says "no" or wants changes, go back to the relevant step.


Step 4: Execute PTQ

Goal: Run the quantization command and report results.

Precondition

This step runs ONLY after the user explicitly confirms in Step 3.

Actions
  1. Create the output directory:

    bash
    mkdir -p <OUTPUT_DIR>
  2. Pick the accelerator and GPU. Read env_context.json for the backend. On ROCm, pin with HIP_VISIBLE_DEVICES (not CUDA_VISIBLE_DEVICES) — --device cuda still works on ROCm torch. On a shared host, check for a free GPU first and pin to it.

  3. Run the quantization command from Step 3. Monitor for:

    • CUDA/ROCm OOM → suggest reducing --num_calib_data or using --multi_gpu
    • Transformers version errors → report the version mismatch
    • Model loading failures → check trust_remote_code or model path
  4. After completion, verify outputs exist:

    bash
    ls -lh <OUTPUT_DIR>/
  5. Report results:

    text
    Quantization complete:
      Output:      <OUTPUT_DIR>/
      Model size:  X.X GB
      Format:      HuggingFace SafeTensors
      Perplexity:  X.XX (wikitext)
Error Recovery
  • If quantization fails, do NOT retry blindly. Report the error and suggest fixes based on quark-torch-debug patterns.
  • If OOM occurs, suggest: reduce --num_calib_data, reduce --batch_size 1, or use --multi_gpu auto.
  • If transformers version is wrong, show the required version range.

Complete Example

For an end-to-end walkthrough (FP8 quantization of Qwen3-8B), see example-fp8-qwen3-8b.md alongside this file.

Recovery

  • If any upstream artifact is missing, stop and name the missing producer skill. Do not improvise a partial artifact.
  • If the workflow hits a blocker (model cannot be loaded, scheme incompatible, execution fails), report it with diagnostic context rather than attempting ad-hoc fixes.
  • If the user wants to change a decision mid-workflow (e.g., switch from FP8 to INT4 after seeing model analysis), go back to the relevant step — do not restart from scratch.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow of amd/Quark.

  • SKILL.md
  • example-fp8-qwen3-8b.md

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch LLM Ptq Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch LLM Ptq Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch LLM Ptq Workflow this skillamd/Quark181—~2.8kAutomated safety check: PassMIT
Diffusion Perf Optvllm-project/vllm-omni7.1k—~7.5kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Check Modelguoqingbao/xinfer334—~3.8kAutomated safety check: PassMIT

Similar skills

  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    334 GitHub stars~3.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    334 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 12 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 12 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 12 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 12 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Quark Torch LLM Ptq Workflow

What does Quark Torch LLM Ptq Workflow do?

Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output. Quark Torch LLM Ptq Workflow is an agent skill from amd/Quark. Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output.

When should I use Quark Torch LLM Ptq Workflow?

Quark Torch LLM Ptq Workflow fits situations like: the user wants a complete PTQ pipeline: model inspection; quantization planning; script generation; optional execution.

How do I install Quark Torch LLM Ptq Workflow in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a claude-code`. Or copy the skill folder (.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow in amd/Quark) into .claude/skills/quark-torch-llm-ptq-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch LLM Ptq Workflow in Codex?

Run `npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a codex`. Or copy the skill folder (.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow in amd/Quark) into .agents/skills/quark-torch-llm-ptq-workflow in your project. Codex loads it when a task matches its description.

Can I use Quark Torch LLM Ptq Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-llm-ptq-workflow, .gemini/skills/quark-torch-llm-ptq-workflow, .github/skills/quark-torch-llm-ptq-workflow and .opencode/skills/quark-torch-llm-ptq-workflow in your project.

What does Quark Torch LLM Ptq Workflow need to run?

Going by SKILL.md and its folder, Quark Torch LLM Ptq Workflow needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Quark Torch LLM Ptq Workflow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch LLM Ptq Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch LLM Ptq Workflow use?

Quark Torch LLM Ptq Workflow is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch LLM Ptq Workflow use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Torch LLM Ptq Workflow?

Skills that share tags, products or a category with Quark Torch LLM Ptq Workflow: Diffusion Perf Opt (vllm-project/vllm-omni, 7.1k stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Add Model (guoqingbao/xinfer, 334 stars) and Resolve (alexziskind1/model-shelf, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch LLM Ptq Workflow?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.