Agent skill

Quark Torch LLM Ptq Eval

by amd in amd/Quark

L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate.

MITAuto-check passedAI & LLM Engineering

Install Quark Torch LLM Ptq Eval

skills CLI
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-llm-ptq-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l3-recipes/torch/quark-torch-llm-ptq-eval .claude/skills/quark-torch-llm-ptq-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-llm-ptq-eval
GitHub stars
182
Token cost
~2.6k tokens
SKILL.md length
1,209 words
Files
2
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate.

  • Works in 3 steps: Quantize (delegate to quark-torch-ptq) → Validate Output → Optional Accuracy Evaluation
  • The user wants to quantize and validate
  • SKILL.md covers Purpose, Inputs, Outputs and CRITICAL RULES, plus 6 more sections
  • Calls docker

What it does

Quark Torch LLM Ptq Eval is an agent skill from amd/Quark. L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate. Phase 1 delegates the full PTQ path (model intake → quantization planning → manifest generation → confirmed execution) to the quark-torch-ptq workflow; Phase 2 runs mandatory structural validation via quark-torch-result-validator; Phase 3 runs opt-in accuracy evaluation via quark-torch-llm-eval. Use when the user wants to "quantize and validate"…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `example-fp8-qwen3-8b.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving, Deep learning and End-to-end testing. It works with ONNX, PyTorch, Mistral AI and Qwen. The licence is MIT.

When your agent uses it

  • The user wants to quantize and validate
  • Quantize and evaluate
  • Run PTQ end to end with accuracy check
  • Full PTQ pipeline including validation and eval

Example prompts

  • “quantize and validate”
  • “quantize and evaluate”
  • “run PTQ end to end with accuracy check”
  • “/quark-torch-llm-ptq-eval”

Requirements

  • Python 3
  • Docker

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Quantize (delegate to quark-torch-ptq)
  2. Validate Output
  3. Optional Accuracy Evaluation

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch LLM Ptq Eval loads about 2.6k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 1,209 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~224
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,209 words, ~2,619 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-llm-ptq-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
quark-torch-llm-ptq-eval
description
L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate. Phase 1 delegates the full PTQ path (model intake → quantization planning → manifest generation → confirmed execution) to the `quark-torch-ptq` workflow; Phase 2 runs mandatory structural validation via `quark-torch-result-validator`; Phase 3 runs opt-in accuracy evaluation via `quark-torch-llm-eval`. Use when the user wants to "quantize and validate", "quantize and evaluate", "run PTQ end to end with accuracy check", "full PTQ pipeline including validation and eval", or "quantize Llama/Qwen/Mistral with FP8/INT4 and measure accuracy". For PTQ only (stop at the quantized output, no validation/eval), use `quark-torch-ptq`. Not for .onnx input models — use `quark-onnx-ptq` / `quark-onnx-autosearch-pro`.
layer
l3-recipes
backend
torch
primary_artifact
eval_report.md
source_knowledge
examples/torch/language_modeling/llm_ptq/quantize_quark.py, examples/torch/language_modeling/llm_ptq/example_quark_torch_llm_ptq.rst

quark-torch-llm-ptq-eval

Purpose

Run the complete PTQ lifecycle for a Torch LLM in one recipe: quantize → validate → evaluate. This recipe composes existing skills rather than re-implementing them — Phase 1 hands off to the quark-torch-ptq workflow (L2) for the PTQ steps, then Phases 2 and 3 chain the atomic validation and evaluation skills (L1). The end result is identical to running the PTQ workflow and then the validator and eval skills by hand; this recipe just drives the whole chain so the user does not have to invoke them separately.

Use this recipe when the user wants quantization plus a correctness and/or accuracy check in a single flow. If they only want the quantized model (no validation, no eval), route to quark-torch-ptq instead — do not run this recipe and then skip its later phases.

Inputs

  • Model path (HuggingFace ID or local directory)
  • User goal (target precision, hardware, accuracy)
  • Output directory for the quantized model
  • session_context.json for user goal and constraints
  • env_context.json for hardware facts
  • workspace_context.json for validated paths
  • pytorch_install_result.json and quark_install_result.json to confirm runtime is ready

Outputs

  • run_manifest.yaml — executed PTQ command and config (from Phase 1)
  • quantized model files in the user's output directory (from Phase 1)
  • validation_report.md — structural validation result (from Phase 2, mandatory)
  • eval_report.md — accuracy evaluation result (from Phase 3, optional)

CRITICAL RULES

  1. Compose, do not re-implement. Phase 1 is the quark-torch-ptq workflow run verbatim, with all its checkpoints. Do not inline or re-derive the intake/plan/manifest/execute steps here.
  2. STOP at every checkpoint surfaced by the delegated workflow and at this recipe's eval checkpoint.
  3. Validation (Phase 2) is mandatory and automatic. Evaluation (Phase 3) requires explicit user opt-in.
  4. NEVER modify Quark's own source code or examples. The Quark repo (quark/, examples/, tools/, docs/, tests/) is read-only. See the same rule in the quark-torch-ptq workflow.

Required Artifact Flow

text
Phase 1 (PTQ via quark-torch-ptq)
  Step 1 (Intake)   ──► model_analysis.json
  Step 2 (Plan)     ──► quant_plan.json
  Step 3 (Manifest) ──► run_manifest.yaml      (contains the exact command)
  Step 4 (Execute)  ──► quantized model output (only after user says yes)
Phase 2 (Validate)  ──► validation_report.md   (auto, mandatory)
Phase 3 (Eval)      ──► eval_report.md         (optional, requires user opt-in + ROCm for Tier 3)

Phase 1: Quantize (delegate to quark-torch-ptq)

Goal: Produce the quantized model plus model_analysis.json, quant_plan.json, and run_manifest.yaml.

Run the quark-torch-ptq workflow end-to-end (its Steps 1–4, including CHECKPOINT 1/2/3 and the Step 3 "shall I run this?" confirmation). Carry forward, for the later phases:

  • <MODEL_PATH> — source model, from intake
  • <OUTPUT_DIR> — quantized model directory, from the manifest
  • quant_plan.json — exclude rules and model paths, for the validator

Do not proceed to Phase 2 until Phase 1 has produced a quantized model in <OUTPUT_DIR>. If the PTQ workflow stops at a checkpoint or fails, stop here too — there is nothing to validate or evaluate.


Phase 2: Validate Output

Goal: Verify the quantized model is structurally sound. Runs automatically after Phase 1.

Actions

Call quark-torch-result-validator with:

  • source_model_dir: from intake (<MODEL_PATH>)
  • quantized_model_dir: from manifest (<OUTPUT_DIR>)
  • quant_config: from quant_plan.json (exclude rules for the MD5 check)

The validator runs four checks (fuzzy header → aux files → config.json → MD5) and emits validation_report.md.

Result handling
  • All four checks pass → one-line summary, proceed to Phase 3.
  • fuzzy layout or MD5 fails → STOP. These are correctness checks; report mismatches and suggest fixes via quark-torch-debug. Do not proceed.
  • only config.json or aux files fail → if the diff is benign metadata drift from a newer transformers re-serializing the model (renamed config keys, repackaged tokenizer files), mark it benign and continue. Otherwise treat as a real failure and STOP.
  • a real failure with an identifiable cause → diagnose the root cause: is it a Quark issue (export/quantization bug) or an environment issue (transformers / accelerate / torch version, model loading, paths)? Report the cause briefly. If it is fixable, ask the user whether they want help resolving it, and on approval apply the fix and re-run validation. Fix the environment or regenerate the artifact — never edit upstream Quark code (Critical Rule 4). If the cause is unclear or not safely fixable, STOP and hand off to quark-torch-debug.

No checkpoint — validation is mandatory and automatic.


Phase 3: Optional Accuracy Evaluation

Goal: Measure post-quantization accuracy and judge whether quantization hurt the model. Opt-in only.

Actions

Ask the user which evaluation tier to run (or skip):

TierMethodSpeedNotes
1. Quick — PPLperplexity (e.g. wikitext)fast, ~minutessanity check; runs on the quant device, no serving
2. Medium — lm_evalPython lm-eval library, direct generationslow — warn the user it can take a very long time (no high-throughput serving)task benchmark, e.g. gsm8k
3. Full — acceleratedhand off to quark-torch-llm-eval: vLLM in docker + lm_eval over the OpenAI endpointfast + comprehensiveROCm only
Skip"no"—finalize the recipe
Show full SKILL.md (512 more words)Show less
>>> CHECKPOINT: User must explicitly pick a tier (or skip) before eval runs
  1. Estimate the runtime before launching, based on tier, model size, benchmark size, and hardware. Confirm before a long run. Rough guide:

    • Tier 1 PPL: ~1-3 min
    • Tier 2 lm_eval (direct generation): can be hours for a full benchmark — warn explicitly
    • Tier 3 vLLM gsm8k: ~10-20 min on MI300X-class

    Tier 3 requires ROCm. If the host is not ROCm, only tiers 1-2 are available — note this and offer the deferred command:

    text
    Full accelerated eval requires AMD ROCm. To run later on a ROCm host:
      /quark-torch-llm-eval model_path=<OUTPUT_DIR> benchmark=gsm8k
  2. Run the chosen tier:

    • Tier 1 → run a PPL check on <OUTPUT_DIR>.

    • Tier 2 → run host lm-eval directly on the model (warn: slow).

    • Tier 3 → first check whether the current environment is already inside a suitable docker container (ROCm + vLLM available):

      • already inside a suitable container → run the eval there; do not create a new one.
      • not in a suitable container → ask the user whether to create a new docker container for the eval. Proceed only on approval; if declined, fall back to Tier 1/2 or skip.

      Then hand off to quark-torch-llm-eval with model_path=<OUTPUT_DIR>, benchmark=<choice>, backend=vLLM. It runs its own flow and produces eval_report.md.

    Known eval gotchas (apply while driving the eval skill / lm_eval):

    • vllm/vllm-openai-rocm image ENTRYPOINT is vllm → start a holder container with --entrypoint sleep, then docker exec the vllm serve.
    • Host lm_eval needs the [api] extra for OpenAI endpoints (a missing tenacity errors out immediately).
    • Reasoning models (e.g. Qwen3 <think>) need a large max_gen_toks on gsm8k; read the flexible-extract score.
  3. Report scores, then assess impact (impact assessment is opt-in).

    • Always: report the quantized scores. If a published/known reference exists (model card, leaderboard), compare against it for free — no extra compute.
    • For an exact apples-to-apples delta you must run the BF16 baseline under the same config, which roughly doubles eval time. Ask the user before running a baseline — do not run it automatically.
    • Verdict (using whatever reference is available):
      • small delta (within ~1-2%) → quantization is safe, model looks healthy
      • large drop (> a few %) → quantization hurt the model; flag it and suggest remedies (try smoothquant / awq, exclude more sensitive layers, or raise num_calib_data)
    • Record the verdict in eval_report.md. If no reference exists and the user declines the baseline, report scores only and state the delta is unknown.
  4. If skip → finalize, report all artifacts, recipe complete.


Complete Example

For an end-to-end walkthrough (FP8 quantization of Qwen3-8B, then validate + eval), see example-fp8-qwen3-8b.md alongside this file.

Recovery

  • If any upstream artifact is missing, stop and name the missing producer skill. Do not improvise a partial artifact.
  • If Phase 1 (quark-torch-ptq) stops or fails, do not start Phase 2/3 — there is nothing to validate or evaluate.
  • If the recipe hits a blocker (model cannot be loaded, scheme incompatible, execution fails), report it with diagnostic context rather than attempting ad-hoc fixes.
  • If the user wants to change a decision mid-recipe (e.g., switch from FP8 to INT4 after seeing model analysis), go back to the relevant step in Phase 1 — do not restart from scratch.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills-impl/l3-recipes/torch/quark-torch-llm-ptq-eval of amd/Quark.

  • SKILL.md
  • example-fp8-qwen3-8b.md

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch LLM Ptq Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch LLM Ptq Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch LLM Ptq Eval this skillamd/Quark182—~2.6kAutomated safety check: PassMIT
Model Builderqualcomm/qai-appbuilder247—~4.1kAutomated safety check: PassBSD-3-Clause
Atc Model Converterascend-ai-coding/awesome-ascend-skills174—~4.6kAutomated safety check: PassNone
Tao Port Huggingface ModelNVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.0
Model Inference Optimizemajiayu000/spellbook287—~1.1kAutomated safety check: PassMIT
Liger Kernel Devlinkedin/Liger-Kernel6.7k—~799Automated safety check: PassBSD-2-Clause

Similar skills

  • Model Builder

    qualcomm/qai-appbuilder

    QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

    247 GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Atc Model Converter

    ascend-ai-coding/awesome-ascend-skills

    Complete toolkit for Huawei Ascend NPU model conversion and end-to-end inference adaptation.

    174 GitHub stars~4.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

    3.6k GitHub stars~4.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Model Inference Optimize

    majiayu000/spellbook

    优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

    287 GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Liger Kernel Dev

    linkedin/Liger-Kernel

    Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.

    6.7k GitHub stars~799 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    182 GitHub stars~3.1k tokensUpdated 12 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    182 GitHub stars~3k tokensUpdated 12 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    182 GitHub stars~2.9k tokensUpdated 12 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    182 GitHub stars~1.4k tokensUpdated 12 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    182 GitHub stars~1.8k tokensUpdated 12 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    182 GitHub stars~3.4k tokensUpdated 12 days ago
    Auto-check passed

Questions about Quark Torch LLM Ptq Eval

What does Quark Torch LLM Ptq Eval do?

L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate. Quark Torch LLM Ptq Eval is an agent skill from amd/Quark. L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate.

When should I use Quark Torch LLM Ptq Eval?

Quark Torch LLM Ptq Eval fits situations like: the user wants to quantize and validate; quantize and evaluate; run PTQ end to end with accuracy check; full PTQ pipeline including validation and eval.

How do I install Quark Torch LLM Ptq Eval in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-llm-ptq-eval -a claude-code`. Or copy the skill folder (.claude/skills-impl/l3-recipes/torch/quark-torch-llm-ptq-eval in amd/Quark) into .claude/skills/quark-torch-llm-ptq-eval in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch LLM Ptq Eval in Codex?

Run `npx skills add amd/Quark --skill quark-torch-llm-ptq-eval -a codex`. Or copy the skill folder (.claude/skills-impl/l3-recipes/torch/quark-torch-llm-ptq-eval in amd/Quark) into .agents/skills/quark-torch-llm-ptq-eval in your project. Codex loads it when a task matches its description.

Can I use Quark Torch LLM Ptq Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-llm-ptq-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-llm-ptq-eval, .gemini/skills/quark-torch-llm-ptq-eval, .github/skills/quark-torch-llm-ptq-eval and .opencode/skills/quark-torch-llm-ptq-eval in your project.

What does Quark Torch LLM Ptq Eval need to run?

Going by SKILL.md and its folder, Quark Torch LLM Ptq Eval needs the command-line tools its instructions call (docker). Our summary lists: Python 3; Docker.

Does Quark Torch LLM Ptq Eval access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Quark Torch LLM Ptq Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch LLM Ptq Eval use?

Quark Torch LLM Ptq Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch LLM Ptq Eval use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Torch LLM Ptq Eval?

Skills that share tags, products or a category with Quark Torch LLM Ptq Eval: Model Builder (qualcomm/qai-appbuilder, 247 stars), Atc Model Converter (ascend-ai-coding/awesome-ascend-skills, 174 stars), Tao Port Huggingface Model (NVIDIA/skills, 3.6k stars) and Model Inference Optimize (majiayu000/spellbook, 287 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch LLM Ptq Eval?

amd (a GitHub organization) maintains it in amd/Quark, which has 182 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.