Agent skill

Quark Torch Quant Perf

by amd in amd/Quark

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

MITAuto-check passedAI & LLM Engineering

Install Quark Torch Quant Perf

skills CLI
$ npx skills add amd/Quark --skill quark-torch-quant-perf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-quant-perf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/quark-torch-quant-perf .claude/skills/quark-torch-quant-perf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-quant-perf
GitHub stars
181
Token cost
~3k tokens
SKILL.md length
1,319 words
Files
4 (incl. references)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

  • Works in 4 steps: Run quark-quant-perf --help for the… → Read CLI mapping for intent-to-option… → Read session lifecycle before resume, → …
  • The user asks for mixed-precision search
  • SKILL.md covers Purpose, Inputs, Outputs: session_report.md and Interaction Flow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quark Torch Quant Perf is an agent skill from amd/Quark. Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. Use whenever the user asks for mixed-precision search, Quark MXFP4/FP8/PTPC-FP8 quantization with accuracy and performance validation, vLLM throughput, TraceLens bottleneck analysis, GEAK kernel optimization, PerfOpt retries, workspace source discovery, or a natural-language request that must become a quark-quant-perf command. Not for ONNX models.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `evals/evals.json`, `references/cli-mapping.md` and `references/session-lifecycle.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Deep learning. It works with PyTorch, vLLM, ONNX and Transformers. The licence is MIT.

When your agent uses it

  • The user asks for mixed-precision search
  • Quark MXFP4/FP8/PTPC-FP8 quantization with accuracy and performance validation
  • VLLM throughput
  • TraceLens bottleneck analysis

Example prompts

  • “/quark-torch-quant-perf”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run quark-quant-perf --help for the installed CLI options and defaults.
  2. Read CLI mapping for intent-to-option mapping.
  3. Read session lifecycle before resume,
  4. Inspect the session's state.json and progress.json when resuming or

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch Quant Perf loads about 3k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 123 tokens; SKILL.md has 1,319 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,319 words, ~2,981 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-quant-perf/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
quark-torch-quant-perf
description
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. Use whenever the user asks for mixed-precision search, Quark MXFP4/FP8/PTPC-FP8 quantization with accuracy and performance validation, vLLM throughput, TraceLens bottleneck analysis, GEAK kernel optimization, PerfOpt retries, workspace source discovery, or a natural-language request that must become a quark-quant-perf command. Not for ONNX models.
layer
l2-workflows
backend
torch
primary_artifact
session_report.md
source_knowledge
quark/experimental/torch/quant_perf/README.md, quark/experimental/torch/quant_perf/cli.py, quark/experimental/torch/quant_perf/orchestration/orchestrator.py…

Quark Torch Quant-Perf

Purpose

Translate a user's PyTorch or HuggingFace quantization request into the current quark-quant-perf CLI, run or resume its managed Orchestrator, monitor the session to a terminal state, and report the generated artifacts. Preserve the user's accuracy, workload, source, and explicit performance requirements throughout.

Use the managed pipeline:

text
input and runtime validation
→ baseline health
→ quantize or mixed-precision search
→ real quantized accuracy gate
→ [measure or optimize] baseline and quantized throughput
→ [optimize, when the target is missed] conditional TraceLens / vendor tuning /
  GEAK → candidate validation and retention
→ FINAL reports

Performance modes skip inapplicable optional stages; they do not create separate FINAL stages. The Orchestrator enters FINAL after the selected stages complete or a terminal failure is recorded. The report subcommand may later regenerate terminal artifacts without rerunning GPU work.

Do not replace a stage with a custom evaluation, benchmark, or proxy metric.

Inputs

  • Model path or HuggingFace model ID.
  • Accuracy-gap requirement and optional performance intent.
  • GPU type, GPU index, TP, ISL, OSL, and concurrency when specified.
  • Fixed quantization intent, or a mixed-precision search space.
  • Framework and kernel source policy: explicit, auto, or readonly.
  • A discoverable Quark checkout containing quark-torch-ptq for fixed direct PTQ; QUARK_ROOT may select one explicitly.
  • Existing session directory when resuming.
  • Optional TraceLens architecture JSON and backend overrides.

Only --model is required by the CLI. Preserve CLI defaults for omitted options rather than inventing values.

Before constructing a command:

  1. Run quark-quant-perf --help for the installed CLI options and defaults.
  2. Read CLI mapping for intent-to-option mapping.
  3. Read session lifecycle before resume, monitoring, diagnosis, or final reporting.
  4. Inspect the session's state.json and progress.json when resuming or reporting.

Treat the installed CLI help as authoritative if a local reference differs.

Outputs: session_report.md

The managed FINAL stage produces:

text
<session>/reports/final.json
<session>/reports/final.md
<session>/session_breakdown.json
<session>/session_report.md

Treat session_breakdown.json as the complete structured fact source and session_report.md as the detailed human-readable result. A terminal failure may still have complete reports and useful quantized artifacts.

Interaction Flow

  1. Intake

    • Resolve local paths to absolute paths.
    • Preserve the requested model, accuracy gap, performance mode, target gain, workload, GPU, tensor parallelism, search modes, and repository policy.
    • Treat explicitly named quantized precision modes as a closed set, separate from any runtime-required native fallback. Do not infer compound modes from their components. Normalize aliases one-to-one; for example, use mxfp4_fp8 only when the user explicitly requests that mode, W4A8, or MXFP4 weights with FP8 activations.
    • If the user does not name precision modes, omit --layer-precision-candidates and preserve the GPU-specific automatic search space.
    • If the user does not request throughput or performance optimization, leave performance at the default off.
    • Inspect existing session state before creating a duplicate run.
  2. Route

    • Omit --quant-strategy for mixed-precision search.
    • Use --quant-strategy "<normalized intent>" for one fixed PTQ recipe.
    • For fixed PTQ, use the automatically discovered Quark checkout when available. Set QUARK_ROOT only when discovery fails or the user selects a specific checkout.
    • When writable framework or kernel repositories are supplied, use --workspace-source explicit with those repositories.
    • When repositories are not supplied, use --workspace-source auto so runtime repair and PerfOpt can discover or materialize writable sources.
    • Use --workspace-source readonly only when the user explicitly requests evaluation or inspection without source modification.
    • Treat fresh or no-history execution as an experience-store policy, not a source policy. It never implies readonly; keep auto unless the user separately prohibits source modification.
    • Before launch, if the command contains --workspace-source readonly, identify the user's explicit no-source-modification request. If there is none, remove the option and use the default auto.
    • Apply the source policy to every run. It governs runtime activation and eligible repair work even when performance mode is off or measure; PerfOpt source modification remains specific to optimize.
  3. Plan

    • Maintain the native task plan for every multi-stage run. Use update_plan in Codex and TaskCreate / TaskUpdate in Claude Code.
    • Before launch, show every applicable top-level task derived from the resolved run intent. Always include validation, baseline health, quantization or search, the real accuracy gate, and FINAL reporting. Add throughput for measure or optimize, and add the PerfOpt decision plus conditional bottleneck, optimization, and retention tasks for optimize. Do not use one universal task list for every performance mode.
    • Keep exactly one plan step in progress.
    • Mark an applicable conditional task as completed with the skip reason when the Orchestrator does not need it. Add candidate, retry, repair, and kernel subtasks dynamically when command output, state.json, progress.json, or an artifact shows that they exist.
    • On resume or context compaction, reconstruct the complete run-specific plan from the persisted run specification and current session evidence rather than conversation memory.
  4. Execute

    • Launch through quark-quant-perf; do not call internal stage functions as a replacement pipeline.
    • Keep long runs in a managed terminal or tool session that can be polled; do not detach them with nohup or shell &.
    • Keep the exact command, session directory, and process handle visible.
    • Do not pause, signal, or terminate an active process without explicit user approval unless immediate system safety or data loss is at risk.
  5. Monitor and summarize

    • Poll process output while the run is active. When no new output arrives, inspect state.json and progress.json about every 30 seconds. Do not redraw an unchanged plan, but always refresh it before a status response or wait.
    • Use progress.json for the current stage and stage_detail, and use state.json for completed, failed, resumed, and terminal facts.
    • During mixed-precision search, show the persisted total_configs_evaluated / total_configs_available counts. Also account for candidate_cursor, candidate_queue, partial_timeout, and termination_reason when describing export fallback or a salvaged search.
    • Keep all applicable top-level tasks visible throughout the run. Enrich the active task and add retry, repair, candidate, and kernel subtasks when their corresponding attempts, measurements, journeys, or artifacts appear.
    • Use terminal output to enrich the current task, not as the sole evidence for completion.
    • Distinguish search accuracy from the authoritative real accuracy gate.
    • Distinguish local kernel speedup from final retained end-to-end gain.
    • Continue through FINAL even when accuracy or performance targets fail.
    • Report the terminal status, selected quantization, accuracy, throughput, retained patches, rejected attempts, and artifact paths.
Show full SKILL.md (364 more words)Show less

Command Rules

Use current canonical option names:

text
--max-search-candidates
--search-timeout
--layer-precision-candidates
--kv-cache-precision-candidates
--search-gsm8k-num-samples
--geak-direction-budget

Do not emit removed historical names:

text
--max-rounds
--layer-mode
--kv-cache-mode
--geak-budget

Convert percentages to ratios:

  • 3% maximum accuracy drop → --accuracy-gap 0.03
  • 35% target speedup → --target-gain 1.35
  • 2x throughput target → --target-gain 2.0

Performance intent:

  • No performance request → omit both options; the effective mode is off.
  • Throughput measurement only → --performance-mode measure.
  • Optimization target → --target-gain MULTIPLIER; this implies --performance-mode optimize.
  • Explicit optimize without a target uses the CLI default target of 1.2.

Precision and backend intent:

  • No requested layer modes → omit --layer-precision-candidates and use the GPU-specific defaults.
  • Explicit layer modes → pass only the named quantized modes; native remains an implicit fallback.
  • No requested backend → omit the MXFP4/W4A8 backend flags and preserve their CLI or environment defaults.
  • Explicit backend → pass only the corresponding backend option; backend selection does not add a precision mode to the search space.

Use --tracelens-gpu-arch-json when the installed TraceLens package lacks the requested GPU architecture data. Do not copy architecture data into site-packages.

Use --retry-accuracy-gate when the real accuracy gate must be reopened. Use --retry-perfopt only when reusable accuracy and quant-only throughput evidence still match the current runtime fingerprint.

Monitoring Rules

  • A status request, parameter correction, turn interruption, or new message is not permission to stop an active task.
  • Do not infer current state from old warnings in logs.
  • Use state.json, terminal command output, and FINAL reports for conclusions.
  • A local GEAK result is not a final performance result.
  • A patch is retained only after the managed correctness, accuracy, and end-to-end performance gates accept it.
  • Preserve unrelated user changes and dirty worktrees.

Recovery

  • Reuse the same session when the user asks to continue or resume.
  • Diagnose the failing boundary before changing code or configuration.
  • Use --recheck-baseline only to bypass an exact cached baseline failure.
  • Use --retry-accuracy-gate after a failed accuracy stage or after a runtime change invalidates the saved accuracy fingerprint.
  • Use --retry-perfopt after perf_failed only when upstream evidence remains reusable.
  • Do not silently change TP, model, ISL, OSL, concurrency, quantization modes, KV-cache modes, or benchmark implementation during recovery.
  • If source paths or runtime identity change, expect fingerprints to invalidate cached measurements and rerun the required managed gates.
  • Preserve failure logs and candidate evidence even when the candidate source change is reverted.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/quark-torch-quant-perf of amd/Quark.

  • SKILL.md
  • evals/evals.json
  • references/cli-mapping.md
  • references/session-lifecycle.md

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch Quant Perf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch Quant Perf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch Quant Perf this skillamd/Quark181—~3kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Hqq QuantizationOrchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT
Model Builderqualcomm/qai-appbuilder246—~4.1kAutomated safety check: PassBSD-3-Clause
Magpie Kernel Evaluatoramd/skills398—~2.3kAutomated safety check: PassMIT
Spark Environment Setupwshobson/agents40k—~2kAutomated safety check: PassMIT

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Hqq Quantization

    Orchestra-Research/AI-Research-SKILLs

    Half-Quadratic Quantization for LLMs without calibration data.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Builder

    qualcomm/qai-appbuilder

    QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

    246 GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    398 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

    40k GitHub stars~2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Pytorch CI Triage

    pytorch/test-infra

    Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra.

    113 GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 10 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 10 days ago
    Auto-check passed
  • Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts.

    181 GitHub stars~4.8k tokensUpdated 10 days ago
    Auto-check passed

Questions about Quark Torch Quant Perf

What does Quark Torch Quant Perf do?

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. Quark Torch Quant Perf is an agent skill from amd/Quark. Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

When should I use Quark Torch Quant Perf?

Quark Torch Quant Perf fits situations like: the user asks for mixed-precision search; quark MXFP4/FP8/PTPC-FP8 quantization with accuracy and performance validation; VLLM throughput; traceLens bottleneck analysis.

How do I install Quark Torch Quant Perf in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-quant-perf -a claude-code`. Or copy the skill folder (skills/quark-torch-quant-perf in amd/Quark) into .claude/skills/quark-torch-quant-perf in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch Quant Perf in Codex?

Run `npx skills add amd/Quark --skill quark-torch-quant-perf -a codex`. Or copy the skill folder (skills/quark-torch-quant-perf in amd/Quark) into .agents/skills/quark-torch-quant-perf in your project. Codex loads it when a task matches its description.

Can I use Quark Torch Quant Perf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-quant-perf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-quant-perf, .gemini/skills/quark-torch-quant-perf, .github/skills/quark-torch-quant-perf and .opencode/skills/quark-torch-quant-perf in your project.

What does Quark Torch Quant Perf need to run?

SKILL.md names no scripts, command-line tools or credentials: Quark Torch Quant Perf is instructions for the agent only.

Does Quark Torch Quant Perf access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch Quant Perf safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch Quant Perf use?

Quark Torch Quant Perf is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch Quant Perf use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Quark Torch Quant Perf?

Skills that share tags, products or a category with Quark Torch Quant Perf: Graphsignal (graphsignal/graphsignal, 257 stars), Hqq Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars), Model Builder (qualcomm/qai-appbuilder, 246 stars) and Magpie Kernel Evaluator (amd/skills, 398 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch Quant Perf?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.