Diffusion Perf Opt
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output.
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/Quark quark-torch-llm-ptq-workflow --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow .claude/skills/quark-torch-llm-ptq-workflow && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "quark-torch-llm-ptq-workflow" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow into .claude/skills/quark-torch-llm-ptq-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-ptq-workflow", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflowType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/Quark quark-torch-llm-ptq-workflow --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow .agents/skills/quark-torch-llm-ptq-workflow && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "quark-torch-llm-ptq-workflow" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow into .agents/skills/quark-torch-llm-ptq-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-ptq-workflow", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/Quark quark-torch-llm-ptq-workflow --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow .cursor/skills/quark-torch-llm-ptq-workflow && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "quark-torch-llm-ptq-workflow" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow into .cursor/skills/quark-torch-llm-ptq-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-ptq-workflow", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/Quark.git --path .claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/Quark quark-torch-llm-ptq-workflow --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow .gemini/skills/quark-torch-llm-ptq-workflow && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "quark-torch-llm-ptq-workflow" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow into .gemini/skills/quark-torch-llm-ptq-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-ptq-workflow", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/Quark quark-torch-llm-ptq-workflowInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow .github/skills/quark-torch-llm-ptq-workflow && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-llm-ptq-workflow" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow into .github/skills/quark-torch-llm-ptq-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-ptq-workflow", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/Quark quark-torch-llm-ptq-workflow --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow .opencode/skills/quark-torch-llm-ptq-workflow && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-llm-ptq-workflow" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow into .opencode/skills/quark-torch-llm-ptq-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-ptq-workflow", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
quark-torch-llm-ptq-workflowTorch LLM PTQ workflow for AMD Quark — from model selection to quantized output.
Quark Torch LLM Ptq Workflow is an agent skill from amd/Quark. Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output. Use when the user wants a complete PTQ pipeline: model inspection, quantization planning, script generation, and optional execution. Stops at the quantized output. Trigger for "quantize my model", "run PTQ", "run model quantization", "full quantization pipeline", "quantize Llama/Qwen/Mistral with FP8/INT4", or any request that spans more than one PTQ step. When in doubt between routing to an atomic skill vs. the workflow, prefer this…
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `example-fp8-qwen3-8b.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Mistral AI and Qwen. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Quark Torch LLM Ptq Workflow loads about 2.8k tokens when it runs. Until then it costs about 159 tokens; SKILL.md has 1,077 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,077 words, ~2,794 tokens.
.claude/skills/quark-torch-llm-ptq-workflow/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Chain the full PTQ path — model intake → quantization planning → manifest generation → confirmed execution — while keeping the user informed at each checkpoint. This workflow orchestrates the atomic skills so the user does not have to manually chain them. It stops at the quantized output; for a run that also validates and evaluates the result, use the quark-torch-llm-ptq-eval recipe.
session_context.json for user goal and constraintsenv_context.json for hardware factsworkspace_context.json for validated pathspytorch_install_result.json and quark_install_result.json to confirm runtime is readyRecords the executed command and config. Side artifacts: quantized model files written to the user's output directory. The manifest is built in Step 3 above.
Schema: run_manifest.schema.json
quark-torch-model-intake → model_analysis.jsonquark-torch-quant-plan → quant_plan.jsonrun_manifest.yaml (stop for user approval)quantize_quark.py directly. Always go through the 4 steps below.quark/, examples/, tools/, docs/, tests/) is read-only from this workflow's perspective. See Upstream Quark Code is Read-Only below.The Quark repository (quark/, examples/, tools/, docs/, tests/, pyproject.toml, requirements.txt) may be read freely but never modified. This includes "just to make the script accept my flag" patches to examples/torch/language_modeling/llm_ptq/quantize_quark.py — they break reproducibility against a clean Quark install.
When the shipped example does not cover the user's needs, write a fresh standalone script in the user's working directory (or /tmp/) that imports from quark:
# user_workspace/my_custom_ptq.py — NOT inside the Quark repo
from quark.torch import ModelQuantizer
from quark.torch.quantization.config.config import Config
# ... user-specific logic ...Reference that script in run_manifest.yaml. The shipped quantize_quark.py stays untouched; the run remains reproducible against any Quark version. If the user actually needs an upstream Quark change, surface it as a contribution — do not silently patch their local checkout.
Step 1 (Intake) ──► model_analysis.json
Step 2 (Plan) ──► quant_plan.json
Step 3 (Manifest) ──► run_manifest.yaml (contains the exact command)
Step 4 (Execute) ──► quantized model output (only after user says yes)Goal: Produce model_analysis.json.
Locate the Quark PTQ script (quantize_quark.py) under <Quark repo>/examples/torch/language_modeling/llm_ptq/ — find / -name quantize_quark.py -path "*/llm_ptq/*" 2>/dev/null | head -1 if the path is unknown. Record it for Step 3.
Call quark-torch-model-intake with the model path. It handles config parsing (no weight load), supported-template matching, and risk identification (MoE, >70B, transformers version constraints) and emits model_analysis.json.
Render the summary from model_analysis.json:
Model Analysis:
Model path: Qwen/Qwen3-8B
Model type: qwen3
Hidden layers: 36
Linear layers: ~224 (quantization targets)
MoE: No
Exclude defaults: [lm_head]
Risks: <list>
Compatibility: OKGoal: Build quant_plan.json from the model analysis and user's stated preferences.
Determine the scheme. If the user stated a scheme (e.g., "FP8"), use it. Otherwise, recommend based on their priority:
| Priority | Recommended Scheme | Algorithm |
|---|---|---|
| Best accuracy | fp8 or ptpc_fp8 | smoothquant (optional) |
| Smallest model | int4_wo_32 | awq or gptq |
| CPU deployment | int8 | none |
| AMD MI300X/MI355X | fp8 or amdfp4 | none |
| NVIDIA H100 | fp8 | none |
| GGUF export | uint4_wo_32 | awq |
Fill the decision table. Show ALL decisions with defaults:
| Decision | Value | Reason |
|---|---|---|
global_scheme | fp8 | User requested FP8 |
kv_cache_scheme | fp8 | Recommended for FP8 inference |
exclude_layers | ["lm_head"] | Standard — lm_head stays full precision |
layer_quant_config | {} | No per-pattern overrides (or e.g. {"*self_attn*": "fp8"} if the user asked to quantize attention with a non-global scheme) |
algorithm | null | RTN baseline (fastest) |
calibration_dataset | pileval | Fast default |
num_calib_data | 128 | Standard default |
seq_len | 512 | Standard default |
evaluation_intent | smoke | Quick PPL check after quantization |
Ask the user if they want to change anything.
Wait for the user to say "ok", "confirm", "looks good", "continue", or similar. If they request changes (e.g., "use smoothquant", "increase calibration to 256"), update the table and re-present.
Goal: Translate the confirmed plan into the exact quantize_quark.py command and produce run_manifest.yaml.
Build the command. Map plan fields to CLI arguments:
| Plan Field | CLI Argument |
|---|---|
| model path | --model_dir |
| output dir | --output_dir |
global_scheme | --quant_scheme |
kv_cache_scheme | --kv_cache_dtype (only if non-null) |
layer_quant_config | one --layer_quant_scheme PATTERN SCHEME per dict entry (only if non-empty) |
exclude_layers | --exclude_layers |
algorithm | --quant_algo (only if non-null) |
num_calib_data | --num_calib_data |
seq_len | --seq_len |
| export format | --model_export (default: hf_format) |
| data type | --data_type auto |
| device | --device cuda |
Present the exact command:
python3 <path_to_quantize_quark.py> \
--model_dir <MODEL> \
--output_dir <OUTPUT> \
--quant_scheme <SCHEME> \
--kv_cache_dtype <KV_SCHEME> \
--num_calib_data <N> \
--seq_len <LEN> \
--model_export hf_format \
--data_type auto \
--device cudaNote on layer_quant_config patterns. Patterns are wildcard module-name matches against the model's named_modules(). Common LLaMA-style picks: '*self_attn*' (attention block), '*experts*' (MoE experts — covers all expert FFN submodules in one entry), 'lm_head' (output head). For models with different naming (e.g. attention, attn, self_attention, DeepSeek MLA), the pattern matches nothing silently and no override is applied — inspect named_modules() and adjust the pattern before running.
Show expected output — config.json + model.safetensors (possibly sharded) + tokenizer files + quark_profile.yaml under <OUTPUT>/.
Do NOT proceed to execution unless the user explicitly confirms. Acceptable confirmations: "yes", "run it", "go", "execute", or similar.
If the user says "no" or wants changes, go back to the relevant step.
Goal: Run the quantization command and report results.
This step runs ONLY after the user explicitly confirms in Step 3.
Create the output directory:
mkdir -p <OUTPUT_DIR>Pick the accelerator and GPU. Read env_context.json for the backend. On ROCm, pin with HIP_VISIBLE_DEVICES (not CUDA_VISIBLE_DEVICES) — --device cuda still works on ROCm torch. On a shared host, check for a free GPU first and pin to it.
Run the quantization command from Step 3. Monitor for:
--num_calib_data or using --multi_gputrust_remote_code or model pathAfter completion, verify outputs exist:
ls -lh <OUTPUT_DIR>/Report results:
Quantization complete:
Output: <OUTPUT_DIR>/
Model size: X.X GB
Format: HuggingFace SafeTensors
Perplexity: X.XX (wikitext)quark-torch-debug patterns.--num_calib_data, reduce --batch_size 1, or use --multi_gpu auto.For an end-to-end walkthrough (FP8 quantization of Qwen3-8B), see
example-fp8-qwen3-8b.md alongside this file.
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow of amd/Quark.
Open the folder on GitHubat commit 313cb0b
Quark Torch LLM Ptq Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Quark Torch LLM Ptq Workflow this skillamd/Quark | 181 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Diffusion Perf Optvllm-project/vllm-omni | 7.1k | — | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Add Modelguoqingbao/xinfer | 334 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Resolvealexziskind1/model-shelf | 130 | — | ~792 | Automated safety check: Pass | MIT | |
| Check Modelguoqingbao/xinfer | 334 | — | ~3.8k | Automated safety check: Pass | MIT |
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
guoqingbao/xinfer
Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.
guoqingbao/xinfer
Test LLM models served by xinfer for correctness, output quality, and performance.
amd/Quark
Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.
amd/Quark
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.
amd/Quark
Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/Quark
Install or verify the AMD Quark package and its dependencies.
amd/Quark
L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…
Works with
Categories
Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output. Quark Torch LLM Ptq Workflow is an agent skill from amd/Quark. Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output.
Quark Torch LLM Ptq Workflow fits situations like: the user wants a complete PTQ pipeline: model inspection; quantization planning; script generation; optional execution.
Run `npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a claude-code`. Or copy the skill folder (.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow in amd/Quark) into .claude/skills/quark-torch-llm-ptq-workflow in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a codex`. Or copy the skill folder (.claude/skills-impl/l2-workflows/torch/quark-torch-llm-ptq-workflow in amd/Quark) into .agents/skills/quark-torch-llm-ptq-workflow in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-llm-ptq-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-llm-ptq-workflow, .gemini/skills/quark-torch-llm-ptq-workflow, .github/skills/quark-torch-llm-ptq-workflow and .opencode/skills/quark-torch-llm-ptq-workflow in your project.
Going by SKILL.md and its folder, Quark Torch LLM Ptq Workflow needs the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Quark Torch LLM Ptq Workflow is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Quark Torch LLM Ptq Workflow: Diffusion Perf Opt (vllm-project/vllm-omni, 7.1k stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Add Model (guoqingbao/xinfer, 334 stars) and Resolve (alexziskind1/model-shelf, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.
Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.