SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent.
$ npx skills add amd/Quark --skill quark-torch-quant-plan -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/Quark quark-torch-quant-plan --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan .claude/skills/quark-torch-quant-plan && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "quark-torch-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan into .claude/skills/quark-torch-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-quant-plan", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-planType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/Quark --skill quark-torch-quant-plan -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/Quark quark-torch-quant-plan --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan .agents/skills/quark-torch-quant-plan && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "quark-torch-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan into .agents/skills/quark-torch-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-quant-plan", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-quant-plan -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/Quark quark-torch-quant-plan --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan .cursor/skills/quark-torch-quant-plan && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "quark-torch-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan into .cursor/skills/quark-torch-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-quant-plan", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/Quark.git --path .claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/Quark --skill quark-torch-quant-plan -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/Quark quark-torch-quant-plan --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan .gemini/skills/quark-torch-quant-plan && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "quark-torch-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan into .gemini/skills/quark-torch-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-quant-plan", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/Quark quark-torch-quant-planInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/Quark --skill quark-torch-quant-plan -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan .github/skills/quark-torch-quant-plan && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan into .github/skills/quark-torch-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-quant-plan", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-quant-plan -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/Quark quark-torch-quant-plan --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan .opencode/skills/quark-torch-quant-plan && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-quant-plan" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan into .opencode/skills/quark-torch-quant-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-quant-plan", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
quark-torch-quant-planBuild a Quark Torch LLM PTQ quantization plan from model analysis and user intent.
Quark Torch Quant Plan is an agent skill from amd/Quark. Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent. Use when the user needs quantization scheme recommendations, exclusion lists, algorithm selection, KV cache decisions, per-layer overrides, or a draft quantplan. Trigger for "quantize with FP8", "what scheme should I use", "plan PTQ", "INT4 quantization", "choose quantization config", "quantization plan", or when the user has a model analysis and needs to decide how to quantize it.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json and bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Quark Torch Quant Plan loads about 2.3k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 905 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 905 words, ~2,284 tokens.
.claude/skills/quark-torch-quant-plan/SKILL.md (or your agent's skills folder).Convert a model analysis plus the user's intent into a confirmed quant_plan.json. This skill makes the quantization decisions — which scheme, which algorithm, what to exclude — without generating scripts or executing PTQ. The plan is the contract between the user's intent and the execution step.
model_analysis.json from quark-torch-model-intakeenv_context.json for accelerator-aware scheme recommendationsRecords the chosen scheme, algorithm, layer overrides, calibration settings, and evaluation intent.
Schema: quant_plan.schema.json
{
"model": {
"model_type": "qwen3",
"analysis_ref": "./model_analysis.json"
},
"global_scheme": "fp8",
"kv_cache_scheme": "fp8",
"exclude_layers": ["lm_head"],
"layer_quant_config": {},
"algorithm": null,
"calibration": {
"dataset": "pileval",
"num_calib_data": 128,
"seq_len": 512
},
"evaluation_intent": "smoke",
"requires_confirmation": false
}| Scheme | Description | Use Case |
|---|---|---|
int4_wo_32 | INT4, group size 32 | Highest accuracy among INT4 |
int4_wo_64 | INT4, group size 64 | Good balance |
int4_wo_128 | INT4, group size 128 | Smaller overhead |
int4_wo_per_channel | INT4, per-channel | Least overhead |
uint4_wo_32/64/128/per_channel | Unsigned INT4 variants | GGUF export compatibility |
| Scheme | Description | Use Case |
|---|---|---|
int8 | INT8 per-tensor for both W and A | CPU deployment, good accuracy |
| Scheme | Description | Use Case |
|---|---|---|
fp8 | FP8 E4M3 per-tensor | Standard GPU quantization |
ptpc_fp8 | Per-Token-Per-Channel FP8 | Higher accuracy, dynamic activation quantization |
| Scheme | Description | Use Case |
|---|---|---|
mxfp4 | OCP MXFP4 | Aggressive compression |
mxfp6_e3m2 | OCP MXFP6 (E3M2) | Better range |
mxfp6_e2m3 | OCP MXFP6 (E2M3) | Better precision |
mxfp4_mxfp6_e2m3 | MXFP4 weights + MXFP6 activations | Mixed precision |
mxfp4_fp8 | MXFP4 weights + FP8 activations | Mixed precision |
| Scheme | Description | Use Case |
|---|---|---|
amdfp4 | amdfp4, group size 16 | AMD MI300X optimized |
amdfp4_g32 | amdfp4, group size 32 | AMD MI300X, less overhead |
| Scheme | Description | Use Case |
|---|---|---|
nvfp4 | NVFP4: FP4 group_size=16 with FP8 E4M3 scale | NVIDIA Blackwell/Hopper |
mx6 | MX6 format | Experimental |
bfp16 | Block Floating Point 16-bit | Experimental |
int4_wa_64 | INT4 weights + activations, group 64 | Research |
| Algorithm | Compatible Schemes | Description |
|---|---|---|
awq | INT4/UINT4 weight-only | Activation-aware weight quantization — finds optimal per-channel scaling |
gptq | INT4/UINT4 weight-only | Second-order weight optimization — often better than AWQ for small models |
smoothquant | INT8, FP8 | Migrates quantization difficulty from activations to weights |
autosmoothquant | INT8, FP8 | Automatic SmoothQuant with optimal alpha search |
rotation | Various | Rotation-based optimization to equalize weight distribution |
gptaq | INT4/UINT4 | GPTAQ variant combining GPTQ with activation quantization |
qronos | Various | Custom algorithm for time-series-aware quantization |
Algorithms can be combined: --quant_algo awq,smoothquant
fp8 is supported for KV cache (--kv_cache_dtype fp8)--min_kv_scale option (default 0.0) to prevent extreme scale values--kv_cache_post_rope quantizes KV cache after RoPE (inside cache) instead of at k_proj/v_proj outputs — can improve accuracy for some modelsHelp the user choose based on their priorities:
"I want the best accuracy" → fp8 or ptpc_fp8, optionally with smoothquant
"I want the smallest model" → int4_wo_32 with awq or gptq
"I need CPU deployment" → int8 (the only scheme that works well on CPU)
"I need GGUF format" → uint4_wo_32 with awq, export as GGUF
"I'm on AMD MI300X" → amdfp4 for best hardware utilization
"I'm on NVIDIA H100/Blackwell" → fp8 or nvfp4
"I want to experiment" → mxfp4 for aggressive compression research
ALWAYS present this table to the user and WAIT for confirmation before finalizing. Do not skip this step.
Fill in the "Value" column based on the user's request and model analysis, then show:
| Decision | Value | Reason |
|---|---|---|
global_scheme | (fill) | (why this scheme) |
kv_cache_scheme | (fill: fp8 or null) | (explain) |
exclude_layers | ["lm_head"] | Standard — lm_head stays full precision |
layer_quant_config | (fill: dict of pattern -> scheme, or {} if none) | (explain which patterns and why) |
algorithm | (fill: algorithm or null) | (explain) |
calibration_dataset | pileval | Fast default |
num_calib_data | 128 | Standard default |
seq_len | 512 | Standard default |
evaluation_intent | smoke | Quick PPL check post-quantization |
After showing the table, ask: "Confirm this plan? Any changes?"
Do NOT proceed until the user confirms.
layer_quant_configThe layer_quant_config plan field is a dict of pattern -> scheme pairs. It is the
single mechanism for any "quantize layer/module X with scheme Y" intent — including
attention modules, MoE experts, lm_head, etc. Each entry emits one
--layer_quant_scheme PATTERN SCHEME CLI argument.
"layer_quant_config": {
"*self_attn*": "fp8",
"lm_head": "int8",
"*experts*": "fp8"
}translates to:
--quant_scheme <global_scheme> \
--layer_quant_scheme '*self_attn*' fp8 \
--layer_quant_scheme lm_head int8 \
--layer_quant_scheme '*experts*' fp8When a user asks for attention-module quantization (e.g. "self_attn in fp8"), populate
this field with the appropriate pattern (commonly *self_attn* for LLaMA-style models;
adjust for models whose attention submodule has a different name). Do NOT introduce a
dedicated attention field — keep all per-pattern overrides in layer_quant_config.
model_analysis.json is missing, route back to quark-torch-model-intake.mxfp4 on a model where accuracy loss may be significant), keep the user's choice but record the risk in the plan.model_analysis.json available? If not, route to quark-torch-model-intake first.quant_plan.json.requires_confirmation: true and note what facts are missing.awq with fp8 — AWQ is designed for INT4), explain why it might not work well and suggest alternatives, but respect the user's choice if they insist.pileval with 128 samples — it is the fastest option and works for most models.© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan of amd/Quark.
Open the folder on GitHubat commit 313cb0b
Quark Torch Quant Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Quark Torch Quant Plan this skillamd/Quark | 181 | — | ~2.3k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| CI Fails Buildkiteguqiong96/Lvllm | 464 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
guqiong96/Lvllm
Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.
perminder-klair/subwave
Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…
amd/Quark
Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.
amd/Quark
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.
amd/Quark
Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/Quark
Install or verify the AMD Quark package and its dependencies.
amd/Quark
L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…
Categories
Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent. Quark Torch Quant Plan is an agent skill from amd/Quark. Build a Quark Torch LLM PTQ quantization plan from model analysis and user intent.
Quark Torch Quant Plan fits situations like: the user needs quantization scheme recommendations; exclusion lists; algorithm selection; KV cache decisions.
Run `npx skills add amd/Quark --skill quark-torch-quant-plan -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan in amd/Quark) into .claude/skills/quark-torch-quant-plan in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/Quark --skill quark-torch-quant-plan -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-quant-plan in amd/Quark) into .agents/skills/quark-torch-quant-plan in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-quant-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-quant-plan, .gemini/skills/quark-torch-quant-plan, .github/skills/quark-torch-quant-plan and .opencode/skills/quark-torch-quant-plan in your project.
SKILL.md names no scripts, command-line tools or credentials: Quark Torch Quant Plan is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Quark Torch Quant Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Quark Torch Quant Plan: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.
Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.