Agent skill

Quark Torch Export

by amd in amd/Quark

Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run.

MITAuto-check passedAI & LLM Engineering

Install Quark Torch Export

skills CLI
$ npx skills add amd/Quark --skill quark-torch-export -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-export --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-export .claude/skills/quark-torch-export && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-export
GitHub stars
181
Token cost
~1.5k tokens
SKILL.md length
458 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run.

  • Works in 3 steps: HuggingFace SafeTensors (hf_format) —… → ONNX (onnx) → GGUF (gguf)
  • The user wants to export a quantized model to HuggingFace SafeTensors
  • SKILL.md covers Purpose, Inputs, Outputs: run_manifest.yaml and Supported Export Formats, plus 6 more sections
  • Calls python

What it does

Quark Torch Export is an agent skill from amd/Quark. Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run. Use when the user wants to export a quantized model to HuggingFace SafeTensors, ONNX, or GGUF format, package for deployment, or set up post-quantization evaluation. Trigger for "export model", "save quantized model", "convert to GGUF", "export quantized model", "export to HuggingFace format", or when the user has a completed or planned PTQ run and needs deployment outputs.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving, Model hubs and datasets and Deployment. It works with llama.cpp, ONNX and Hugging Face. The licence is MIT.

When your agent uses it

  • The user wants to export a quantized model to HuggingFace SafeTensors
  • Package for deployment
  • Set up post-quantization evaluation
  • Save quantized model

Example prompts

  • “export model”
  • “save quantized model”
  • “convert to GGUF”
  • “/quark-torch-export”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. HuggingFace SafeTensors (hf_format) — Default
  2. ONNX (onnx)
  3. GGUF (gguf)

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch Export loads about 1.5k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 458 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 458 words, ~1,464 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-export/SKILL.md (or your agent's skills folder).
name
quark-torch-export
description
Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run. Use when the user wants to export a quantized model to HuggingFace SafeTensors, ONNX, or GGUF format, package for deployment, or set up post-quantization evaluation. Trigger for "export model", "save quantized model", "convert to GGUF", "export quantized model", "export to HuggingFace format", or when the user has a completed or planned PTQ run and needs deployment outputs.
layer
l1-atomic
primary_artifact
run_manifest.yaml
source_knowledge
examples/torch/language_modeling/llm_ptq/quantize_quark.py, quark/torch/export/api.py, examples/contrib/llm_eval/llm_eval.py

quark-torch-export

Purpose

Translate a confirmed quantization plan into export expectations and downstream evaluation requirements. Export is a post-quantization step — the model must be quantized first (or have a plan to be quantized) before export decisions make sense. This skill ensures the right export format is chosen and evaluation is properly configured.

Inputs

  • quant_plan.json from quark-torch-quant-plan
  • workspace_context.json for the output directory
  • run_manifest.yaml (optional, existing manifest to extend)

Outputs: run_manifest.yaml

Export does not own this artifact — it updates the workflow's manifest with export and evaluation config.

Schema: run_manifest.schema.json

(Export updates the workflow's run_manifest.yaml with export and evaluation fields rather than producing a new artifact.)

yaml
export:
  formats:
    - hf_format
  output_dir: ./output/qwen3-8b-fp8
  weight_format: real_quantized
  custom_mode: quark

evaluation:
  skip: false
  metrics:
    - ppl
  dataset: wikitext
  tasks: null
  batch_size: auto

Supported Export Formats

1. HuggingFace SafeTensors (hf_format) — Default
  • Produces: config.json + *.safetensors files with quantization_config metadata
  • Compatible with: HuggingFace transformers loading, vLLM, TGI
  • CLI flag: --model_export hf_format
  • Weight format options:
    • real_quantized (default) — compressed, actual quantized weights
    • fake_quantized — full-precision weights with quantization metadata only
2. ONNX (onnx)
  • Produces: quark_model.onnx with optimization passes applied
  • Compatible with: ONNX Runtime, TensorRT (with conversion)
  • CLI flag: --model_export onnx
  • Supports INT4/UINT4 conversion pass automatically
3. GGUF (gguf)
  • Produces: GGUF format file for llama.cpp and compatible inference engines
  • Requires: gguf>=0.10.0 package and tokenizer path
  • CLI flag: --model_export gguf
  • Best with: uint4_wo_32 scheme + AWQ algorithm

Multiple formats can be exported simultaneously: --model_export hf_format --model_export gguf

Export CLI Arguments

bash
python quantize_quark.py \
  --model_dir /path/to/model \
  --output_dir /path/to/output \
  --quant_scheme fp8 \
  --model_export hf_format \                    # Export format(s)
  --export_weight_format real_quantized \        # Compression mode
  --custom_mode quark \                          # Export mode: quark|awq|fp8
  --pack_method reorder                          # Weight packing: order|reorder

Evaluation Options

Post-quantization evaluation can be configured as part of the export step:

Perplexity (PPL)
  • Default dataset: wikitext
  • Flag: included by default unless --skip_evaluation is set
Task-Based Evaluation (via lm-eval harness)
bash
--tasks hellaswag,winogrande,arc_easy
--eval_batch_size auto
--num_fewshot 0
ROUGE/METEOR (for generation models)
bash
# Evaluated on cnn_dailymail by default
--use_mlperf_rouge  # For MLPerf-compatible ROUGE scoring
KV Cache Evaluation
bash
--use_ppl_eval_for_kv_cache
--ppl_eval_for_kv_cache_context_size 1024
--ppl_eval_for_kv_cache_sample_size 512
Skip Evaluation
bash
--skip_evaluation  # Skip all post-quantization evaluation

Model Reload for Separate Evaluation

If quantization and evaluation are done in separate steps:

bash
# Step 1: Quantize and export
python quantize_quark.py --model_dir MODEL --quant_scheme fp8 \
  --model_export hf_format --output_dir output/ --skip_evaluation

# Step 2: Reload and evaluate
python quantize_quark.py --model_dir MODEL --model_reload \
  --output_dir output/ --skip_quantization
Show full SKILL.md (193 more words)Show less

Rules

  • Never invent export paths that conflict with the existing run_manifest.yaml or quant_plan.json.
  • Export depends on quantization. If the model has not been quantized yet, this skill produces an export plan attached to the manifest — it does not run quantization.
  • Match export format to deployment target. Ask the user where the model will run: HuggingFace ecosystem → hf_format, llama.cpp → gguf, ONNX Runtime → onnx.
  • GGUF works best with UINT4. If the user wants GGUF but the plan uses FP8, flag the mismatch — GGUF is primarily designed for integer quantization.

Interaction Flow

  1. Check prerequisites: Is there a quant_plan.json? Has quantization been run or is this plan-only?
  2. Choose format: Ask where the model will be deployed and recommend the right export format.
  3. Configure evaluation: Ask if the user wants post-quantization evaluation and which metrics.
  4. Emit: Update run_manifest.yaml with export and evaluation configuration.

Recovery

  • If export is requested before quantization, return the missing prerequisites and attach the export request to the manifest under pending_exports.
  • If GGUF export fails, check that gguf>=0.10.0 is installed and that the tokenizer is accessible.
  • If ONNX export fails on a complex model, suggest trying hf_format first as a fallback.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-export of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch Export next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch Export compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch Export this skillamd/Quark181—~1.5kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Configure G1 Sim2realEGalahad/sim2real145—~1.5kAutomated safety check: PassNone
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Configure G1 Sim2real

    EGalahad/sim2real

    Install, repair, and verify sim2real on G1 robot computers. An agent skill from EGalahad/sim2real.

    145 GitHub stars~1.5k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 9 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 9 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 9 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 9 days ago
    Auto-check passed

Questions about Quark Torch Export

What does Quark Torch Export do?

Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run. Quark Torch Export is an agent skill from amd/Quark. Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run.

When should I use Quark Torch Export?

Quark Torch Export fits situations like: the user wants to export a quantized model to HuggingFace SafeTensors; package for deployment; set up post-quantization evaluation; save quantized model.

How do I install Quark Torch Export in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-export -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-export in amd/Quark) into .claude/skills/quark-torch-export in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch Export in Codex?

Run `npx skills add amd/Quark --skill quark-torch-export -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-export in amd/Quark) into .agents/skills/quark-torch-export in your project. Codex loads it when a task matches its description.

Can I use Quark Torch Export in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-export -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-export, .gemini/skills/quark-torch-export, .github/skills/quark-torch-export and .opencode/skills/quark-torch-export in your project.

What does Quark Torch Export need to run?

Going by SKILL.md and its folder, Quark Torch Export needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Quark Torch Export access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch Export safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch Export use?

Quark Torch Export is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch Export use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Torch Export?

Skills that share tags, products or a category with Quark Torch Export: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Configure G1 Sim2real (EGalahad/sim2real, 145 stars) and Add Model (guoqingbao/xinfer, 333 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch Export?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.