Agent skill

Quark Torch File2file Quantization

by amd in amd/Quark

Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.

MITAuto-check passedAI & LLM Engineering

Install Quark Torch File2file Quantization

skills CLI
$ npx skills add amd/Quark --skill quark-torch-file2file-quantization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-file2file-quantization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization .claude/skills/quark-torch-file2file-quantization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-file2file-quantization
GitHub stars
181
Token cost
~2.3k tokens
SKILL.md length
724 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.

  • Works in 6 steps: Inspect checkpoint → Choose adaptation path → Generate wrapper or conversion script → …
  • The user wants to run file2file quantization
  • SKILL.md covers Purpose, Inputs, Outputs: run_manifest.yaml and Interaction Flow, plus 3 more sections
  • Calls python

What it does

Quark Torch File2file Quantization is an agent skill from amd/Quark. Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. Use when the user wants to run file2file quantization, adapt a new safetensors checkpoint without loading the full model, register an external LLMTemplate, inspect sharded checkpoint naming, generate wrapper or conversion scripts, or validate low-memory sharded quantization outputs. Trigger for "run file2file quantization", "quantize without loading the model", "large safetensors low-memory quantization", "file2file for…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with DeepSeek and Qwen. The licence is MIT.

When your agent uses it

  • The user wants to run file2file quantization
  • Adapt a new safetensors checkpoint without loading the full model
  • Register an external LLMTemplate
  • Inspect sharded checkpoint naming

Example prompts

  • “run file2file quantization”
  • “quantize without loading the model”
  • “large safetensors low-memory quantization”
  • “/quark-torch-file2file-quantization”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inspect checkpoint
  2. Choose adaptation path
  3. Generate wrapper or conversion script
  4. Minimum-scale experiment (mandatory gate)
  5. Full file2file
  6. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch File2file Quantization loads about 2.3k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 724 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 724 words, ~2,325 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-file2file-quantization/SKILL.md (or your agent's skills folder).
name
quark-torch-file2file-quantization
description
Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. Use when the user wants to run file2file quantization, adapt a new safetensors checkpoint without loading the full model, register an external LLMTemplate, inspect sharded checkpoint naming, generate wrapper or conversion scripts, or validate low-memory sharded quantization outputs. Trigger for "run file2file quantization", "quantize without loading the model", "large safetensors low-memory quantization", "file2file for DeepSeek/Qwen/MoE", "safetensors naming incompatible", "register LLMTemplate externally".
layer
l1-atomic
primary_artifact
run_manifest.yaml
source_knowledge
quark/torch/quantization/file2file_quantization.py, quark/torch/quantization/weight_convert.py, quark/torch/utils/llm/model_preparation.py…

quark-file2file-quantization-runner

Purpose

Quantize very large safetensors checkpoints without loading the full model into memory, using ModelQuantizer.direct_quantize_checkpoint(). Covers the full adaptation path: checkpoint inspection, external LLMTemplate registration, optional naming normalization via a one-time conversion script, minimum-scale experiment gating, full file2file execution, and output validation via quark-torch-result-validator.

Default policy: solve naming or layer-selection mismatches with external adapters (LLMTemplate.register_template(), weight_converters) or a temporary conversion script. Do not modify Quark source unless the required capability is absent, the external path has been ruled out, and the user explicitly agrees.

Inputs

  • pretrained_model_path — local directory of the safetensors checkpoint
  • save_path — output directory
  • quant_scheme — quantization scheme (e.g., w_fp8_a_fp8, w_int4_a_bf16)
  • device — e.g., cuda:0, cpu
  • User intent: direct quantization, checkpoint conversion first, HF-format export, or script generation only
  • model_analysis.json from quark-torch-model-intake (optional but recommended)

Outputs: run_manifest.yaml

Records the adaptation path taken, scripts generated, experiment results, and validation status.

yaml
pretrained_model_path: /models/DeepSeek-V3
save_path: /output/DeepSeek-V3-fp8
quant_scheme: w_fp8_a_fp8
device: cuda:0
adaptation_path: external_template          # direct | external_template | weight_converters | conversion_script
conversion_script: null                     # path if generated
wrapper_script: /tmp/run_ds_v3_f2f.py
min_experiment:
  status: passed                            # passed | failed | skipped
  layers_covered: [layer_0_expert_0, ...]
  moe_covered: true
full_run_status: completed                  # pending | completed | failed
validation_status: passed                  # passed | failed | not_run

Interaction Flow

Step 1 — Inspect checkpoint

Read keys and config without loading weights:

bash
python - <<'PY'
import json, os
from glob import glob
from safetensors.torch import safe_open

model_dir = "<pretrained_model_path>"
index_path = os.path.join(model_dir, "model.safetensors.index.json")
print("config:", os.path.exists(os.path.join(model_dir, "config.json")))
print("index:", os.path.exists(index_path))
files = sorted(glob(os.path.join(model_dir, "*.safetensors")))
print("safetensors:", len(files))
if files:
    with safe_open(files[0], framework="pt", device="cpu") as f:
        keys = list(f.keys())
    print("sample_keys (first 80):")
    for k in keys[:80]: print(" ", k)
if os.path.exists(index_path):
    with open(index_path) as f:
        wm = json.load(f).get("weight_map", {})
    print("index_keys:", len(wm))
PY

Verify: model_type, weight-name suffixes (*.weight, *_scale_inv, *.scale), shard count, MoE expert / shared-expert / gate naming, and whether scale tensors are co-located with weights.

Step 2 — Choose adaptation path
SituationAction
Names already match Quark templateDirect file2file; tune exclude_layers only
Layer naming differs from built-in templateExternal LLMTemplate.register_template()
Only weight suffixes differweight_converters / _apply_weight_converters
Scale naming or dtype incompatible pre-recoveryGenerate normalization conversion script first

_apply_weight_converters limits: suited for post-recovery single-suffix rename or one-source split. Not suited for multi-source merge, cross-shard scale pairing, or FP4→FP8 dtype conversion.

Step 3 — Generate wrapper or conversion script

For external template registration, generate a wrapper script (do NOT modify quantize_quark.py):

python
from quark.torch import ModelQuantizer
from quark.torch.utils.llm import LLMTemplate

template = LLMTemplate(
    model_type="<model_type>",
    kv_layers_name=["<pattern>"],
    q_layer_name=["<pattern>"],
    exclude_layers_name=["embed", "head", "<other>"],
)
LLMTemplate.register_template(template)

quantizer = ModelQuantizer(config)
quantizer.direct_quantize_checkpoint(
    pretrained_model_path=pretrained_model_path,
    save_path=save_path,
    keep_excluded_layers_as_original_model_state=False,
    weight_converters=weight_converters,
    device=device,
)

Wrapper must print: registered model_type, input/output dirs, quant scheme, and exclude rules. Conversion scripts must stream safetensors (no full-model load), include explicit remap_name(), scale/weight pairing validation, shard output in HF style, index rebuild, and atomic output.

Step 4 — Minimum-scale experiment (mandatory gate)

This step is not optional. Full file2file must not run until the minimum experiment passes.

Construct the minimum input:

  1. If num_hidden_layers is safely reducible, copy config.json with the smallest value that still covers at least one MoE layer (for MoE models, use first_moe_layer_id + 1).
  2. If not, generate a subset checkpoint filtered by --key-regex covering at least one complete MoE expert + its scale tensor + shared expert/router/gate + adjacent non-quantized tensors.
  3. Never set num_hidden_layers=1 for a MoE model if layer 0 is dense — verify from config or key patterns which layer is the first actual MoE layer.

Run the minimum experiment, then call quark-torch-result-validator with: inspect_safetensors, summarize_dtypes, check_index_consistency, check_scale_pairs, get_fuzzy_tensor_names, and auxiliary-file copy check (if source dir available).

If validation fails, fix template / naming / subset and re-run. Do not proceed to full file2file.

Show full SKILL.md (292 more words)Show less
Step 5 — Full file2file

Set cache paths to avoid polluting home quota:

bash
export TMPDIR=/path/to/run/tmp
export TORCH_EXTENSIONS_DIR=/path/to/run/torch_extensions
export TRITON_CACHE_DIR=/path/to/run/triton_cache

Run the wrapper script. After completion, re-run quark-torch-result-validator on the final output with the same checks as Step 4.

Step 6 — Deliver

Emit run_manifest.yaml and present to the user:

  • Script paths and execution commands
  • Minimum experiment summary: layers selected, MoE coverage, validation outcome
  • Full run validation summary
  • Any outstanding risk items

Recovery

FailureAction
Missing Triton / compressed-tensorsInstall dependency, retry
Incomplete source shardsRepair shards and index before proceeding
Template not recognizedRegister via external LLMTemplate; do not patch Quark source
Scale/weight cannot be pairedNormalize checkpoint first; add unit test if Quark recovery is extended
Output index inconsistentRebuild index or fix shard write logic; re-validate
Minimum experiment failsFix adapter/naming/subset; never skip to full run

Rules

  • Never load full weights for inspection — read safetensors headers only.
  • External adapters first: use LLMTemplate.register_template() or weight_converters before any Quark source change.
  • Minimum experiment is a gate, not a hint — full file2file is blocked until it passes.
  • MoE minimum experiment must include actual MoE expert weights + their scale tensors. Setting num_hidden_layers=1 when layer 0 is dense is invalid.
  • No cluster paths in shared scripts — keep one-time paths in temporary wrapper scripts only.
  • Source changes require tests — if Quark source must change, add a test/test_for_torch/ unit test and run the relevant pytest before committing.

Notes

  • DeepSeek-V4-family: prefer external LLMTemplate; check whether inference-format naming needs a conversion script (embed_tokens→embed, self_attn→attn, q_proj→wq, etc.) before file2file.
  • FP4/e2m1fn expert weights must be converted to FP8/e4m3fn in the conversion script, not via _apply_weight_converters.
  • Sibling scale naming ({base}.scale) must be resolvable before recovery; post-recovery suffix conversion cannot substitute for pre-recovery scale identification.
  • MTP embedding/head, attention, gate, hc_* auxiliary tensors must be explicitly excluded in the template or conversion script.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch File2file Quantization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch File2file Quantization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch File2file Quantization this skillamd/Quark181—~2.3kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Serving LLMs On Instinctamd/skills398—~4kAutomated safety check: NotesMIT
Miles Rl TrainingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.2kAutomated safety check: PassMIT
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~3.9kAutomated safety check: PassNone
Update Ollama Cloud Modelsheypinchy/pinchy182—~3.9kAutomated safety check: NotesAGPL-3.0

Similar skills

  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    398 GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Miles Rl Training

    Orchestra-Research/AI-Research-SKILLs

    Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

    13k GitHub starsUsed in 3 repos~2.2k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    911 GitHub stars~3.9k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.

    182 GitHub stars~3.9k tokensUpdated 17 days ago
    AI & LLM EngineeringAuto-check: notes
  • Vllm Ascend

    ascend-ai-coding/awesome-ascend-skills

    vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

    174 GitHub stars~2.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 10 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 10 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about Quark Torch File2file Quantization

What does Quark Torch File2file Quantization do?

Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. Quark Torch File2file Quantization is an agent skill from amd/Quark. Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.

When should I use Quark Torch File2file Quantization?

Quark Torch File2file Quantization fits situations like: the user wants to run file2file quantization; adapt a new safetensors checkpoint without loading the full model; register an external LLMTemplate; inspect sharded checkpoint naming.

How do I install Quark Torch File2file Quantization in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-file2file-quantization -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization in amd/Quark) into .claude/skills/quark-torch-file2file-quantization in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch File2file Quantization in Codex?

Run `npx skills add amd/Quark --skill quark-torch-file2file-quantization -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization in amd/Quark) into .agents/skills/quark-torch-file2file-quantization in your project. Codex loads it when a task matches its description.

Can I use Quark Torch File2file Quantization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-file2file-quantization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-file2file-quantization, .gemini/skills/quark-torch-file2file-quantization, .github/skills/quark-torch-file2file-quantization and .opencode/skills/quark-torch-file2file-quantization in your project.

What does Quark Torch File2file Quantization need to run?

Going by SKILL.md and its folder, Quark Torch File2file Quantization needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Quark Torch File2file Quantization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch File2file Quantization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch File2file Quantization use?

Quark Torch File2file Quantization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch File2file Quantization use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Torch File2file Quantization?

Skills that share tags, products or a category with Quark Torch File2file Quantization: Add Model (guoqingbao/xinfer, 333 stars), Serving LLMs On Instinct (amd/skills, 398 stars), Miles Rl Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch File2file Quantization?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.