Add Model
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.
$ npx skills add amd/Quark --skill quark-torch-file2file-quantization -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/Quark quark-torch-file2file-quantization --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization .claude/skills/quark-torch-file2file-quantization && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "quark-torch-file2file-quantization" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization into .claude/skills/quark-torch-file2file-quantization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-file2file-quantization", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantizationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/Quark --skill quark-torch-file2file-quantization -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/Quark quark-torch-file2file-quantization --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization .agents/skills/quark-torch-file2file-quantization && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "quark-torch-file2file-quantization" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization into .agents/skills/quark-torch-file2file-quantization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-file2file-quantization", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-file2file-quantization -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/Quark quark-torch-file2file-quantization --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization .cursor/skills/quark-torch-file2file-quantization && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "quark-torch-file2file-quantization" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization into .cursor/skills/quark-torch-file2file-quantization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-file2file-quantization", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/Quark.git --path .claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/Quark --skill quark-torch-file2file-quantization -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/Quark quark-torch-file2file-quantization --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization .gemini/skills/quark-torch-file2file-quantization && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "quark-torch-file2file-quantization" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization into .gemini/skills/quark-torch-file2file-quantization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-file2file-quantization", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/Quark quark-torch-file2file-quantizationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/Quark --skill quark-torch-file2file-quantization -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization .github/skills/quark-torch-file2file-quantization && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-file2file-quantization" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization into .github/skills/quark-torch-file2file-quantization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-file2file-quantization", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-file2file-quantization -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/Quark quark-torch-file2file-quantization --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization .opencode/skills/quark-torch-file2file-quantization && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-file2file-quantization" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization into .opencode/skills/quark-torch-file2file-quantization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-file2file-quantization", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
quark-torch-file2file-quantizationLow-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.
Quark Torch File2file Quantization is an agent skill from amd/Quark. Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. Use when the user wants to run file2file quantization, adapt a new safetensors checkpoint without loading the full model, register an external LLMTemplate, inspect sharded checkpoint naming, generate wrapper or conversion scripts, or validate low-memory sharded quantization outputs. Trigger for "run file2file quantization", "quantize without loading the model", "large safetensors low-memory quantization", "file2file for…
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with DeepSeek and Qwen. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Quark Torch File2file Quantization loads about 2.3k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 724 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 724 words, ~2,325 tokens.
.claude/skills/quark-torch-file2file-quantization/SKILL.md (or your agent's skills folder).Quantize very large safetensors checkpoints without loading the full model into memory, using
ModelQuantizer.direct_quantize_checkpoint(). Covers the full adaptation path: checkpoint
inspection, external LLMTemplate registration, optional naming normalization via a one-time
conversion script, minimum-scale experiment gating, full file2file execution, and output
validation via quark-torch-result-validator.
Default policy: solve naming or layer-selection mismatches with external adapters
(LLMTemplate.register_template(), weight_converters) or a temporary conversion script.
Do not modify Quark source unless the required capability is absent, the external path has
been ruled out, and the user explicitly agrees.
pretrained_model_path — local directory of the safetensors checkpointsave_path — output directoryquant_scheme — quantization scheme (e.g., w_fp8_a_fp8, w_int4_a_bf16)device — e.g., cuda:0, cpumodel_analysis.json from quark-torch-model-intake (optional but recommended)Records the adaptation path taken, scripts generated, experiment results, and validation status.
pretrained_model_path: /models/DeepSeek-V3
save_path: /output/DeepSeek-V3-fp8
quant_scheme: w_fp8_a_fp8
device: cuda:0
adaptation_path: external_template # direct | external_template | weight_converters | conversion_script
conversion_script: null # path if generated
wrapper_script: /tmp/run_ds_v3_f2f.py
min_experiment:
status: passed # passed | failed | skipped
layers_covered: [layer_0_expert_0, ...]
moe_covered: true
full_run_status: completed # pending | completed | failed
validation_status: passed # passed | failed | not_runRead keys and config without loading weights:
python - <<'PY'
import json, os
from glob import glob
from safetensors.torch import safe_open
model_dir = "<pretrained_model_path>"
index_path = os.path.join(model_dir, "model.safetensors.index.json")
print("config:", os.path.exists(os.path.join(model_dir, "config.json")))
print("index:", os.path.exists(index_path))
files = sorted(glob(os.path.join(model_dir, "*.safetensors")))
print("safetensors:", len(files))
if files:
with safe_open(files[0], framework="pt", device="cpu") as f:
keys = list(f.keys())
print("sample_keys (first 80):")
for k in keys[:80]: print(" ", k)
if os.path.exists(index_path):
with open(index_path) as f:
wm = json.load(f).get("weight_map", {})
print("index_keys:", len(wm))
PYVerify: model_type, weight-name suffixes (*.weight, *_scale_inv, *.scale), shard count,
MoE expert / shared-expert / gate naming, and whether scale tensors are co-located with weights.
| Situation | Action |
|---|---|
| Names already match Quark template | Direct file2file; tune exclude_layers only |
| Layer naming differs from built-in template | External LLMTemplate.register_template() |
| Only weight suffixes differ | weight_converters / _apply_weight_converters |
| Scale naming or dtype incompatible pre-recovery | Generate normalization conversion script first |
_apply_weight_converters limits: suited for post-recovery single-suffix rename or one-source split.
Not suited for multi-source merge, cross-shard scale pairing, or FP4→FP8 dtype conversion.
For external template registration, generate a wrapper script (do NOT modify quantize_quark.py):
from quark.torch import ModelQuantizer
from quark.torch.utils.llm import LLMTemplate
template = LLMTemplate(
model_type="<model_type>",
kv_layers_name=["<pattern>"],
q_layer_name=["<pattern>"],
exclude_layers_name=["embed", "head", "<other>"],
)
LLMTemplate.register_template(template)
quantizer = ModelQuantizer(config)
quantizer.direct_quantize_checkpoint(
pretrained_model_path=pretrained_model_path,
save_path=save_path,
keep_excluded_layers_as_original_model_state=False,
weight_converters=weight_converters,
device=device,
)Wrapper must print: registered model_type, input/output dirs, quant scheme, and exclude rules.
Conversion scripts must stream safetensors (no full-model load), include explicit remap_name(),
scale/weight pairing validation, shard output in HF style, index rebuild, and atomic output.
This step is not optional. Full file2file must not run until the minimum experiment passes.
Construct the minimum input:
num_hidden_layers is safely reducible, copy config.json with the smallest value that
still covers at least one MoE layer (for MoE models, use first_moe_layer_id + 1).--key-regex covering at least one complete
MoE expert + its scale tensor + shared expert/router/gate + adjacent non-quantized tensors.num_hidden_layers=1 for a MoE model if layer 0 is dense — verify from config or
key patterns which layer is the first actual MoE layer.Run the minimum experiment, then call quark-torch-result-validator with:
inspect_safetensors, summarize_dtypes, check_index_consistency, check_scale_pairs,
get_fuzzy_tensor_names, and auxiliary-file copy check (if source dir available).
If validation fails, fix template / naming / subset and re-run. Do not proceed to full file2file.
Set cache paths to avoid polluting home quota:
export TMPDIR=/path/to/run/tmp
export TORCH_EXTENSIONS_DIR=/path/to/run/torch_extensions
export TRITON_CACHE_DIR=/path/to/run/triton_cacheRun the wrapper script. After completion, re-run quark-torch-result-validator on the
final output with the same checks as Step 4.
Emit run_manifest.yaml and present to the user:
| Failure | Action |
|---|---|
| Missing Triton / compressed-tensors | Install dependency, retry |
| Incomplete source shards | Repair shards and index before proceeding |
| Template not recognized | Register via external LLMTemplate; do not patch Quark source |
| Scale/weight cannot be paired | Normalize checkpoint first; add unit test if Quark recovery is extended |
| Output index inconsistent | Rebuild index or fix shard write logic; re-validate |
| Minimum experiment fails | Fix adapter/naming/subset; never skip to full run |
LLMTemplate.register_template() or weight_converters before
any Quark source change.num_hidden_layers=1 when layer 0 is dense is invalid.test/test_for_torch/
unit test and run the relevant pytest before committing.LLMTemplate; check whether inference-format naming needs
a conversion script (embed_tokens→embed, self_attn→attn, q_proj→wq, etc.) before file2file._apply_weight_converters.{base}.scale) must be resolvable before recovery; post-recovery suffix
conversion cannot substitute for pre-recovery scale identification.hc_* auxiliary tensors must be explicitly excluded in
the template or conversion script.© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization of amd/Quark.
Open the folder on GitHubat commit 313cb0b
Quark Torch File2file Quantization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Quark Torch File2file Quantization this skillamd/Quark | 181 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Add Modelguoqingbao/xinfer | 333 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Serving LLMs On Instinctamd/skills | 398 | — | ~4k | Automated safety check: Notes | MIT | |
| Miles Rl TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.2k | Automated safety check: Pass | MIT | |
| LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~3.9k | Automated safety check: Pass | None | |
| Update Ollama Cloud Modelsheypinchy/pinchy | 182 | — | ~3.9k | Automated safety check: Notes | AGPL-3.0 |
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
amd/skills
Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.
Orchestra-Research/AI-Research-SKILLs
Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.
BBuf/AI-Infra-Auto-Driven-SKILLS
Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.
heypinchy/pinchy
A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.
ascend-ai-coding/awesome-ascend-skills
vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.
amd/Quark
Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.
amd/Quark
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.
amd/Quark
Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/Quark
Install or verify the AMD Quark package and its dependencies.
amd/Quark
L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…
Categories
Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. Quark Torch File2file Quantization is an agent skill from amd/Quark. Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.
Quark Torch File2file Quantization fits situations like: the user wants to run file2file quantization; adapt a new safetensors checkpoint without loading the full model; register an external LLMTemplate; inspect sharded checkpoint naming.
Run `npx skills add amd/Quark --skill quark-torch-file2file-quantization -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization in amd/Quark) into .claude/skills/quark-torch-file2file-quantization in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/Quark --skill quark-torch-file2file-quantization -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-file2file-quantization in amd/Quark) into .agents/skills/quark-torch-file2file-quantization in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-file2file-quantization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-file2file-quantization, .gemini/skills/quark-torch-file2file-quantization, .github/skills/quark-torch-file2file-quantization and .opencode/skills/quark-torch-file2file-quantization in your project.
Going by SKILL.md and its folder, Quark Torch File2file Quantization needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Quark Torch File2file Quantization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Quark Torch File2file Quantization: Add Model (guoqingbao/xinfer, 333 stars), Serving LLMs On Instinct (amd/skills, 398 stars), Miles Rl Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.
Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.