Agent skill

Quark Torch Model Intake

by amd in amd/Quark

Inspect a target model and prepare metadata for Quark PTQ planning.

MITAuto-check passedAI & LLM Engineering

Install Quark Torch Model Intake

skills CLI
$ npx skills add amd/Quark --skill quark-torch-model-intake -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-torch-model-intake --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-model-intake .claude/skills/quark-torch-model-intake && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-torch-model-intake
GitHub stars
181
Token cost
~2k tokens
SKILL.md length
605 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Inspect a target model and prepare metadata for Quark PTQ planning.

  • Works in 5 steps: Confirm model source: Is it a local… → Read config: Extract architecture facts… → Assess compatibility: Check model type… → …
  • The user needs model path validation
  • SKILL.md covers Purpose, Inputs, Outputs: model_analysis.json and Supported Model Architectures, plus 5 more sections
  • Calls python3

What it does

Quark Torch Model Intake is an agent skill from amd/Quark. Inspect a target model and prepare metadata for Quark PTQ planning. Use when the user needs model path validation, architecture detection, quantization target discovery, layer counting, risk assessment, or transformer compatibility checks before planning PTQ. Trigger for "analyze my model", "check this model", "what architecture is this", "can Quark quantize X", "is this model supported", or when any quantization step needs model facts that are missing.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Qwen and DeepSeek. The licence is MIT.

When your agent uses it

  • The user needs model path validation
  • Architecture detection
  • Quantization target discovery
  • Risk assessment

Example prompts

  • “analyze my model”
  • “check this model”
  • “what architecture is this”
  • “/quark-torch-model-intake”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Confirm model source: Is it a local directory or a HuggingFace ID? Check if quark-workspace-validate already confirmed the path.
  2. Read config: Extract architecture facts from config.json. Present a summary table to the user.
  3. Assess compatibility: Check model type against Quark's template list. Flag any transformer version requirements.
  4. Identify risks: Note anything that could affect downstream PTQ planning.
  5. Emit: Write model_analysis.json. Surface any new constraints back to quark-torch-router so they land in session_context.json.

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Torch Model Intake loads about 2k tokens when it runs. Until then it costs about 121 tokens; SKILL.md has 605 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~121
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 605 words, ~1,976 tokens.

Download SKILL.mdSave it as .claude/skills/quark-torch-model-intake/SKILL.md (or your agent's skills folder).
name
quark-torch-model-intake
description
Inspect a target model and prepare metadata for Quark PTQ planning. Use when the user needs model path validation, architecture detection, quantization target discovery, layer counting, risk assessment, or transformer compatibility checks before planning PTQ. Trigger for "analyze my model", "check this model", "what architecture is this", "can Quark quantize X", "is this model supported", or when any quantization step needs model facts that are missing.
layer
l1-atomic
primary_artifact
model_analysis.json
source_knowledge
examples/torch/language_modeling/llm_ptq/quantize_quark.py, quark/torch/utils/llm/model_preparation.py, quark/torch/quantization/config/template.py

quark-torch-model-intake

Purpose

Validate the target model and extract the structural facts that quark-torch-quant-plan needs to make correct quantization decisions. This step exists because different model architectures have different quantization requirements — MoE models need expert module replacement, some models require trust_remote_code, and certain architectures have known compatibility issues with specific transformers versions.

Inputs

  • Model path (HuggingFace ID or local directory)
  • env_context.json for Python and accelerator constraints
  • workspace_context.json for the validated model path

Outputs: model_analysis.json

Captures architecture facts, quantizable layer count, transformer compatibility, and risks.

Schema: model_analysis.schema.json

json
{
  "analysis_status": "complete",
  "model": {
    "model_path": "Qwen/Qwen3-8B",
    "model_type": "qwen3",
    "trust_remote_code": false,
    "transformers_version_required": ">=4.48.0,<5.3",
    "loading_class": "AutoModelForCausalLM",
    "is_moe": false,
    "num_hidden_layers": 36,
    "hidden_size": 4096,
    "estimated_size_gb": 16.0
  },
  "quantization_targets": {
    "linear_layer_count": 224,
    "has_non_linear_experts": false,
    "exclude_defaults": ["lm_head"],
    "needs_moe_preparation": false
  },
  "risks": [
    {
      "severity": "low",
      "message": "Model size fits in single GPU with 24GB+ VRAM for FP8/INT8 schemes.",
      "recovery_hint": "Use --multi_gpu for INT4 with large calibration datasets if OOM occurs."
    }
  ]
}

Supported Model Architectures

Quark has built-in templates for 36 model families:

CategoryModels
Standard LLMsllama, mistral, opt, phi, phi3, qwen, qwen2, gptj, cohere, olmo
Advanced LLMsqwen3, qwen3_next, deepseek, deepseek_v2, deepseek_v3, deepseek_v32
MoE Modelsmixtral, dbrx, llama4, qwen2_moe, qwen3_moe, qwen3_5_moe, gpt_oss, granitemoehybrid, glm4_moe
Vision-Languagemllama, deepseek_vl_v2, qwen3_vl_moe
Otherchatglm, gemma2, gemma3, gemma3_text, grok-1, instella, kimi_k25, minimax_m2

Models not in this list may still work if they follow standard HuggingFace AutoModelForCausalLM patterns, but need extra attention.

What to Extract

From config.json
  • model_type — must match a Quark template name (e.g., "llama", "qwen3", "mistral")
  • num_hidden_layers — determines the number of quantizable linear layers
  • hidden_size, intermediate_size — affects memory estimates
  • num_attention_heads, num_key_value_heads — relevant for KV cache quantization
  • Architecture-specific fields (e.g., num_experts for MoE models)
Quantization Targets
  • Count total linear layers (these are what gets quantized)
  • Identify default exclusions — lm_head is almost always excluded
  • For MoE models, note that expert modules need special preparation via prepare_for_moe_quant()
Transformer Compatibility

Some models require specific minimum transformers versions:

  • llama4 → transformers >= 4.51.0
  • gpt_oss, granitemoehybrid → transformers >= 4.55.1
  • qwen3_vl_moe → transformers >= 4.57.0
  • qwen3_5_moe → transformers >= 5.2.0
  • General requirement for LLM PTQ: transformers < 5.3
Special Loading Requirements
  • deepseek_vl_v2 → uses AutoModel instead of AutoModelForCausalLM
  • mllama → uses MllamaForConditionalGeneration
  • gpt_oss → needs Mxfp4Config(dequantize=True) for loading
  • Models with custom code → need trust_remote_code=True
Risks

Flag anything that could cause failures downstream:

  • Model type not in Quark's template list
  • Transformer version incompatibility
  • Very large models that may need --multi_gpu or --file2file_quantization
  • MoE models that need module replacement

Concrete Actions

Action 1: Read model config (NEVER load full weights)
bash
python3 -c "
from transformers import AutoConfig
import json
config = AutoConfig.from_pretrained('<MODEL_PATH>', trust_remote_code=True)
info = {
    'model_type': config.model_type,
    'num_hidden_layers': config.num_hidden_layers,
    'hidden_size': config.hidden_size,
    'intermediate_size': getattr(config, 'intermediate_size', None),
    'num_attention_heads': config.num_attention_heads,
    'num_key_value_heads': getattr(config, 'num_key_value_heads', None),
    'vocab_size': config.vocab_size,
    'num_experts': getattr(config, 'num_local_experts', getattr(config, 'num_experts', None)),
    'torch_dtype': str(getattr(config, 'torch_dtype', 'unknown')),
}
print(json.dumps(info, indent=2))
"
Action 2: Check model size (for local models)
bash
du -sh /path/to/model/
ls -lh /path/to/model/*.safetensors
Action 3: Present summary table to user

After running the above, format the results as:

text
Model Analysis:
  Model path:       <path or HuggingFace ID>
  Model type:       <model_type from config>
  Hidden layers:    <num_hidden_layers>
  Linear layers:    ~<estimated count>
  MoE:              Yes/No
  Exclude defaults: [lm_head]
  Risks:            <list or "None">
  Compatibility:    OK / <version constraints>

This table is what the user sees at Checkpoint 1 of quark-torch-llm-ptq-workflow.

Show full SKILL.md (239 more words)Show less

Rules

  • Run or reference quark-workspace-validate first to confirm that model paths are valid before attempting to read config.json.
  • Do not load the full model during intake. Reading config.json and listing files is enough — loading weights is expensive and belongs to the quantization step.
  • Preserve ambiguity when a model reference could be local or remote. Note both possibilities and let the user resolve.
  • Surface new environment constraints (e.g., model requires transformers >= 4.57.0) by recording them in model_analysis.json under risks and asking quark-torch-router to add them to session_context.json's open_questions. Do not write directly to env_context.json.

Interaction Flow

  1. Confirm model source: Is it a local directory or a HuggingFace ID? Check if quark-workspace-validate already confirmed the path.
  2. Read config: Extract architecture facts from config.json. Present a summary table to the user.
  3. Assess compatibility: Check model type against Quark's template list. Flag any transformer version requirements.
  4. Identify risks: Note anything that could affect downstream PTQ planning.
  5. Emit: Write model_analysis.json. Surface any new constraints back to quark-torch-router so they land in session_context.json.

Recovery

  • If analysis_status: "partial" — some facts were extracted but the model could not be fully inspected. Common cause: model needs trust_remote_code=True but the user has not approved it.
  • If model type is not in Quark's template list — report this clearly. The user may need to register a custom template (see LLMTemplate.register_template() in quantize_quark.py).
  • If transformer version is incompatible — show the exact version mismatch and the upgrade/downgrade command.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/l1-atomic/torch/quark-torch-model-intake of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Torch Model Intake next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Torch Model Intake compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Torch Model Intake this skillamd/Quark181—~2kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Serving LLMs On Instinctamd/skills398—~4kAutomated safety check: NotesMIT
Miles Rl TrainingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.2kAutomated safety check: PassMIT
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~3.9kAutomated safety check: PassNone
Update Ollama Cloud Modelsheypinchy/pinchy182—~3.9kAutomated safety check: NotesAGPL-3.0

Similar skills

  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    398 GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Miles Rl Training

    Orchestra-Research/AI-Research-SKILLs

    Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

    13k GitHub starsUsed in 3 repos~2.2k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    911 GitHub stars~3.9k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.

    182 GitHub stars~3.9k tokensUpdated 17 days ago
    AI & LLM EngineeringAuto-check: notes
  • Vllm Ascend

    ascend-ai-coding/awesome-ascend-skills

    vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

    174 GitHub stars~2.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 10 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 10 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 10 days ago
    Auto-check passed

Works with

Questions about Quark Torch Model Intake

What does Quark Torch Model Intake do?

Inspect a target model and prepare metadata for Quark PTQ planning. Quark Torch Model Intake is an agent skill from amd/Quark. Inspect a target model and prepare metadata for Quark PTQ planning.

When should I use Quark Torch Model Intake?

Quark Torch Model Intake fits situations like: the user needs model path validation; architecture detection; quantization target discovery; risk assessment.

How do I install Quark Torch Model Intake in Claude Code?

Run `npx skills add amd/Quark --skill quark-torch-model-intake -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-model-intake in amd/Quark) into .claude/skills/quark-torch-model-intake in your project. Claude Code loads it when a task matches its description.

How do I install Quark Torch Model Intake in Codex?

Run `npx skills add amd/Quark --skill quark-torch-model-intake -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-model-intake in amd/Quark) into .agents/skills/quark-torch-model-intake in your project. Codex loads it when a task matches its description.

Can I use Quark Torch Model Intake in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-model-intake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-model-intake, .gemini/skills/quark-torch-model-intake, .github/skills/quark-torch-model-intake and .opencode/skills/quark-torch-model-intake in your project.

What does Quark Torch Model Intake need to run?

Going by SKILL.md and its folder, Quark Torch Model Intake needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Quark Torch Model Intake access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Torch Model Intake safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Torch Model Intake use?

Quark Torch Model Intake is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Torch Model Intake use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Torch Model Intake?

Skills that share tags, products or a category with Quark Torch Model Intake: Add Model (guoqingbao/xinfer, 333 stars), Serving LLMs On Instinct (amd/skills, 398 stars), Miles Rl Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Torch Model Intake?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.