Agent skill

Quark Onnx Subgraph Partitioner

by amd in amd/Quark

Partition an ONNX model graph into named functional subgraphs and emit a subgraphpartition.json file.

MITAuto-check passedAI & LLM Engineering

Install Quark Onnx Subgraph Partitioner

skills CLI
$ npx skills add amd/Quark --skill quark-onnx-subgraph-partitioner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-onnx-subgraph-partitioner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/_legacy_impl/l1-atomic/onnx/quark-onnx-subgraph-partitioner .claude/skills/quark-onnx-subgraph-partitioner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-onnx-subgraph-partitioner
GitHub stars
182
Token cost
~2.7k tokens
SKILL.md length
1,075 words
Files
3
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Partition an ONNX model graph into named functional subgraphs and emit a subgraphpartition.json file.

  • Works in 7 steps: Load and inspect the model → Identify structural landmarks → Assign blocks → …
  • The user wants to understand a models high-level structure
  • SKILL.md covers Purpose, Inputs, Outputs: subgraph_partition.json and Analysis Procedure, plus 3 more sections
  • Runs Python scripts from its folder; calls python3 and pip

What it does

Quark Onnx Subgraph Partitioner is an agent skill from amd/Quark. Partition an ONNX model graph into named functional subgraphs and emit a subgraphpartition.json file. Use when the user wants to understand a model's high-level structure, document its architectural blocks, or prepare a partition for downstream workflows (mixed-precision quantization, layer-wise profiling, partial deployment, graph visualization). Triggers on "partition my ONNX model", "generate a subgraph JSON", "which nodes belong to the attention blocks", "show me the model architecture", "split the model into…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `inspect_model.py` and `validate_partition.py`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with ONNX. The licence is MIT.

When your agent uses it

  • The user wants to understand a models high-level structure
  • Document its architectural blocks
  • Prepare a partition for downstream workflows (mixed-precision quantization
  • Layer-wise profiling

Example prompts

  • “partition my ONNX model”
  • “generate a subgraph JSON”
  • “which nodes belong to the attention blocks”
  • “/quark-onnx-subgraph-partitioner”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Load and inspect the model
  2. Identify structural landmarks
  3. Assign blocks
  4. Choose start and end nodes
  5. Validate the partition
  6. Emit the JSON file
  7. Summary table

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Onnx Subgraph Partitioner loads about 2.7k tokens when it runs. Until then it costs about 244 tokens; SKILL.md has 1,075 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~244
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,075 words, ~2,672 tokens.

Download SKILL.mdSave it as .claude/skills/quark-onnx-subgraph-partitioner/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
quark-onnx-subgraph-partitioner
description
Partition an ONNX model graph into named functional subgraphs and emit a subgraph_partition.json file. Use when the user wants to understand a model's high-level structure, document its architectural blocks, or prepare a partition for downstream workflows (mixed-precision quantization, layer-wise profiling, partial deployment, graph visualization). Triggers on "partition my ONNX model", "generate a subgraph JSON", "which nodes belong to the attention blocks", "show me the model architecture", "split the model into blocks", "group layers for AMP", or any request that needs a named block decomposition of an ONNX graph. Handles any ONNX architecture: CNNs (ResNet, EfficientNet, YOLOv5/v8, DenseNet, MobileNet), Transformers (BERT, ViT, Swin, LLMs), multi-modal (CLIP, BLIP), detection models with FPN/PAN necks, BEV perception models (BEVFormer), and custom or unfamiliar architectures identified from node-name prefixes and op sequences.
layer
l1-atomic
primary_artifact
subgraph_partition.json
source_knowledge
quark/onnx/algorithm/mprecision/subgraph_parser.py, quark/onnx/algorithm/mprecision/mprecision_config.py…

quark-onnx-subgraph-partitioner

Purpose

Produce a subgraph_partition.json that groups ONNX graph nodes into named functional blocks (stem, backbone stage, attention layer, FFN, FPN neck, etc.). The partition can serve as input to AMP sensitivity analysis, layer-wise profiling, partial deployment, or graph documentation.

Two helper scripts live alongside this file:

bash
SKILL_DIR=skills/_legacy_impl/l1-atomic/onnx/quark-onnx-subgraph-partitioner

Inputs

  • Model path — a single .onnx file supplied as $ARGUMENTS
  • Optional architecture hints from the user (layer count, naming convention, block granularity preference)

Outputs: subgraph_partition.json

Schema: subgraph_partition.schema.json

json
{
  "quantized":     false,
  "num_subgraphs": 4,
  "subgraphs": [
    {
      "name": "backbone_stem",
      "description": "Initial stride-2 Conv → SiLU activation, produces P1 features",
      "start_nodes": ["/model.0/conv/Conv"],
      "end_nodes":   ["/model.0/act/Mul"]
    },
    {
      "name": "attention_layer_0",
      "description": "Full encoder layer 0: self-attention + FFN + LayerNorm + residuals",
      "start_nodes": ["/encoder/layer.0/attention/self/query/MatMul"],
      "end_nodes":   ["/encoder/layer.0/output/LayerNorm"]
    }
  ]
}

Every entry requires all four fields: name, description, start_nodes, end_nodes. Never include resolved_nodes — node resolution happens at runtime. num_subgraphs must equal len(subgraphs); set it last.

Analysis Procedure

Step 1 — Load and inspect the model
bash
python3 "$SKILL_DIR/inspect_model.py" <MODEL_PATH>

For quantized models (detected automatically via Q/DQ op presence), Q/DQ wrapper nodes are suppressed so structural landmarks remain visible; Q/DQ counts are reported separately.

Text .onnxtxt — use grep instead:

bash
grep -n 'op_type:\|  name:\|input:\|output:' <MODEL_PATH> | head -400

Quantized model rules: use compute nodes (Conv, MatMul, Add, …) as boundaries — never Q/DQ nodes. BFS from a compute node captures surrounding Q/DQ wrappers automatically.

Step 2 — Identify structural landmarks
SignalLikely boundary
First Conv consuming the graph inputBackbone / model stem
Repeated /model.N/ or /layer.N/ prefix groupsStage or layer repetition
MaxPool or strided Conv (stride 2)Resolution downsampling
Resize (upsample) + Concat on lateral pathFPN / PAN top-down merge
Split → N parallel Conv → Concat + 1×1 ConvC2f / CSP bottleneck
GlobalAveragePool → FC → activation → MulSE block
MatMul×3 → Softmax → MatMul (QKV pattern)Self-attention block
LayerNormalization or InstanceNormalizationTransformer / diffusion layer
MatMul/Gemm → activation → MatMul/Gemm (hidden × 4)FFN / MLP block
Add immediately after attention + FFNResidual add (layer end)
GridSample with sampling offsetsDeformable attention
Gather on index 0 → LayerNorm → GemmCLS token / pooler / head
Step 3 — Assign blocks
Block typeDescription
StemInitial convs before first downsampling
BackboneStageStage at one resolution (N residual / bottleneck blocks)
SEBlockSqueeze-and-Excitation (GAP → FC → activation → Mul)
PatchEmbeddingViT patch projection (Conv with patch-stride kernel)
SelfAttentionBlockMulti-head self-attention (QKV + Softmax + out proj)
FFNBlockFeed-forward network (Linear → activation → Linear)
TransformerLayerFull encoder/decoder layer (attn + FFN + norms + residuals)
SPPFSpatial Pyramid Pooling Fast
FPNNeckFeature Pyramid Network (lateral convs + upsample + merge)
PANNeckPath Aggregation Network (down-path convs + concat + merge)
DetectionHeadConv layers + decode logic for bbox / class output
ClassificationHeadMLP predicting class logits

Granularity rules:

  • No single subgraph > ~25 % of total nodes; split if needed.
  • No trivial single-activation subgraphs; minimum unit is one compute op.
  • Split every repeated layer individually (12 transformer layers → 12 entries).
  • Merge inseparable op chains (Conv + BN + ReLU with no branch points).
  • Assign Concat / Add / Resize to the subgraph that produces the feature map.
Step 4 — Choose start and end nodes

start_nodes: first node(s) receiving data from outside the block. Only list a node if it is not already reachable from another listed start node — downstream-reachable additions are redundant. When a block has multiple independent entry paths (e.g., parallel detection-head branches), list all.

end_nodes: last node(s) whose output feeds the next block.

Residual fork hazard. BFS stops when it visits an end node, but still enqueues all consumers of every preceding node. If any node before the end node fans out to an external consumer (a residual skip, a lateral FPN branch), BFS cascades through the rest of the model. Fix: move end_nodes back to the last node whose output has no external consumers, i.e. the node immediately before the fork. Check 5 in validate_partition.py catches this automatically by flagging subgraphs that resolve to > 25 % of total nodes.

Copy all node names verbatim from the Step 1 listing — never reconstruct them.

Step 5 — Validate the partition
bash
python3 "$SKILL_DIR/validate_partition.py" <MODEL_PATH> <DRAFT_JSON>

Runs five checks (see validate_partition.py docstring for details). Fix all ERROR: lines before proceeding. Common fixes:

  • Replace Constant start nodes with the Resize / Conv that consumes them.
  • Replace initializer names with the node that reads them.
  • Adjust end_nodes so all end nodes are reachable from start_nodes.
  • Move end_nodes one step earlier to avoid residual-fork BFS explosion.

Also verify coverage: confirm the union of resolved subgraphs covers the model's quantizable ops (Conv / MatMul / Gemm). Explain any ops that land in __ungrouped__.

Show full SKILL.md (418 more words)Show less
Step 6 — Emit the JSON file

Write subgraph_partition.json beside the model file (or in the working directory). Before writing, verify:

  • All four fields on every entry; no duplicate name values.
  • num_subgraphs equals len(subgraphs) (count programmatically).
  • All Step 5 validation checks pass with zero errors.
Step 6 — Summary table

Always finish with:

#NameBlock typeStart nodeEnd nodeEst. nodes

Include an Architectural note: overall architecture family, number and type of repeated blocks, and any non-obvious design choices visible in the graph.

Rules

  • All four fields required on every entry; omitting description is not acceptable.
  • Never fabricate node names. Copy from Step 1 listing. If unconfirmable, omit rather than guess.
  • Never use Constant / ConstantOfShape as start/end node (shape side-branches).
  • Never use Q/DQ wrapper nodes as start/end node; always anchor on compute nodes.
  • Never use initializer or graph-input names as node names.
  • Never use an empty string as a node name.
  • Never include resolved_nodes / nodes in the output.
  • Split every repeated layer separately (N layers → N entries).
  • Cover the whole model and explain any __ungrouped__ nodes.
  • Run Step 5 validation and fix all errors before writing the file.
  • Always write the file to disk and report its absolute path.

Interaction Flow

  1. Confirm the model path from $ARGUMENTS; resolve to absolute path.
  2. Run Step 1 — print the node listing.
  3. Run Steps 2–4 — identify landmarks, assign blocks, select node names. Ask one focused question if granularity is ambiguous.
  4. Run Step 5 — fix all validation errors before proceeding.
  5. Run Step 6 — write subgraph_partition.json; confirm zero errors.
  6. Run Step 7 — print summary table and architectural note.
  7. Confirm output path and mention typical next steps:
    • AMP: AutoMixprecisionConfig(subgraph_json="<path>")
    • Profiling: iterate over resolved_nodes per block
    • Documentation: open alongside Netron

Recovery

FailureAction
onnx not installedRun pip install onnx and retry Step 1
Model > 2 GB, Python OOMPrint only n.name and n.op_type; skip weight shapes
All node names are anonymous (/Add_7)Use op-type sequences and output shapes; use index ranges as block names
Validation: node not foundRe-check Step 1 listing; never guess — omit if unconfirmable
Validation: name is an initializerUse the graph node that reads the initializer instead
Validation: start node is ConstantReplace with the data-path node consuming it
Validation: start/end is a Q/DQ wrapperReplace with the adjacent compute node
Validation: end not reachable from startAdd the branching node to start_nodes, or adjust end_nodes
Check 5: subgraph resolves to > 25 %Residual-fork hazard; move end_nodes one step earlier to the DQL before the fork

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/_legacy_impl/l1-atomic/onnx/quark-onnx-subgraph-partitioner of amd/Quark.

  • SKILL.md
  • inspect_model.py
  • validate_partition.py

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Onnx Subgraph Partitioner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Onnx Subgraph Partitioner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Onnx Subgraph Partitioner this skillamd/Quark182—~2.7kAutomated safety check: PassMIT
Aipc Toolkitqualcomm/qai-appbuilder247—~5.7kAutomated safety check: NotesCustom licence
Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassCustom licence
Onboard Jetpack5 Inference BackendsEGalahad/sim2real146—~1.1kAutomated safety check: PassNone
Model Builderqualcomm/qai-appbuilder247—~4.1kAutomated safety check: PassBSD-3-Clause
Engine Performancescragnog/HOT-Step-CPP174—~4.9kAutomated safety check: PassMIT

Similar skills

  • Aipc Toolkit

    qualcomm/qai-appbuilder

    AIPC, AI Porting Conversion. An agent skill from qualcomm/qai-appbuilder.

    247 GitHub stars~5.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Matlab Use Visual Inspection

    matlab/matlab-agentic-toolkit

    Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

    1.1k GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

    146 GitHub stars~1.1k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Model Builder

    qualcomm/qai-appbuilder

    QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

    247 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Engine Performance

    scragnog/HOT-Step-CPP

    Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

    174 GitHub stars~4.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Running Openmed Ondevice

    maziyarpanahi/openmed

    Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows.

    5.5k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    182 GitHub stars~3.1k tokensUpdated 13 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    182 GitHub stars~3k tokensUpdated 13 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    182 GitHub stars~2.9k tokensUpdated 13 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    182 GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    182 GitHub stars~1.8k tokensUpdated 13 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    182 GitHub stars~3.4k tokensUpdated 13 days ago
    Auto-check passed

Works with

Questions about Quark Onnx Subgraph Partitioner

What does Quark Onnx Subgraph Partitioner do?

Partition an ONNX model graph into named functional subgraphs and emit a subgraphpartition.json file. Quark Onnx Subgraph Partitioner is an agent skill from amd/Quark.json file.

When should I use Quark Onnx Subgraph Partitioner?

Quark Onnx Subgraph Partitioner fits situations like: the user wants to understand a models high-level structure; document its architectural blocks; prepare a partition for downstream workflows (mixed-precision quantization; layer-wise profiling.

How do I install Quark Onnx Subgraph Partitioner in Claude Code?

Run `npx skills add amd/Quark --skill quark-onnx-subgraph-partitioner -a claude-code`. Or copy the skill folder (skills/_legacy_impl/l1-atomic/onnx/quark-onnx-subgraph-partitioner in amd/Quark) into .claude/skills/quark-onnx-subgraph-partitioner in your project. Claude Code loads it when a task matches its description.

How do I install Quark Onnx Subgraph Partitioner in Codex?

Run `npx skills add amd/Quark --skill quark-onnx-subgraph-partitioner -a codex`. Or copy the skill folder (skills/_legacy_impl/l1-atomic/onnx/quark-onnx-subgraph-partitioner in amd/Quark) into .agents/skills/quark-onnx-subgraph-partitioner in your project. Codex loads it when a task matches its description.

Can I use Quark Onnx Subgraph Partitioner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-onnx-subgraph-partitioner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-onnx-subgraph-partitioner, .gemini/skills/quark-onnx-subgraph-partitioner, .github/skills/quark-onnx-subgraph-partitioner and .opencode/skills/quark-onnx-subgraph-partitioner in your project.

What does Quark Onnx Subgraph Partitioner need to run?

Going by SKILL.md and its folder, Quark Onnx Subgraph Partitioner needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: Python 3.

Does Quark Onnx Subgraph Partitioner access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Quark Onnx Subgraph Partitioner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Onnx Subgraph Partitioner use?

Quark Onnx Subgraph Partitioner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Onnx Subgraph Partitioner use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Onnx Subgraph Partitioner?

Skills that share tags, products or a category with Quark Onnx Subgraph Partitioner: Aipc Toolkit (qualcomm/qai-appbuilder, 247 stars), Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Onboard Jetpack5 Inference Backends (EGalahad/sim2real, 146 stars) and Model Builder (qualcomm/qai-appbuilder, 247 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Onnx Subgraph Partitioner?

amd (a GitHub organization) maintains it in amd/Quark, which has 182 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.