Agent skill

Quark Onnx Eval Runner

by amd in amd/Quark

Manually verify that the Quark ONNX skill family behaves correctly across the four contract categories (routing, planning, artifact, recovery).

MITAuto-check passedTesting & QA

Install Quark Onnx Eval Runner

skills CLI
$ npx skills add amd/Quark --skill quark-onnx-eval-runner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/Quark quark-onnx-eval-runner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/meta/onnx/quark-onnx-eval-runner .claude/skills/quark-onnx-eval-runner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quark-onnx-eval-runner
GitHub stars
181
Token cost
~2.8k tokens
SKILL.md length
1,258 words
Files
1
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Manually verify that the Quark ONNX skill family behaves correctly across the four contract categories (routing, planning, artifact, recovery).

  • Works in 4 steps: Routing → Planning → Artifact → …
  • Maintainers need to confirm that ONNX routing
  • SKILL.md covers Purpose, Inputs, Outputs: validation_report.md and Manual Verification Protocol, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Quark Onnx Eval Runner is an agent skill from amd/Quark. Manually verify that the Quark ONNX skill family behaves correctly across the four contract categories (routing, planning, artifact, recovery). Use when maintainers need to confirm that ONNX routing, planning, artifact generation, or error recovery skills still work as expected. Trigger for "verify the ONNX skills", "smoke-test ONNX routing", "check ONNX skill behavior", or before tagging a release that touches quark-onnx- skills. This is a governance tool for skill maintainers, not for end users running ONNX…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports. It works with ONNX. The licence is MIT.

When your agent uses it

  • Maintainers need to confirm that ONNX routing
  • Artifact generation
  • Error recovery skills still work as expected
  • Verify the ONNX skills

Example prompts

  • “verify the ONNX skills”
  • “smoke-test ONNX routing”
  • “check ONNX skill behavior”
  • “/quark-onnx-eval-runner”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Routing
  2. Planning
  3. Artifact
  4. Recovery

What it can do on your machine

Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quark Onnx Eval Runner loads about 2.8k tokens when it runs. Until then it costs about 158 tokens; SKILL.md has 1,258 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~158
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 1,258 words, ~2,817 tokens.

Download SKILL.mdSave it as .claude/skills/quark-onnx-eval-runner/SKILL.md (or your agent's skills folder).
name
quark-onnx-eval-runner
description
Manually verify that the Quark ONNX skill family behaves correctly across the four contract categories (routing, planning, artifact, recovery). Use when maintainers need to confirm that ONNX routing, planning, artifact generation, or error recovery skills still work as expected. Trigger for "verify the ONNX skills", "smoke-test ONNX routing", "check ONNX skill behavior", or before tagging a release that touches `quark-onnx-*` skills. This is a governance tool for skill maintainers, not for end users running ONNX model accuracy evaluation — for the latter use the upstream Quark ONNX evaluation tooling.
layer
meta
backend
onnx
primary_artifact
validation_report.md
source_knowledge
examples/onnx/yolo_quantization/quantize_yolo.py, examples/onnx/model_support.md, docs/source/onnx/basic_usage_onnx.rst…

quark-onnx-eval-runner

Purpose

Walk a maintainer through manual verification of the ONNX skill family across the four contract categories: routing, planning, artifact, and recovery. Run this after modifying any quark-onnx-* skill, after a Quark ONNX upgrade, or before tagging a release.

Inputs

  • The current ONNX skill files under .claude/skills-impl/{l1-atomic,l2-workflows,l3-recipes}/onnx/
  • The entry stubs under .claude/skills/quark-onnx-*
  • The contract schemas under .claude/skills-impl/shared/contracts/
  • The example prompts under examples/agent_skills/prompts/ (add ONNX-specific cases as the catalog grows)

Outputs: validation_report.md

A markdown report recording per-category pass/fail and concrete evidence for each finding.

Schema: validation_report.schema.json

markdown
# ONNX Skill Verification Report

## Summary
| Category | Cases Run | Pass | Fail |
|----------|-----------|------|------|
| routing  | N         | N    | 0    |
| planning | N         | N    | 0    |
| artifact | N         | N    | 0    |
| recovery | N         | N    | 0    |

## Failures
### <category> / <case name>
- **Expected**: ...
- **Got**: ...
- **Impact**: ...
- **Fix**: ...

Manual Verification Protocol

For each category below, run at least one case and record the result in the report. As ONNX prompts are not yet enumerated in examples/agent_skills/prompts/, the cases below double as the seed catalog — add more as the ONNX skill set grows.

1. Routing

Goal: verify that quark-onnx-router (and Claude's auto-routing via the descriptions in .claude/skills/quark-onnx-*) maps natural-language ONNX goals to the correct downstream skill and never silently routes ONNX requests through a torch skill.

Manual procedure:

  1. Pick a user-style prompt that names a .onnx artifact or ONNX-specific vocabulary.

  2. In a fresh Claude Code session at the Quark repo root, paste the prompt.

  3. Observe which skill Claude invokes first.

  4. Compare against the expected target skill. Examples of expected mappings:

    • "Quantize my ./models/yolov8n.onnx to XINT8 for AMD NPU CNN" → quark-onnx-ptq (which loads quark-onnx-ptq-workflow)
    • "Run AutoSearchPro on this .onnx with the XINT8_SEARCH preset" → quark-onnx-autosearch-pro
    • "Analyze this .onnx — what opset is it, is it NPU-compatible, is it already QDQ?" → quark-onnx-model-intake
    • "Validate my quantized model.onnx — did QDQ insertion happen, are the non-quantized initializers byte-identical?" → quark-onnx-result-validator
    • "onnxruntime-gpu import fails, CUDAExecutionProvider not in providers list" → quark-onnx-install (or quark-onnx-debug if the user already attempted install)
    • "quantize_static failed with custom-op library load failure for BFPQuantizeDequantize" → quark-onnx-debug
    • "Is onnxruntime-rocm installed correctly? Show me the install matrix" → quark-onnx-install
  5. Cross-backend guard: also run one negative prompt that mentions a .onnx path and confirm Claude does not route to quark-torch-* (e.g., "quantize ./models/foo.onnx with FP8" must not land on quark-torch-ptq).

Pass criteria: the first skill invoked matches the expected target, and no ONNX prompt is routed to a torch skill.

2. Planning

Goal: verify that quark-onnx-quant-plan produces internally consistent plans for typical ONNX inputs and that the deployment-target gates are respected.

Manual procedure:

  1. Construct (or take from a prior session) a model_analysis.json produced by quark-onnx-model-intake for a representative model (e.g., YOLOv8n exported at opset 17, Conv-heavy, 6.2 MB inline).
  2. Hand it to quark-onnx-quant-plan with a target preset (e.g., XINT8) and a deployment target (e.g., AMD NPU CNN).
  3. Inspect the produced quant_plan.json for:
    • preset matches the requested preset
    • activation_spec and weight_spec are consistent with the preset (e.g., both XInt8Spec for XINT8)
    • EnableNPUCnn=True is set when the target is AMD NPU CNN
    • use_external_data_format is True iff the model is >2 GB
    • algo_config is a non-empty list when CLE or AdaRound was requested or recommended
    • exclude is a list (may be empty) and never contains an op the plan also quantizes
    • requires_confirmation is set when the plan deviates from preset defaults
    • Negative gate: the plan refuses incompatible combos (e.g., BFP16 + AMD NPU CNN) rather than silently downgrading

Pass criteria: the plan validates against quant_plan.schema.json, contains no internal contradictions, and explicitly rejects unsupported deployment-target / preset combinations.

3. Artifact

Goal: verify that ONNX workflow output artifacts conform to their JSON schemas and that the generated standalone script + manifest produced by quark-onnx-ptq-workflow agree with each other.

Manual procedure:

  1. Take the quant_plan.json from the planning case.
  2. Run quark-onnx-ptq-workflow to produce a run_manifest.yaml and the standalone <name>_ptq.py script in the user's working directory.
  3. Validate the manifest against .claude/skills-impl/shared/contracts/run_manifest.schema.json (use any JSON-schema validator, e.g., the jsonschema Python package).
  4. Spot-check that:
    • All required fields are present in the manifest
    • The manifest's command references the generated script path and uses python3
    • The manifest's resolved QConfig matches the plan's preset / algo_config / EnableNPUCnn / use_external_data_format
    • The generated script imports only from quark.onnx and the standard ORT calibration API (no editing of upstream examples/onnx/ or quark/onnx/ files)
    • Output paths reflect any .onnx_data sidecar when external data is enabled

Pass criteria: schema validation passes; the manifest's resolved config matches the plan; the generated script is self-contained in the user's working directory.

Show full SKILL.md (565 more words)Show less
4. Recovery

Goal: verify that quark-onnx-debug correctly diagnoses known ONNX-side error patterns and that handoffs to quark-onnx-install happen for runtime/provider issues.

Manual procedure:

  1. Pick a known ONNX error scenario. Examples:
    • RuntimeError: CUDAExecutionProvider not in available providers after installing onnxruntime (CPU build) instead of onnxruntime-gpu.
    • Custom-op library load failure for BFPQuantizeDequantize or MXQuantizeDequantize (missing C++ build, ABI mismatch).
    • model.onnx >2 GB and the run fails with "external data not found" because use_external_data_format was not set or the sibling .onnx_data was not staged.
    • OOM during calibration on a vision model with num_calib_data=1000 and batch_size=4.
    • AdaRound divergence with default learning rate.
    • NPU CNN run fails because activation scales are not power-of-two.
  2. Present the error (full traceback) to quark-onnx-debug.
  3. Verify the diagnosis:
    • Root cause is correctly identified
    • A concrete fix command or config change is provided
    • For runtime/provider issues, the skill explicitly hands off to quark-onnx-install instead of silently swapping execution providers
    • For OOM, the skill suggests the documented ladder: reduce num_calib_data → drop batch_size to 1 → move calibration to CPU (OptimDevice="cpu")
    • For custom-op load failures, the skill names the expected library path / build step rather than recommending the user ignore the op

Pass criteria: the diagnosis names the actual root cause, suggests a fix that would actually work, and respects the "never silently fall back to CPU" rule from quark-onnx-ptq-workflow.

Rules

  • Run at least one case per category before any ONNX release — even small wording changes in preset names or custom-op names can shift routing or break generated scripts.
  • A failing case means the skill is broken, not the procedure — investigate the skill first. Only update the expected behavior if the skill change was intentional.
  • Report all results, not just failures. A clean run is positive evidence and worth recording.
  • Document new cases inline in the report. As the ONNX skill set grows (e.g. new deployment target, new preset, new AutoSearchPro preset), add cases that cover the new triggers.
  • Cross-backend isolation is part of the protocol. Every routing case must include a guard that confirms ONNX requests stay on quark-onnx-* skills.

Interaction Flow

  1. Select scope: which categories to verify — all four, or a subset affected by a recent change?
  2. Pick or write cases: start from the examples above; add ONNX prompt files under examples/agent_skills/prompts/ as the catalog grows.
  3. Run each case manually following the protocol above in a fresh Claude Code session.
  4. Record results in validation_report.md using the template.
  5. Hand off failures to quark-onnx-skill-sync if they look like upstream drift, to quark-onnx-doc-drift-check if they look like stale user-facing facts, or directly to the affected skill's owner if it's a content bug.

Recovery

  • If a case can't run because a prerequisite artifact is missing (e.g., no model_analysis.json for the planning case), report the missing producer skill and stop — do not fabricate the input.
  • If the ONNX custom-op binaries are not built locally, the recovery case for the custom-op load failure may produce a real failure rather than a simulated one — note this distinction in the report so future runs don't confuse genuine environment gaps with skill bugs.
  • If Claude Code itself is misbehaving (ONNX skill not discovered, stub not loading), that's an infrastructure issue separate from skill quality — record it distinctly in the report.
  • If an ONNX prompt routes to a torch skill (or vice versa), this is a routing-category failure and must block the release — cross-backend mis-routing produces wrong artifacts silently.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills-impl/meta/onnx/quark-onnx-eval-runner of amd/Quark.

Open the folder on GitHubat commit 313cb0b

Compare with similar skills

Quark Onnx Eval Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quark Onnx Eval Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quark Onnx Eval Runner this skillamd/Quark181—~2.8kAutomated safety check: PassMIT
Dogfood Exploratory QAvercel-labs/agent-browser44k8 repos~2.7kAutomated safety check: PassApache-2.0
Codex Plugin QAcode-yeongyu/oh-my-openagent70k1 repos~1.9kAutomated safety check: PassCustom licence
DeerFlow Smoke Testbytedance/deer-flow83k—~2.5kAutomated safety check: NotesMIT
CodexBar Live QAsteipete/CodexBar22k—~1.2kAutomated safety check: PassMIT
Diagnose Playwright Failure as Product Bugappsmithorg/appsmith41k—~1.5kAutomated safety check: PassApache-2.0

Similar skills

  • Dogfood Exploratory QA

    vercel-labs/agent-browser

    Official

    Explores a web app with the agent-browser CLI to find bugs and UX problems, then writes a report with screenshots, repro videos and step-by-step reproduction for each issue.

    44k GitHub starsUsed in 8 repos~2.7k tokens
    Testing & QAAuto-check passed
  • Codex Plugin QA

    code-yeongyu/oh-my-openagent

    Tests the omo Codex plugin in an isolated CODEX_HOME with a local mock model, proving hooks fired through app-server notifications without touching ~/.codex.

    70k GitHub starsUsed in 1 repo~1.9k tokens
    Testing & QAAuto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    83k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • CodexBar Live QA

    steipete/CodexBar

    Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

    22k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Investigates a stubbornly failing Playwright test as a possible product bug, using error output, screenshots, traces and server code, and writes a structured bug report.

    41k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check passed
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed

More from amd/Quark

All 37 skills in this repo
  • Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.

    181 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    181 GitHub stars~3k tokensUpdated 10 days ago
    Auto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    181 GitHub stars~1.8k tokensUpdated 10 days ago
    Auto-check: notes
  • L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

    181 GitHub stars~3.4k tokensUpdated 10 days ago
    Auto-check passed

Works with

Categories

Questions about Quark Onnx Eval Runner

What does Quark Onnx Eval Runner do?

Manually verify that the Quark ONNX skill family behaves correctly across the four contract categories (routing, planning, artifact, recovery). Quark Onnx Eval Runner is an agent skill from amd/Quark. Manually verify that the Quark ONNX skill family behaves correctly across the four contract categories (routing, planning, artifact, recovery).

When should I use Quark Onnx Eval Runner?

Quark Onnx Eval Runner fits situations like: maintainers need to confirm that ONNX routing; artifact generation; error recovery skills still work as expected; verify the ONNX skills.

How do I install Quark Onnx Eval Runner in Claude Code?

Run `npx skills add amd/Quark --skill quark-onnx-eval-runner -a claude-code`. Or copy the skill folder (.claude/skills-impl/meta/onnx/quark-onnx-eval-runner in amd/Quark) into .claude/skills/quark-onnx-eval-runner in your project. Claude Code loads it when a task matches its description.

How do I install Quark Onnx Eval Runner in Codex?

Run `npx skills add amd/Quark --skill quark-onnx-eval-runner -a codex`. Or copy the skill folder (.claude/skills-impl/meta/onnx/quark-onnx-eval-runner in amd/Quark) into .agents/skills/quark-onnx-eval-runner in your project. Codex loads it when a task matches its description.

Can I use Quark Onnx Eval Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-onnx-eval-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-onnx-eval-runner, .gemini/skills/quark-onnx-eval-runner, .github/skills/quark-onnx-eval-runner and .opencode/skills/quark-onnx-eval-runner in your project.

What does Quark Onnx Eval Runner need to run?

SKILL.md names no scripts, command-line tools or credentials: Quark Onnx Eval Runner is instructions for the agent only. Our summary lists: Python 3.

Does Quark Onnx Eval Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quark Onnx Eval Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quark Onnx Eval Runner use?

Quark Onnx Eval Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quark Onnx Eval Runner use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Quark Onnx Eval Runner?

Skills that share tags, products or a category with Quark Onnx Eval Runner: Dogfood Exploratory QA (vercel-labs/agent-browser, 44k stars), Codex Plugin QA (code-yeongyu/oh-my-openagent, 70k stars), DeerFlow Smoke Test (bytedance/deer-flow, 83k stars) and CodexBar Live QA (steipete/CodexBar, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quark Onnx Eval Runner?

amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.

Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.