SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.
$ npx skills add amd/Quark --skill quark-torch-llm-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/Quark quark-torch-llm-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval .claude/skills/quark-torch-llm-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "quark-torch-llm-eval" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval into .claude/skills/quark-torch-llm-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/Quark --skill quark-torch-llm-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/Quark quark-torch-llm-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval .agents/skills/quark-torch-llm-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "quark-torch-llm-eval" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval into .agents/skills/quark-torch-llm-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-llm-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/Quark quark-torch-llm-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval .cursor/skills/quark-torch-llm-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "quark-torch-llm-eval" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval into .cursor/skills/quark-torch-llm-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/Quark.git --path .claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/Quark --skill quark-torch-llm-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/Quark quark-torch-llm-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval .gemini/skills/quark-torch-llm-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "quark-torch-llm-eval" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval into .gemini/skills/quark-torch-llm-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/Quark quark-torch-llm-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/Quark --skill quark-torch-llm-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval .github/skills/quark-torch-llm-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-llm-eval" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval into .github/skills/quark-torch-llm-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/Quark --skill quark-torch-llm-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/Quark quark-torch-llm-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/Quark.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval .opencode/skills/quark-torch-llm-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "quark-torch-llm-eval" agent skill from https://github.com/amd/Quark/tree/release%2F0.13/.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval into .opencode/skills/quark-torch-llm-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quark-torch-llm-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
quark-torch-llm-evalEnd-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.
Quark Torch LLM Eval is an agent skill from amd/Quark. End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. Use when the user wants to evaluate, benchmark, or compare an LLM's accuracy. Trigger for "evaluate this model", "run gsm8k/mmlu/mmlupro/aime/gpqa/hellaswag/arc", "test accuracy", "measure perplexity", "compare quantized model accuracy", "does this mxfp4 model lose accuracy". For evaluating Quark Agent Skills themselves, use quark-torch-eval-runner instead.
Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `backends.md`, `eval-frameworks.md` and `model-inspection.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving, Web search and Model hubs and datasets. It works with vLLM, SGLang, Perplexity and Docker. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 313cb0b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockercurlpythonpython3pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.corecipes.vllm.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Quark Torch LLM Eval loads about 6.2k tokens when it runs. Until then it costs about 132 tokens; SKILL.md has 2,923 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/Quark at commit 313cb0b, republished under its MIT licence (© amd). 2,923 words, ~6,171 tokens.
.claude/skills/quark-torch-llm-eval/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.End-to-end LLM accuracy evaluation on AMD ROCm — from Docker container to final accuracy table.
Trigger this skill when the user asks to:
| Parameter | Required | Description |
|---|---|---|
model_path or hf_id | Yes | Local path to model weights, or HuggingFace model ID |
benchmark | No | Benchmark to run (default: gsm8k) |
backend | No | Serving backend: vLLM (default), SGLang, or ATOM |
image | No | Docker image (default: vllm/vllm-openai-rocm:latest; overridden by model card / Recipes / user). Ignored when runtime=host. |
runtime | No | docker (default) or host. Auto-detected in Phase 0.5: docker when the daemon is reachable, else host when a native vllm / sglang is importable. ATOM is docker-only. Stop and report if neither path is available. |
$EVAL_STATE_DIR/eval_report.md — see templates.md §2 for the full skeleton, formatting rules, and the example used as the canonical shape.
HIP_VISIBLE_DEVICES. Do not preemptively stack flags.troubleshooting.md is a diagnostic guide, not a prevention checklist. Launch first, consult on failure.After Phase 0 + 0.5 + 1 finish collecting all information, present a single confirmation summary to the user before any container or server operations begin. Use the template in templates.md §1 — it defines all required fields, command format, and display rules.
Wait for explicit user approval before executing any of these commands. This avoids piecemeal confirmations and lets the user review the full pipeline at once.
After approval, the following are pre-approved by the unified gate — execute without re-asking:
docker run / docker exec / docker pull of the planned container/imagevllm serve / python -m sglang.launch_server / python -m atom.entrypoints.openai_server command shown in the plan (with the planned GPU/TP/port)lm_eval / lighteval / evalscope command shown in the planpip install of frameworks listed in eval-frameworks.md §Installation — docker runtime only (inside the planned container). In host runtime, pip install is never pre-approved; see Phase 0.5 host-mode constraints.Each of the following still requires per-command confirmation, even after the unified gate:
config.json or backend source code (show full diff first)kill -9 / sudo / rm -rf / mv outside $EVAL_STATE_DIR and outputsvllm serve / eval command that deviates from the plan (different flags, different port, retry with new args)docker stop / docker rm (Phase 7 cleanup — always ask)Read-only probes (ls, cat, du, rocm-smi, docker ps, docker images, curl -sI, curl -fsSL, Read) do not need confirmation.
pkill -f <pattern> or ps | grep | awk | kill broad matchingkill -0 <pid> + ps -fp <pid> (check USER and CMD columns) before sending a signal — PIDs are recycledSIGTERM first, wait up to 30 s, then SIGKILL only if still aliveAll HuggingFace HTTP requests follow model-inspection.md §2 — curl -fsSL, ${HF_ENDPOINT:-https://huggingface.co}, token from ~/.hf_token then $HF_TOKEN. Read that section before the first HF request in any phase.
Prerequisite: user must provide model_path (local path) or hf_id (HuggingFace ID).
ls <model_path>/config.json) → proceedconfig.json and tokenizer_config.json from HF API (use the HuggingFace fetch rules above), then continue to Phase 0.5 → 1. Before Phase 2, ask: A "I have a local copy" or B "Download from HuggingFace"Once a local path is confirmed, read model-inspection.md and follow its rules (infer hf_id, parse config.json, detect model type, estimate size, compute TP candidates). Write findings to $EVAL_STATE_DIR/current-eval.yaml.
Before Phase 1, run a 60-second environment probe to fail fast:
which curl python3
# Probe both runtimes; do NOT fail if docker is missing — host mode covers that.
docker --version >/dev/null 2>&1 && docker info >/dev/null 2>&1 \
&& echo "docker_available=true" || echo "docker_available=false"
python3 -c "import vllm; print('vllm', vllm.__version__)" 2>/dev/null \
&& echo "host_vllm_available=true" || echo "host_vllm_available=false"
python3 -c "import sglang; print('sglang', sglang.__version__)" 2>/dev/null \
&& echo "host_sglang_available=true" || echo "host_sglang_available=false"Runtime selection (record runtime: in current-eval.yaml):
runtime explicitly (docker or host), honour it; validate the matching tool exists.docker_available=true, use docker (default — production path).backend=vLLM and host_vllm_available=true, or backend=SGLang and host_sglang_available=true, fall back to runtime=host and report the fallback in the unified confirmation.backend=ATOM and docker is missing, stop here regardless of host imports.Host-mode constraints (announce in the unified confirmation):
image is ignored; the host's currently-installed framework version is what runs.docker exec ${CONTAINER} <cmd> become <cmd> directly on the host (the "Host runtime variant" call-out under each phase below).pip install of frameworks is NOT performed — the host environment is used as-is. Missing a benchmark framework (e.g., lm_eval) → stop and ask the user to install it explicitly; do not pip install on the host without consent.Session path setup (first thing in this phase): generate a unique, per-session state directory to avoid collisions between concurrent evals.
bash shell state does NOT persist across separate Bash tool calls — EVAL_USER / EVAL_TS exported in one call are gone in the next. The path must therefore be resolved once and persisted to disk, not re-exported each phase.
EVAL_USER=$(whoami)
EVAL_TS=$(date +%Y%m%d_%H%M%S)
EVAL_STATE_DIR="/tmp/eval-state/${EVAL_USER}/${EVAL_TS}"
mkdir -p "${EVAL_STATE_DIR}"
echo "eval_state_dir: ${EVAL_STATE_DIR}" > "${EVAL_STATE_DIR}/current-eval.yaml"
echo "${EVAL_STATE_DIR}" > /tmp/eval-state/.last-session-path-${EVAL_USER}All subsequent state files live under $EVAL_STATE_DIR:
$EVAL_STATE_DIR/current-eval.yaml — phase progress (machine-readable YAML)$EVAL_STATE_DIR/server-runtime.md — server state$EVAL_STATE_DIR/vllm-server.log — server log (redirected here)$EVAL_STATE_DIR/eval-output.log — eval framework output$EVAL_STATE_DIR/monitor.pid — background monitor task IDs (Phase 3)Recovering $EVAL_STATE_DIR in later phases: every later Bash call starts with
EVAL_STATE_DIR=$(cat /tmp/eval-state/.last-session-path-$(whoami))or reads the value from the first line of the most recent current-eval.yaml. Do not call date +%Y%m%d_%H%M%S again — it produces a different timestamp and a different (empty) directory.
Record results in the env_preflight section of the state file (capture all probe outputs plus the chosen runtime). If neither docker run nor a native backend is usable per the Runtime selection rules above → stop and report, listing what is missing.
Do not ask the user any questions during Phase 0, 0.5, or 1. Collect all information autonomously, apply sensible defaults for anything missing, and present everything in a single unified confirmation at the end. The user makes one decision — approve or adjust.
Defaults for missing info:
image: no model card / recipe specifies one → default by backend: vLLM vllm/vllm-openai-rocm:latest; ATOM rocm/atom-dev:latest; SGLang lmsysorg/sglang:v0.5.11-rocm720-mi35xbenchmark: unspecified and no model card recommendation → default to gsm8kbackend: unspecified → default to vLLMImage availability check: run docker images --format '{{.Repository}}:{{.Tag}}' and check if the chosen image already exists locally. Report the result in the unified confirmation (e.g., "Image: vllm/vllm-openai-rocm:latest (already pulled)" or "Image: vllm/vllm-openai-rocm:latest (needs pull)").
Fetch model card and look for: recommended Docker image, serving / launch command, evaluation command, recommended benchmark.
Model card fetch: always use curl to get the raw README.md directly — returns raw markdown without needing HTML parsing or JS rendering. Apply the HuggingFace fetch rules above:
curl -fsSL "${HF_BASE}/<hf_id>/raw/main/README.md" \
${HF_TOKEN:+-H "Authorization: Bearer $HF_TOKEN"}Fetch strategy — per-item priority chain with two-stage model card lookup:
Five pieces of info are collected: image, launch_cmd, eval_cmd, eval_settings, and reference_scores. Each one independently walks the priority chain below. Finding one item at a given level does NOT skip that level for other items — only skip lower levels for the specific item already found.
eval_settings includes: num_fewshot, evaluation mode (base/chat, i.e. local-completions/local-chat-completions), and prompting strategy (CoT/direct). These settings determine how the eval command is constructed and are critical for reproducing published results.
Example: model card provides launch_cmd but not eval_cmd → launch_cmd is settled. For eval_cmd, continue down the chain until found.
Priority levels:
Quantized model's own model card — curl -sL https://huggingface.co/<hf_id>/raw/main/README.md (e.g., amd/MiniMax-M2.7-mxfp4). Highest trust: quant-specific commands verified by the quantized model's author.
Sibling quant variant's model card — if the quantized model has no HF page or no useful commands, search for sibling quant variants under the same org:
curl -fsSL "${HF_BASE}/api/models?search=<org>/<base-model-name>&author=<org>" \
${HF_TOKEN:+-H "Authorization: Bearer $HF_TOKEN"}If a variant with a different quant strategy is found (e.g., amd/Model-fp8 when evaluating amd/Model-mxfp4), fetch its model card. Commands from a sibling variant share the same architecture and only differ in quantization flags. Record launch_cmd_source: similar_quant_card and reference_model: <sibling_hf_id>.
Base model's model card — curl -sL https://huggingface.co/<base_hf_id>/raw/main/README.md (e.g., MiniMaxAI/MiniMax-M2.7). Commands need adaptation: replace model path, may need quantization-specific adjustments. Record launch_cmd_source: base_model_card.
vLLM Recipes — use base_hf_id (base model org/name, not quantized variant). Recipes pages are JS-rendered; try curl -sL https://recipes.vllm.ai/<org>/<base-model-name> and extract from __NEXT_DATA__ or HTML. If not accessible, skip to next level.
ATOM Recipes / SGLang Cookbook — backend-specific sources
Framework repos / InferenceX — eval commands only
Paper / technical report — for eval_cmd and eval_settings only. When levels 1–6 do not provide an explicit eval command or eval settings, extract settings from the model's arxiv paper. See eval-frameworks.md §Paper Eval Settings Extraction for details. Record eval_cmd_source: paper and eval_settings_source: paper.
Generic templates from eval-frameworks.md — last resort. Record launch_cmd_source: template.
Record each item and its source in the state file (e.g., launch_cmd_source: quant_model_card, eval_cmd_source: template).
Reference scores collection: while fetching model cards and papers in the priority chain above, also collect benchmark reference scores for the target benchmark(s). Walk the lookup chain in eval-frameworks.md §Accuracy Comparison — it covers model card, base model card, paper/technical report, Open LLM Leaderboard, and similar-scale fallback. Record the score, source, and eval setting (num_fewshot, CoT/direct, base/chat) for each benchmark. This enables the unified confirmation to show expected accuracy and flag setting mismatches before the eval runs.
Note: The local path may be a quantized variant (e.g., /shareddata/amd/Model-mxfp4), but Recipes uses the base model ID (base_hf_id). In Phase 3, replace the base model ID in the command with the local path.
Host runtime variant: if runtime=host (set in Phase 0.5), skip this phase entirely. Set CONTAINER="" (sentinel) in the state file so Phase 3+ branches detect host mode. Jump straight to Phase 3.
Docker runtime (default): Read backends.md for the full docker run template. Name the container eval-<hf_id_safe>-<username>-<date> (username from whoami).
Starting point — use the highest-priority source found in Phase 1:
vllm serve <model_path> --tensor-parallel-size <TP>python -m atom.entrypoints.openai_server --model <model_path> --kv_cache_dtype fp8 -tp <TP>python -m sglang.launch_server --model-path <model_path> --tp <TP>Model card command rule: when a model card provides a launch command, preserve ALL flags from that command. The allowed/forbidden modifications table below only applies to changes YOU make on top of the model card command — do not strip flags that the model card author included (e.g., --enforce-eager, --reasoning-parser, --mm-encoder-tp-mode, env vars like VLLM_ROCM_USE_AITER=1).
Adaptation rules: see backends.md §Cross-Backend Launch Adaptation Rules — applies to all three backends. Only the changes in that table are allowed on top of the source command; reactive flags (--enforce-eager etc.) come from troubleshooting.md only when an error matches.
TP priority: see model-inspection.md §5 — recipe TP takes precedence over auto-inferred.
ATOM-specific notes: ATOM uses --kv_cache_dtype fp8 by default for memory efficiency. ATOM auto-detects quantization configs from HuggingFace model configs (FP8, MXFP4, INT8, INT4). For MTP speculative decoding, add --method mtp --num-speculative-tokens 3. ATOM also supports running as a vLLM plugin — if atom is installed alongside vllm, vllm serve automatically picks up ATOM's optimized kernels.
Pre-launch checks (immediately before launching):
model-inspection.md §4 — build the rocm-smi → HIP index mapping via PCIe bus addresses. rocm-smi indices and HIP_VISIBLE_DEVICES indices are NOT the same.rocm-smi --showmeminfo vram for exact free VRAM, --showuse for GPU%, --showpids for processes. Identify idle GPUs by rocm-smi index, then convert to HIP indices using the mapping.lsof -i :<port> — if occupied, try 8081, 8082, etc. Add --port <port> to the launch command.Launch template — server log path is mandatory. Phase 4/5/7 monitoring all assume the log is reachable at ${EVAL_STATE_DIR}/vllm-server.log; do not deviate.
Both runtimes write the server log directly to ${EVAL_STATE_DIR}/vllm-server.log — for docker, this works because the docker run template in backends.md bind-mounts ${EVAL_STATE_DIR} into the container at the same path.
Docker runtime (default):
docker exec -d ${CONTAINER} bash -c \
"HIP_VISIBLE_DEVICES=${GPUS} ${LAUNCH_CMD} > ${EVAL_STATE_DIR}/vllm-server.log 2>&1"Host runtime variant (when runtime=host):
# Launch directly on the host; PID captured into the state dir for later signals.
( HIP_VISIBLE_DEVICES=${GPUS} ${LAUNCH_CMD} > "${EVAL_STATE_DIR}/vllm-server.log" 2>&1 ) &
echo "$!" > "${EVAL_STATE_DIR}/server.pid"In host mode, every later step that says docker exec ${CONTAINER} <cmd> becomes just <cmd> on the host; every "tail the in-container log" becomes tail -F "${EVAL_STATE_DIR}/vllm-server.log" (same path works in docker mode too, since the log file is bind-mounted).
Write $EVAL_STATE_DIR/server-runtime.md immediately after launch (PID, port, container, image, full launch command).
Readiness wait (up to 300s): use the poll template in backends.md §Readiness Poll Template — polls health every 5s with process liveness check, breaks immediately on process death instead of waiting 300s.
Background server monitor: after the server is ready, launch the monitor via Bash with run_in_background: true and persist the returned task_id:
# Bash, run_in_background: true. Capture task_id, then:
echo "server_monitor=${task_id}" >> "${EVAL_STATE_DIR}/monitor.pid"
# Both runtimes — log is bind-mounted into the container, so tail from host either way:
tail -F "${EVAL_STATE_DIR}/vllm-server.log" \
| grep -E --line-buffered "Avg prompt throughput|Avg generation throughput|ERROR|error|OOM|Killed|CUDA|torch.OutOfMemoryError|engine is dead"TaskStop every ID in monitor.pid at every exit path: Phase 5 retry (before each new launch), Phase 6 early-exit, Phase 7 cleanup. Skipping this leaks a background tail -F per session.
Follow backends.md §Smoke Test Templates — verify model name via /v1/models, then send a simple inference request (chat or base). Both HTTP 200 and reasonable content are required. Either gate fails → enter Phase 5.
Entered on Phase 3 launch failure or Phase 4 gate failure. Mode comes from launch_cmd_source in current-eval.yaml.
Reproduce mode (launch_cmd_source ∈ {quant_model_card, similar_quant_card}) — the command is verified for this arch+quant combo, so do not auto-fix:
TaskStop every ID in $EVAL_STATE_DIR/monitor.pidtroubleshooting.md, (c) suggested fixfall through to bringup (re-enter Phase 5 in Bringup mode) or abort (stop, offer Phase 7 cleanup). Do not silently fall through. Do not retry without instruction.Bringup mode (launch_cmd_source ∈ {base_model_card, recipes, template}) — run the fix cycle below.
Three rules (bringup only): one fix at a time (no flag stacking); rerun the exact failing step (Phase 3 or 4, no skipping); max 3 attempts then stop and report the full log.
Fix cycle:
troubleshooting.mdconfig.json → show full diff, wait for user approvalgithub.com/vllm-project/vllm/issues; same approval rule for any source mod{attempt, error, fix, result} in current-eval.yamlThree steps, in order. Do not present the final report until all three are done.
Step 1 — Run eval. Read eval-frameworks.md for command templates, task type classification, and num_fewshot lookup. Format the raw eval output as shown in the Outputs section above.
Step 2 — Accuracy comparison. Use the reference_scores already collected in Phase 1. If Phase 1 found a reference, use it directly. If not, do a fallback lookup per eval-frameworks.md §Accuracy Comparison (paper → Open LLM Leaderboard → similar-scale model). Build the comparison table with the eval setting from the reference source. If the reference's eval setting differs from what was actually used, note the difference explicitly. No benchmark may be silently omitted — if no reference exists, write "N/A" with reason.
Step 3 — Anomaly detection. Flag any delta > 5% absolute. See eval-frameworks.md §Anomaly detection for the diagnostic table.
After the evaluation report is produced, ask the user how to handle the running resources:
$EVAL_STATE_DIR/monitor.pid, call TaskStop on the recorded task_id. Truncate the file after.docker exec ${CONTAINER} pkill -f "vllm serve|sglang.launch_server|atom.entrypoints"), then verify the port is released via lsof -i :${PORT}.${EVAL_STATE_DIR}/server.pid following the "Never kill other users' processes" rule above (L72-76) — verify the PID is alive and still our backend, then SIGTERM → wait → SIGKILL only if needed. Verify port with lsof -i :${PORT}.docker stop + docker rm. Skip in host runtime.source_edits is non-empty in $EVAL_STATE_DIR/current-eval.yaml, offer to restore each .bak.<timestamp> file: cp <file>.bak.<timestamp> <file>Always ask — never auto-kill. After cleanup (or skip), update $EVAL_STATE_DIR/server-runtime.md.
$EVAL_STATE_DIR is set in Phase 0.5 to /tmp/eval-state/<user>/<timestamp>/ — unique per session and user. Every later phase reads the path from /tmp/eval-state/.last-session-path-$(whoami) rather than re-deriving from whoami + timestamp.
Primary: $EVAL_STATE_DIR/current-eval.yaml — updated after each Phase.
Server: $EVAL_STATE_DIR/server-runtime.md — written at Phase 3 launch, kept current.
Fallback: if /tmp/eval-state/ is not writable, use ./.eval-state/<timestamp>/ in the current working directory. Phase 0.5 preflight probes this and writes the resolved path as the first field of current-eval.yaml and into the .last-session-path-<user> pointer.
| File | When to read |
|---|---|
model-inspection.md | Phase 0, for config parsing / GPU probing / TP math; §2 for all HF HTTP requests |
backends.md | Phase 2-3, for docker setup + serving commands (all backends) |
troubleshooting.md | Phase 5, on launch failure or smoke test failure |
eval-frameworks.md | Phase 6, for eval commands, accuracy comparison lookup, and anomaly detection |
templates.md | Phase 1 (pre-run confirmation) and Phase 6 (assembling eval_report.md) |
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in .claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval of amd/Quark.
Open the folder on GitHubat commit 313cb0b
Quark Torch LLM Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Quark Torch LLM Eval this skillamd/Quark | 181 | — | ~6.2k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Hyperloom SetupAMD-AGI/Hyperloom | 218 | — | ~7.2k | Automated safety check: Notes | Custom licence | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Debug InferenceNVIDIA/OpenShell | 16k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
AMD-AGI/Hyperloom
Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
NVIDIA/OpenShell
Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.
vllm-project/vllm-ascend
Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.
amd/Quark
Author or restructure a Quark Agent Skill so it conforms to this project's template, contracts, and layer rules.
amd/Quark
Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.
amd/Quark
Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/Quark
Install or verify the AMD Quark package and its dependencies.
amd/Quark
L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…
Works with
Categories
End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. Quark Torch LLM Eval is an agent skill from amd/Quark. End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.
Quark Torch LLM Eval fits situations like: the user wants to evaluate; compare an LLMs accuracy; evaluate this model; run gsm8k/mmlu/mmlupro/aime/gpqa/hellaswag/arc.
Run `npx skills add amd/Quark --skill quark-torch-llm-eval -a claude-code`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval in amd/Quark) into .claude/skills/quark-torch-llm-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/Quark --skill quark-torch-llm-eval -a codex`. Or copy the skill folder (.claude/skills-impl/l1-atomic/torch/quark-torch-llm-eval in amd/Quark) into .agents/skills/quark-torch-llm-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/Quark --skill quark-torch-llm-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quark-torch-llm-eval, .gemini/skills/quark-torch-llm-eval, .github/skills/quark-torch-llm-eval and .opencode/skills/quark-torch-llm-eval in your project.
Going by SKILL.md and its folder, Quark Torch LLM Eval needs the command-line tools its instructions call (docker, curl, python, python3 and pip) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker.
SKILL.md names 2 domains. In commands or code: huggingface.co and recipes.vllm.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Quark Torch LLM Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Quark Torch LLM Eval: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Hyperloom Setup (AMD-AGI/Hyperloom, 218 stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/Quark, which has 181 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on September 28, 2026.
Source: amd/Quark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.