Agent skill

Serving LLMs On Epyc

by amd in amd/skills

Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

MITAuto-check: notesAI & LLM Engineering

Install Serving LLMs On Epyc

skills CLI
$ npx skills add amd/skills --skill serving-llms-on-epyc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/skills serving-llms-on-epyc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/serving-llms-on-epyc .claude/skills/serving-llms-on-epyc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
serving-llms-on-epyc
GitHub stars
398
Token cost
~5.7k tokens
SKILL.md length
2,431 words
Files
11 (incl. scripts)
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

  • Works in 8 steps: Detect the CPU → Validate the runtime and environment → Resolve and validate the model → …
  • Zentorch serving
  • SKILL.md covers Data file, Step 1: Detect the CPU, Step 2: Validate the runtime… and Step 3: Resolve and validate…, plus 7 more sections
  • Runs Python scripts from its folder; calls curl and python3; needs HF_TOKEN

What it does

Serving LLMs On Epyc is an agent skill from amd/skills. Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. Use for "vLLM on CPU", "zentorch serving", or an EPYC CPU endpoint, including on a host that also has AMD Instinct GPUs. Detects the EPYC generation, validates the runtime, checks model support and RAM fit, sizes threads/KV/NUMA, confirms the plan, launches, and verifies the endpoint. Runs one instance on one socket and its memory. Reports and stops on failure; does not retry or debug. Use…

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts (for example `data/epyc.json`, `evals/evals.json` and `evals/machine.yml`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Containers. It works with vLLM and Docker. The repository describes itself as: Official AMD catalog of AI agent skills. Empower your AI agents with AMD's optimized SW stack. The licence is MIT.

When your agent uses it

  • Zentorch serving
  • An EPYC CPU endpoint
  • Including on a host that also has AMD Instinct GPUs

Example prompts

  • “vLLM on CPU”
  • “zentorch serving”
  • “Use the serving-llms-on-epyc skill to serve an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda”
  • “/serving-llms-on-epyc”

Requirements

  • Python 3
  • Docker
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Detect the CPU
  2. Validate the runtime and environment
  3. Resolve and validate the model
  4. Check it fits host RAM
  5. Size the CPU runtime from the hardware
  6. Confirm the plan, then launch (container-first)
  7. Poll until up and responsive
  8. On success, hand over the endpoint

What it can do on your machine

Read from SKILL.md and the folder at commit 6c92b41. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • curl
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Serving LLMs On Epyc loads about 5.7k tokens when it runs. Until then it costs about 162 tokens; SKILL.md has 2,431 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~162
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from amd/skills at commit 6c92b41, republished under its MIT licence (© amd). 2,431 words, ~5,681 tokens.

Download SKILL.mdSave it as .claude/skills/serving-llms-on-epyc/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
serving-llms-on-epyc
description
Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. Use for "vLLM on CPU", "zentorch serving", or an EPYC CPU endpoint, including on a host that also has AMD Instinct GPUs. Detects the EPYC generation, validates the runtime, checks model support and RAM fit, sizes threads/KV/NUMA, confirms the plan, launches, and verifies the endpoint. Runs one instance on one socket and its memory. Reports and stops on failure; does not retry or debug. Use serving-llms-on-instinct when the endpoint should run on a GPU. Excludes multi-node, EPYC 4000, and pre-Zen4 EPYC without AVX-512.
allowed-tools
Bash, Read

Serving LLMs on AMD EPYC™ (vLLM + zentorch, CPU)

Bring up a single vLLM OpenAI endpoint on an AMD EPYC™ host with the zentorch CPU backend, sized to the hardware. Container-first (Docker or Podman); conda/host is the fallback. An installed AMD Instinct GPU does not disqualify the host: select this skill when the endpoint itself should run on the EPYC CPU.

This is single-socket serving: one instance pinned to one socket and its memory (vLLM scales poorly across sockets, so we do not span them). On a dual-socket host it runs on a single socket; the multi-socket answer is multiple instances (one per socket), which is out of scope for this single-instance recipe.

Hard rule for this skill: on any failure, report the cause + logs and STOP. Do not retry, do not debug. (Debugging is a separate workflow.)

The agent does the serve flow itself -- pull, configure, launch, poll -- using the runtime validate.py reports. Never hand the user per-serve commands. Like serving-llms-on-instinct, an accessible container runtime is a one-time prerequisite: if validate.py finds none, report its one-time fix (make docker accessible / install podman / provide a conda env) and stop. Do not attempt sudo or privilege escalation.

Data file

Read data/epyc.json directly. It holds the container image, mandatory CPU run flags, supported precision, the model-support policy, the default model, and the verified throughput-flag gotcha. Its vllm_version and image tag are one validated default stack; keep them aligned and do not hardcode either from memory.

Step 1: Detect the CPU

bash
python3 scripts/detect.py            # add --host user@box for a remote host

Returns cpu_model, is_amd_epyc, epyc_generation (Naples/Rome/Milan/Genoa/Bergamo/Siena/Turin/Venice or EPYC 4004/4005), zen_arch, is_supported_epyc, avx512, logical_cores, physical_cores, sockets, numa_nodes, memory_gb.

Route from detect.py -- decide the serving path:

  • is_amd_epyc is false -> stop: this skill targets AMD EPYC. (Other x86 may work but is unsupported here.)
  • avx512 is false -> zentorch cannot run on this CPU (its bf16 path needs AVX-512 BF16, avx512_bf16, which only lands on Zen4+). This is a pre-Zen4 EPYC (Naples / Rome / Milan, 7000 series). Do not dead-end -- it is still an EPYC host: offer the stock vLLM CPU path (plain vLLM, no zentorch acceleration -- slower, but verified working on EPYC 7763/Milan). Proceed only on the user's explicit OK, launching the official stock vLLM CPU image (Step 6 "Stock vLLM" variant); if the user declines, stop.
  • is_supported_epyc is false but avx512 is true (e.g. Bergamo / Siena / EPYC 4004/4005) -> the zentorch path is not validated for this generation. Get explicit confirmation to try zentorch unvalidated, or take the same stock vLLM offer.
  • else (9000 series -- Genoa/Turin/Venice -- with AVX-512 BF16) -> the validated zentorch path. Proceed.

validate.py (Step 2) reports zentorch_capable and sets requires_confirmation for the stock/unvalidated paths, so this routing is enforced there too.

Carry epyc_generation / avx512 through the later phases -- e.g. Venice packs up to 256 cores/socket, which the thread-binding in Step 5 sizes from.

Step 2: Validate the runtime and environment

bash
python3 scripts/validate.py --image <image from data/epyc.json> --generation <epyc_generation from detect> --avx512 <avx512 from detect>

This also hard-enforces the AVX-512 gate: on a CPU without AVX-512, validate.py returns a blocking error (ready: false) so the flow stops here regardless of the Step 1 prose -- no image pull, no launch. (It reads the local CPU itself if --avx512 is omitted, so the gate holds even if the value was not passed through.)

Returns ready, requires_confirmation, zentorch_capable (false -> zentorch can't run; use the stock vLLM variant in Step 6), runtime (docker, podman, or null), runtime_detail, conda_path_available, stack, compatibility, hf_cache (resolved HF cache mount -- use hf_cache.mount at launch), ram_gb, and errors/warnings/advisories. If requires_confirmation is set (stock/unvalidated path), surface that and get the user's OK before launching. Pick the path:

  • runtime is docker or podman -> container path (Step 6), used verbatim.
  • runtime null but conda_path_available: true -> conda/host path.
  • runtime null and no conda -> ready is false. Report the one-time onboarding fix (make docker accessible / install podman / conda env) and stop.

Do not proceed if ready is false.

Stack-compatibility gate. validate.py probes the selected runtime for its exact vllm/zentorch/torch versions and the active vLLM platform, then sets compatibility.status:

  • proceed -> the stack is the validated default (or a validated family on a Zen platform); continue.
  • blocked -> a stock CPU platform is active, so zentorch acceleration is not on (error). Report compatibility.message and stop.
  • confirmation_required (requires_confirmation: true) -> Venice on a vLLM other than the pinned default. This recipe has not been validated on Venice with that version. Surface compatibility.message, recommend the pinned vllm_version image from data/epyc.json, and stop for an explicit user go/no-go before launching. On the pinned default vLLM, Venice proceeds with no warning.

The gate only runs once the image is local. If validate.py reports the image is not pulled, pull it (or let Step 6 pull it) and re-run validate.py so the gate probes the real stack rather than only the tag.

Step 3: Resolve and validate the model

If the user named no model, use default_model from data/epyc.json (Qwen/Qwen3-0.6B -- ungated, tiny, fast first success). Otherwise use theirs.

Check that vLLM actually supports the model (do not blanket-block multimodal). Pass the vLLM version the model will actually run on: use stack.vllm from validate.py when it was probed (the conda env may differ from the pin), else the vllm_version from data/epyc.json.

bash
python3 scripts/check_model.py --model-id <model> --revision <rev or main> --vllm-version <stack.vllm from validate, else vllm_version from data/epyc.json>
  • Exit 0 = vLLM serves it as a generation endpoint, or support is undeterminable (gated/offline) -- proceed; launch confirms.
  • Exit 1 = stop: the architecture is not in vLLM's registry, it is a pooling/embedding/reranker (not a chat/completion endpoint), or it is a multimodal model with no usable chat template (launchable: false). Report the printed message and stop.

The result also carries the client endpoint the model supports:

  • primary_endpoint: "chat_completions" -- a usable chat template is present (chat_template.status: present); serve and hand off /v1/chat/completions.
  • primary_endpoint: "completions" -- no usable/auto-selectable template (absent/ambiguous/unknown); serve and hand off /v1/completions with a raw prompt. Chat can still be enabled by passing --chat-template <file> (or, for ambiguous, choosing one of chat_template.names); never invent one.
  • Carry primary_endpoint, supported_endpoints, and chat_template through to verification (Step 7) and the handoff (Step 8).
  • A multimodal model is allowed; a vLLM-supported multimodal arch may still hit a GPU-only kernel on CPU, which surfaces at load (the no-retry rule then applies).

Precision/dtype: native CPU dtypes are bf16 (default), fp16, fp32. Use bfloat16 unless the user asks otherwise.

For gated models (Llama, Gemma) HF_TOKEN must be set and the license accepted on HuggingFace; if not, stop and say so.

Step 4: Check it fits host RAM

RAM is the ceiling on CPU (weights + KV cache both live in RAM). Run on ONE line:

bash
python3 scripts/estimate_memory.py --model-id <model> --revision <rev or main> --ram-gb <memory_gb from detect> --max-model-len <4096 or user value> --num-prompts <1 or desired concurrency>

Exit 0 = fits, exit 1 = does not fit. If fit.fits is false: do not launch. Tell the user required_gb vs ram_gb and the printed fit.action -- reduce --max-model-len to fit.suggested_max_model_len and retry, or use a smaller model. --max-model-len and --num-prompts are the two knobs that move KV. Extra flag: --weight-gb N overrides weights if a model has no HF metadata (rare). KV cache is bf16-only on zentorch CPU (no fp8 KV).

Step 5: Size the CPU runtime from the hardware

bash
eval "$(python3 scripts/cpu_tune.py)"      # or --format json to inspect

A single instance runs on one socket, with its memory (vLLM scales poorly across sockets). cpu_tune.py exports VLLM_CPU_OMP_THREADS_BIND (the chosen socket's physical cores) and VLLM_CPU_KVCACHE_SPACE (sized from that socket's local RAM, not whole-system, so the KV pool stays on-socket). It does not set OMP_NUM_THREADS (vLLM derives it) or VLLM_CPU_NUM_OF_RESERVED_CPU (vLLM's own default).

Socket choice on a dual-socket host (load-aware): it samples per-socket CPU busy% (~0.5s) and prefers a free socket -- both free → socket 0; one free → that socket; both busy (≥ --busy-threshold, default 15%) → it warnings and proceeds on the least-busy socket. --socket N forces a choice. Single-socket hosts use socket 0.

For the chosen socket it also emits the memory-bound pin: container_cpuset (--cpuset-cpus=<cores> --cpuset-mems=<nodes>) for the container path, and conda_launch_prefix (numactl --cpunodebind/--membind, falling back to taskset CPU-only, or empty-with-note if neither tool exists) for conda. Surface warning to the user if set. On NPS2/NPS4 a socket spans multiple NUMA nodes; memory is bound across them and nps_note flags that finer binding could add performance.

Show full SKILL.md (1,138 more words)Show less

Step 6: Confirm the plan, then launch (container-first)

Before launching, present this summary and wait for the user to confirm -- do not launch unprompted. This is the human gate before anything runs:

FieldValue
Model / kind<model> -- text or multimodal (from check_model.py)
Backendzentorch-accelerated (Zen4+) or stock vLLM CPU / no zentorch (zentorch_capable:false -- unaccelerated; verified on Milan)
Pathcontainer (<runtime>, image from data/epyc.json) or conda/host
Precisionbfloat16 (or the user's choice)
Fitrequired <required_gb> GB vs <ram_gb> GB RAM
CPU sizingsocket <chosen_socket> (<socket_choice_reason>), bind <VLLM_CPU_OMP_THREADS_BIND>, KV <VLLM_CPU_KVCACHE_SPACE> GB (socket-local), mem bound to nodes <numa_nodes_on_socket>
HardwareEPYC <epyc_generation> (<zen_arch>), <physical_cores> cores, AVX-512 <avx512>
Port<port>

If cpu_tune.py returned a warning (e.g. all sockets busy), include it here so the user sees it before confirming.

Proceed only on a clear "go". If the user declines or wants changes (model, --max-model-len, port), stop and adjust -- do not launch.

Build the launch from data/epyc.json. The CLI is vllm serve <model>. Do not pass --device cpu on vLLM >= 0.20 -- the zentorch plugin auto-selects the CPU platform and vllm serve rejects the flag. Only add it if vllm serve --help lists it (older vLLM).

Pick a free port first. With --network=host the port is bound directly on the host, so a busy port is a hard failure (no remapping). Choose one that is free -- e.g. PORT=8000; while ss -ltn "sport = :$PORT" | grep -q LISTEN; do PORT=$((PORT+1)); done -- and use $PORT in the launch, health poll, and handover.

Mount the HF cache that validate.py resolved. Use hf_cache.mount from validate.py (it follows symlinks and flags NFS/root-squash) rather than a raw ~/.cache/huggingface, or the bind-mount can fail at container start on NFS homes.

Container path (runtime from validate.py). The agent runs these itself, including the pull. RT is the resolved runtime verbatim:

bash
RT="<runtime from validate.py: docker | podman>"
$RT rm -f vllm-epyc 2>/dev/null               # clear any leftover container from a prior run (name collision otherwise)
$RT pull <image from data/epyc.json>          # agent pulls; do not ask the user to
$RT run -d --name vllm-epyc \
  <run_flags from data/epyc.json>            # --ipc=host --network=host --cap-add=SYS_NICE (SYS_NICE = NUMA membind; NO --shm-size with --ipc=host)
  <hf_cache.mount from validate.py> \        # resolved real path, e.g. -v /scratch/you/hf:/root/.cache/huggingface
  <container_cpuset from cpu_tune>             # --cpuset-cpus=<cores> --cpuset-mems=<nodes>
  --env VLLM_CPU_OMP_THREADS_BIND="$VLLM_CPU_OMP_THREADS_BIND" \
  --env VLLM_CPU_KVCACHE_SPACE=$VLLM_CPU_KVCACHE_SPACE \
  --env HF_TOKEN=${HF_TOKEN} \
  <image from data/epyc.json> \
  vllm serve <model> --dtype bfloat16 --port <port> --max-model-len <len>

Conda/host path (no container runtime, conda_path_available true). eval-ing cpu_tune already exported the env vars; prefix the launch with conda_launch_prefix from cpu_tune so memory is bound to the chosen socket (empty → unpinned, with a note):

bash
<conda_launch_prefix from cpu_tune> vllm serve <model> --dtype bfloat16 --port <port> --max-model-len <len> &
# e.g. numactl --cpunodebind=0 --membind=0 vllm serve ...

Stock vLLM path (no zentorch) -- only when validate.py reports zentorch_capable: false (pre-Zen4 EPYC like Milan) and the user confirmed the unaccelerated fallback. Use the official stock vLLM CPU image, not the zentorch image: vllm/vllm-openai-cpu:latest-x86_64 (its ENTRYPOINT is vllm serve, so pass <model> --dtype ... --port ... as args). Same sized env + flags as the container launch above (VLLM_CPU_OMP_THREADS_BIND, --cpuset-cpus/--cpuset-mems, --cap-add=SYS_NICE, --ipc=host --network=host, the resolved HF cache mount):

bash
RT="<runtime>"
$RT rm -f vllm-epyc 2>/dev/null
$RT run -d --name vllm-epyc \
  --ipc=host --network=host --cap-add=SYS_NICE \
  <container_cpuset from cpu_tune> <hf_cache.mount from validate.py> \
  --env VLLM_CPU_OMP_THREADS_BIND="$VLLM_CPU_OMP_THREADS_BIND" \
  --env VLLM_CPU_KVCACHE_SPACE=$VLLM_CPU_KVCACHE_SPACE --env HF_TOKEN=${HF_TOKEN} \
  vllm/vllm-openai-cpu:latest-x86_64 \
  <model> --dtype bfloat16 --port <port> --max-model-len <len>

This path is unaccelerated (no zentorch), but verified working on EPYC 7763 (Milan/Zen3) with --dtype bfloat16 -- bf16 runs on stock vLLM CPU without AVX-512, so no fp32 is needed. If it still fails at load, apply the no-retry rule (report + stop).

Optional throughput flags are opt-in and must move together (see Gotchas): TORCHINDUCTOR_FREEZING=1 + VLLM_USE_AOT_COMPILE=0 (+ ZENTORCH_WEIGHT_PREPACK=1). The base launch sets none of them.

Step 7: Poll until up and responsive

A 503 while loading is normal. Poll /health until the server answers, confirm the served model is listed, then prove the selected endpoint works (from primary_endpoint in Step 3). CPU first-token compile can take a minute or two. Track a healthy flag so a timeout is a failure, not a fall-through.

bash
# 1. container alive (conda: process alive) + /health, with a real timeout
healthy=""
for i in $(seq 1 120); do
  $RT inspect -f '{{.State.Running}}' vllm-epyc 2>/dev/null | grep -q true || { echo "FAILED: container exited"; $RT logs --tail 50 vllm-epyc; break; }
  curl -sf http://localhost:<port>/health >/dev/null 2>&1 && { healthy=1; echo "HEALTHY"; break; }
  sleep 3
done
[ -n "$healthy" ] || { echo "FAILED: not healthy before timeout"; $RT logs --tail 50 vllm-epyc; }

# 2. the served model is registered
curl -sf --max-time 30 http://localhost:<port>/v1/models

Then exercise the endpoint the model actually supports. Use deterministic sampling and a small output cap for the smoke check:

bash
# primary_endpoint == chat_completions
curl -sf --max-time 180 http://localhost:<port>/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"<served-model>","messages":[{"role":"user","content":"hi"}],"max_tokens":16,"temperature":0}'

# primary_endpoint == completions  (no chat template)
curl -sf --max-time 180 http://localhost:<port>/v1/completions -H 'Content-Type: application/json' \
  -d '{"model":"<served-model>","prompt":"Hello, world","max_tokens":16,"temperature":0}'

Confirm the response is JSON with a non-error choices[0] (chat: message.content; completion: text). An HTTP 200 that carries an error payload is not success. Resource sanity (your validation list): $RT stats --no-stream vllm-epyc.

If the server never becomes healthy, /v1/models omits the model, or the endpoint returns an error/empty choices: print the container/process logs, state the failing phase, and STOP. Do not retry. Do not start a debugging loop.

Step 8: On success, hand over the endpoint

Give the user everything needed to call the server. Print a connection table:

FieldValue
Base URLhttp://localhost:<port>/v1 (the trailing /v1 matters)
Served model<served-model> (the id from /v1/models)
Endpoint/v1/chat/completions or /v1/completions (from primary_endpoint)
Whychat = a chat template is present; completions = no template (raw prompts)
Runtime / port<runtime> / <port>
SizingOMP threads, KV GB, --max-model-len, socket / NUMA pinning
Stop$RT rm -f vllm-epyc (container) or kill <pid> (conda)

Then a ready-to-run example for the selected endpoint.

Chat model (primary_endpoint: chat_completions):

bash
curl -s http://localhost:<port>/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"<served-model>","messages":[{"role":"user","content":"Hello"}],"max_tokens":128,"temperature":0.7}'

Base/prompt model (primary_endpoint: completions):

bash
curl -s http://localhost:<port>/v1/completions -H 'Content-Type: application/json' \
  -d '{"model":"<served-model>","prompt":"Hello, world","max_tokens":128,"temperature":0.7}'

OpenAI Python client (point base_url at the local server; the SDK requires a non-empty key, so any placeholder works when the server has no auth):

python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:<port>/v1", api_key="EMPTY")
model = client.models.list().data[0].id

# chat model:
r = client.chat.completions.create(
    model=model,
    messages=[{"role": "user", "content": "Hello"}],
    max_tokens=128, temperature=0.7,
)
print(r.choices[0].message.content)

# base/prompt model:
r = client.completions.create(model=model, prompt="Hello, world", max_tokens=128)
print(r.choices[0].text)

Argument guidance to pass along (see reference.md for the full list):

  • max_tokens caps the output; prompt_tokens + max_tokens must be <= --max-model-len.
  • temperature (0 = deterministic/greedy, higher = more random); tune top_p or temperature, not both.
  • stream: true streams tokens (SSE) instead of one blocking response.
  • The model's generation_config.json can set sampling defaults; pass explicit values to be sure.

Offline (single-instance batch)

For a one-shot offline run instead of a server, replace Step 6-8 with a single vllm bench throughput (or an offline LLM.generate) using the same sized env, wait for completion, and report the metrics. Same no-retry / no-debug rule.

Gotchas

See reference.md for the full list. The load-bearing ones:

  • --device cpu was removed from vllm serve in vLLM >= 0.20. The zentorch plugin auto-selects CPU. Passing it makes vllm serve error with "unrecognized arguments: --device cpu".
  • TORCHINDUCTOR_FREEZING=1 alone crashes engine-core init on vLLM 0.23 / zentorch 2.11 (AssertionError: expected OutputCode, got function). It only works with VLLM_USE_AOT_COMPILE=0 set alongside it. Never set one without the other.
  • /dev/shm — use --ipc=host, not --shm-size. vLLM needs a large /dev/shm (the 64MB container default is too small). The base recipe uses --ipc=host, which shares the host's large shared memory. Do not also pass --shm-size: podman errors with "cannot set shmsize when running in the host IPC Namespace", and it is redundant on docker. If you instead isolate IPC (drop --ipc=host), then add --shm-size=16g — one or the other, never both.
  • NUMA / socket: one instance is pinned to one socket plus its memory -- CPU bind + --cpuset-mems (container) / numactl --membind (conda), with KV sized from that socket's local RAM. On a dual-socket host cpu_tune.py picks a free socket by load and warnings if both are busy. NPS2/NPS4 (multi-node socket) gets an nps_note that finer per-node binding could add more.
  • Rootless podman + --cpuset-cpus/--cpuset-mems: these are cgroup limits and may be ignored or rejected on rootless podman without cpuset cgroup delegation (cgroup v1, or v2 without the controller delegated). This is not fatal: CPU thread binding still applies via VLLM_CPU_OMP_THREADS_BIND inside the container; only the container-level memory pin is lost (reduced NUMA locality). If the run errors specifically on the cpuset flags, drop them and proceed -- do not treat it as a launch failure.
  • HF cache mount: the default mounts ~/.cache/huggingface. If HF_HOME points elsewhere (common on shared hosts, e.g. /proj/.../vllm), mount that path to /root/.cache/huggingface instead, or the model re-downloads inside the container.
  • Container name reuse: a leftover vllm-epyc from a prior run makes run fail with "name already in use" -- Step 6 clears it first with $RT rm -f vllm-epyc.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts) in skills/serving-llms-on-epyc of amd/skills.

  • SKILL.md
  • data/epyc.json
  • evals/evals.json
  • evals/machine.yml
  • reference.md
  • scripts/check_model.py
  • scripts/cpu_tune.py
  • scripts/detect.py
  • scripts/estimate_memory.py
  • scripts/validate.py
  • skill-card.md

Open the folder on GitHubat commit 6c92b41

Compare with similar skills

Serving LLMs On Epyc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Serving LLMs On Epyc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Serving LLMs On Epyc this skillamd/skills398—~5.7kAutomated safety check: NotesMIT
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Vss Deployopen-edge-platform/edge-ai-libraries169—~4.1kAutomated safety check: PassApache-2.0
Hyperloom SetupAMD-AGI/Hyperloom217—~7.2kAutomated safety check: NotesCustom licence
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Ascend Model Adapter for vLLMvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Vss Deploy

    open-edge-platform/edge-ai-libraries

    Deploys and manages VSS through setup.sh and its Docker Compose overlays.

    169 GitHub stars~4.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Hyperloom Setup

    AMD-AGI/Hyperloom

    Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

    217 GitHub stars~7.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from amd/skills

All 9 skills in this repo
  • Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…

    398 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.

    398 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    398 GitHub stars~4k tokensUpdated today
    Auto-check: notes
  • Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

    398 GitHub stars~1.7k tokensUpdated today
    Auto-check: notes
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    398 GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Serving LLMs On Epyc

What does Serving LLMs On Epyc do?

Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. Serving LLMs On Epyc is an agent skill from amd/skills. Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

When should I use Serving LLMs On Epyc?

Serving LLMs On Epyc fits situations like: zentorch serving; an EPYC CPU endpoint; including on a host that also has AMD Instinct GPUs.

How do I install Serving LLMs On Epyc in Claude Code?

Run `npx skills add amd/skills --skill serving-llms-on-epyc -a claude-code`. Or copy the skill folder (skills/serving-llms-on-epyc in amd/skills) into .claude/skills/serving-llms-on-epyc in your project. Claude Code loads it when a task matches its description.

How do I install Serving LLMs On Epyc in Codex?

Run `npx skills add amd/skills --skill serving-llms-on-epyc -a codex`. Or copy the skill folder (skills/serving-llms-on-epyc in amd/skills) into .agents/skills/serving-llms-on-epyc in your project. Codex loads it when a task matches its description.

Can I use Serving LLMs On Epyc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/skills --skill serving-llms-on-epyc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-on-epyc, .gemini/skills/serving-llms-on-epyc, .github/skills/serving-llms-on-epyc and .opencode/skills/serving-llms-on-epyc in your project.

What does Serving LLMs On Epyc need to run?

Going by SKILL.md and its folder, Serving LLMs On Epyc needs Python for the scripts in its folder, the command-line tools its instructions call (curl and python3) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read.

Does Serving LLMs On Epyc access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Serving LLMs On Epyc safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Serving LLMs On Epyc use?

Serving LLMs On Epyc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Serving LLMs On Epyc use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Serving LLMs On Epyc?

Skills that share tags, products or a category with Serving LLMs On Epyc: Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Vss Deploy (open-edge-platform/edge-ai-libraries, 169 stars), Hyperloom Setup (AMD-AGI/Hyperloom, 217 stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Serving LLMs On Epyc?

amd (a GitHub organization) maintains it in amd/skills, which has 398 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.

Source: amd/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.