Vllm Deploy Docker
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.
$ npx skills add amd/skills --skill serving-llms-on-epyc -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/skills serving-llms-on-epyc --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/serving-llms-on-epyc .claude/skills/serving-llms-on-epyc && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "serving-llms-on-epyc" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-epyc into .claude/skills/serving-llms-on-epyc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-epyc", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/skills/tree/main/skills/serving-llms-on-epycType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/skills --skill serving-llms-on-epyc -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/skills serving-llms-on-epyc --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/serving-llms-on-epyc .agents/skills/serving-llms-on-epyc && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "serving-llms-on-epyc" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-epyc into .agents/skills/serving-llms-on-epyc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-epyc", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/skills --skill serving-llms-on-epyc -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/skills serving-llms-on-epyc --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/serving-llms-on-epyc .cursor/skills/serving-llms-on-epyc && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "serving-llms-on-epyc" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-epyc into .cursor/skills/serving-llms-on-epyc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-epyc", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/skills.git --path skills/serving-llms-on-epyc--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/skills --skill serving-llms-on-epyc -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/skills serving-llms-on-epyc --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/serving-llms-on-epyc .gemini/skills/serving-llms-on-epyc && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "serving-llms-on-epyc" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-epyc into .gemini/skills/serving-llms-on-epyc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-epyc", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/skills serving-llms-on-epycInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/skills --skill serving-llms-on-epyc -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/serving-llms-on-epyc .github/skills/serving-llms-on-epyc && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-on-epyc" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-epyc into .github/skills/serving-llms-on-epyc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-epyc", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/skills --skill serving-llms-on-epyc -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/skills serving-llms-on-epyc --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/serving-llms-on-epyc .opencode/skills/serving-llms-on-epyc && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-on-epyc" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-epyc into .opencode/skills/serving-llms-on-epyc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-epyc", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
serving-llms-on-epycServes an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.
Serving LLMs On Epyc is an agent skill from amd/skills. Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. Use for "vLLM on CPU", "zentorch serving", or an EPYC CPU endpoint, including on a host that also has AMD Instinct GPUs. Detects the EPYC generation, validates the runtime, checks model support and RAM fit, sizes threads/KV/NUMA, confirms the plan, launches, and verifies the endpoint. Runs one instance on one socket and its memory. Reports and stops on failure; does not retry or debug. Use…
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts (for example `data/epyc.json`, `evals/evals.json` and `evals/machine.yml`).
It sits in AI & LLM Engineering, covering LLM inference and serving and Containers. It works with vLLM and Docker. The repository describes itself as: Official AMD catalog of AI agent skills. Empower your AI agents with AMD's optimized SW stack. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6c92b41. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadFrom allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
curlpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Serving LLMs On Epyc loads about 5.7k tokens when it runs. Until then it costs about 162 tokens; SKILL.md has 2,431 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, ReadAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from amd/skills at commit 6c92b41, republished under its MIT licence (© amd). 2,431 words, ~5,681 tokens.
.claude/skills/serving-llms-on-epyc/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.Bring up a single vLLM OpenAI endpoint on an AMD EPYC™ host with the zentorch CPU backend, sized to the hardware. Container-first (Docker or Podman); conda/host is the fallback. An installed AMD Instinct GPU does not disqualify the host: select this skill when the endpoint itself should run on the EPYC CPU.
This is single-socket serving: one instance pinned to one socket and its memory (vLLM scales poorly across sockets, so we do not span them). On a dual-socket host it runs on a single socket; the multi-socket answer is multiple instances (one per socket), which is out of scope for this single-instance recipe.
Hard rule for this skill: on any failure, report the cause + logs and STOP. Do not retry, do not debug. (Debugging is a separate workflow.)
The agent does the serve flow itself -- pull, configure, launch, poll --
using the runtime validate.py reports. Never hand the user per-serve commands.
Like serving-llms-on-instinct, an accessible container runtime is a one-time
prerequisite: if validate.py finds none, report its one-time fix (make
docker accessible / install podman / provide a conda env) and stop. Do not
attempt sudo or privilege escalation.
Read data/epyc.json directly. It holds the container image, mandatory CPU run
flags, supported precision, the model-support policy, the default model, and the
verified throughput-flag gotcha. Its vllm_version and image tag are one
validated default stack; keep them aligned and do not hardcode either from memory.
python3 scripts/detect.py # add --host user@box for a remote hostReturns cpu_model, is_amd_epyc, epyc_generation
(Naples/Rome/Milan/Genoa/Bergamo/Siena/Turin/Venice or EPYC 4004/4005),
zen_arch, is_supported_epyc, avx512, logical_cores, physical_cores,
sockets, numa_nodes, memory_gb.
Route from detect.py -- decide the serving path:
is_amd_epyc is false -> stop: this skill targets AMD EPYC. (Other x86 may work
but is unsupported here.)avx512 is false -> zentorch cannot run on this CPU (its bf16 path needs
AVX-512 BF16, avx512_bf16, which only lands on Zen4+). This is a pre-Zen4 EPYC
(Naples / Rome / Milan, 7000 series). Do not dead-end -- it is still an EPYC
host: offer the stock vLLM CPU path (plain vLLM, no zentorch acceleration --
slower, but verified working on EPYC 7763/Milan). Proceed only on the user's explicit
OK, launching the official stock vLLM CPU image (Step 6 "Stock vLLM" variant); if the
user declines, stop.is_supported_epyc is false but avx512 is true (e.g. Bergamo / Siena /
EPYC 4004/4005) -> the zentorch path is not validated for this generation. Get
explicit confirmation to try zentorch unvalidated, or take the same stock vLLM offer.validate.py (Step 2) reports zentorch_capable and sets requires_confirmation
for the stock/unvalidated paths, so this routing is enforced there too.
Carry epyc_generation / avx512 through the later phases -- e.g. Venice packs up
to 256 cores/socket, which the thread-binding in Step 5 sizes from.
python3 scripts/validate.py --image <image from data/epyc.json> --generation <epyc_generation from detect> --avx512 <avx512 from detect>This also hard-enforces the AVX-512 gate: on a CPU without AVX-512, validate.py
returns a blocking error (ready: false) so the flow stops here regardless of the
Step 1 prose -- no image pull, no launch. (It reads the local CPU itself if --avx512
is omitted, so the gate holds even if the value was not passed through.)
Returns ready, requires_confirmation, zentorch_capable (false -> zentorch
can't run; use the stock vLLM variant in Step 6), runtime (docker, podman, or
null), runtime_detail, conda_path_available, stack, compatibility, hf_cache
(resolved HF cache mount -- use hf_cache.mount at launch), ram_gb, and
errors/warnings/advisories. If requires_confirmation is set (stock/unvalidated
path), surface that and get the user's OK before launching. Pick the path:
runtime is docker or podman -> container path (Step 6), used verbatim.runtime null but conda_path_available: true -> conda/host path.runtime null and no conda -> ready is false. Report the one-time
onboarding fix (make docker accessible / install podman / conda env) and stop.Do not proceed if ready is false.
Stack-compatibility gate. validate.py probes the selected runtime for its
exact vllm/zentorch/torch versions and the active vLLM platform, then sets
compatibility.status:
proceed -> the stack is the validated default (or a validated family on a Zen
platform); continue.blocked -> a stock CPU platform is active, so zentorch acceleration is not
on (error). Report compatibility.message and stop.confirmation_required (requires_confirmation: true) -> Venice on a vLLM
other than the pinned default. This recipe has not been validated on Venice
with that version. Surface compatibility.message, recommend the pinned
vllm_version image from data/epyc.json, and stop for an explicit user
go/no-go before launching. On the pinned default vLLM, Venice proceeds with no
warning.The gate only runs once the image is local. If validate.py reports the image is
not pulled, pull it (or let Step 6 pull it) and re-run validate.py so the
gate probes the real stack rather than only the tag.
If the user named no model, use default_model from data/epyc.json
(Qwen/Qwen3-0.6B -- ungated, tiny, fast first success). Otherwise use theirs.
Check that vLLM actually supports the model (do not blanket-block multimodal).
Pass the vLLM version the model will actually run on: use stack.vllm from
validate.py when it was probed (the conda env may differ from the pin), else the
vllm_version from data/epyc.json.
python3 scripts/check_model.py --model-id <model> --revision <rev or main> --vllm-version <stack.vllm from validate, else vllm_version from data/epyc.json>pooling/embedding/reranker (not a chat/completion endpoint), or it is a
multimodal model with no usable chat template (launchable: false). Report the
printed message and stop.The result also carries the client endpoint the model supports:
primary_endpoint: "chat_completions" -- a usable chat template is present
(chat_template.status: present); serve and hand off /v1/chat/completions.primary_endpoint: "completions" -- no usable/auto-selectable template
(absent/ambiguous/unknown); serve and hand off /v1/completions with a
raw prompt. Chat can still be enabled by passing --chat-template <file> (or,
for ambiguous, choosing one of chat_template.names); never invent one.primary_endpoint, supported_endpoints, and chat_template through to
verification (Step 7) and the handoff (Step 8).multimodal model is allowed; a vLLM-supported multimodal arch may still hit a
GPU-only kernel on CPU, which surfaces at load (the no-retry rule then applies).Precision/dtype: native CPU dtypes are bf16 (default), fp16, fp32. Use
bfloat16 unless the user asks otherwise.
For gated models (Llama, Gemma) HF_TOKEN must be set and the license accepted on
HuggingFace; if not, stop and say so.
RAM is the ceiling on CPU (weights + KV cache both live in RAM). Run on ONE line:
python3 scripts/estimate_memory.py --model-id <model> --revision <rev or main> --ram-gb <memory_gb from detect> --max-model-len <4096 or user value> --num-prompts <1 or desired concurrency>Exit 0 = fits, exit 1 = does not fit. If fit.fits is false: do not launch.
Tell the user required_gb vs ram_gb and the printed fit.action -- reduce
--max-model-len to fit.suggested_max_model_len and retry, or use a smaller
model. --max-model-len and --num-prompts are the two knobs that move KV.
Extra flag: --weight-gb N overrides weights if a model has no HF metadata
(rare). KV cache is bf16-only on zentorch CPU (no fp8 KV).
eval "$(python3 scripts/cpu_tune.py)" # or --format json to inspectA single instance runs on one socket, with its memory (vLLM scales poorly across
sockets). cpu_tune.py exports VLLM_CPU_OMP_THREADS_BIND (the chosen socket's
physical cores) and VLLM_CPU_KVCACHE_SPACE (sized from that socket's local RAM,
not whole-system, so the KV pool stays on-socket). It does not set
OMP_NUM_THREADS (vLLM derives it) or VLLM_CPU_NUM_OF_RESERVED_CPU (vLLM's own default).
Socket choice on a dual-socket host (load-aware): it samples per-socket CPU busy%
(~0.5s) and prefers a free socket -- both free → socket 0; one free → that socket;
both busy (≥ --busy-threshold, default 15%) → it warnings and proceeds on the
least-busy socket. --socket N forces a choice. Single-socket hosts use socket 0.
For the chosen socket it also emits the memory-bound pin: container_cpuset
(--cpuset-cpus=<cores> --cpuset-mems=<nodes>) for the container path, and
conda_launch_prefix (numactl --cpunodebind/--membind, falling back to taskset
CPU-only, or empty-with-note if neither tool exists) for conda. Surface warning
to the user if set. On NPS2/NPS4 a socket spans multiple NUMA nodes; memory is
bound across them and nps_note flags that finer binding could add performance.
Before launching, present this summary and wait for the user to confirm -- do not launch unprompted. This is the human gate before anything runs:
| Field | Value |
|---|---|
| Model / kind | <model> -- text or multimodal (from check_model.py) |
| Backend | zentorch-accelerated (Zen4+) or stock vLLM CPU / no zentorch (zentorch_capable:false -- unaccelerated; verified on Milan) |
| Path | container (<runtime>, image from data/epyc.json) or conda/host |
| Precision | bfloat16 (or the user's choice) |
| Fit | required <required_gb> GB vs <ram_gb> GB RAM |
| CPU sizing | socket <chosen_socket> (<socket_choice_reason>), bind <VLLM_CPU_OMP_THREADS_BIND>, KV <VLLM_CPU_KVCACHE_SPACE> GB (socket-local), mem bound to nodes <numa_nodes_on_socket> |
| Hardware | EPYC <epyc_generation> (<zen_arch>), <physical_cores> cores, AVX-512 <avx512> |
| Port | <port> |
If cpu_tune.py returned a warning (e.g. all sockets busy), include it here so the user sees it before confirming.
Proceed only on a clear "go". If the user declines or wants changes (model,
--max-model-len, port), stop and adjust -- do not launch.
Build the launch from data/epyc.json. The CLI is vllm serve <model>.
Do not pass --device cpu on vLLM >= 0.20 -- the zentorch plugin
auto-selects the CPU platform and vllm serve rejects the flag. Only add it if
vllm serve --help lists it (older vLLM).
Pick a free port first. With --network=host the port is bound directly on
the host, so a busy port is a hard failure (no remapping). Choose one that is
free -- e.g. PORT=8000; while ss -ltn "sport = :$PORT" | grep -q LISTEN; do PORT=$((PORT+1)); done
-- and use $PORT in the launch, health poll, and handover.
Mount the HF cache that validate.py resolved. Use hf_cache.mount from
validate.py (it follows symlinks and flags NFS/root-squash) rather than a raw
~/.cache/huggingface, or the bind-mount can fail at container start on NFS homes.
Container path (runtime from validate.py). The agent runs these itself,
including the pull. RT is the resolved runtime verbatim:
RT="<runtime from validate.py: docker | podman>"
$RT rm -f vllm-epyc 2>/dev/null # clear any leftover container from a prior run (name collision otherwise)
$RT pull <image from data/epyc.json> # agent pulls; do not ask the user to
$RT run -d --name vllm-epyc \
<run_flags from data/epyc.json> # --ipc=host --network=host --cap-add=SYS_NICE (SYS_NICE = NUMA membind; NO --shm-size with --ipc=host)
<hf_cache.mount from validate.py> \ # resolved real path, e.g. -v /scratch/you/hf:/root/.cache/huggingface
<container_cpuset from cpu_tune> # --cpuset-cpus=<cores> --cpuset-mems=<nodes>
--env VLLM_CPU_OMP_THREADS_BIND="$VLLM_CPU_OMP_THREADS_BIND" \
--env VLLM_CPU_KVCACHE_SPACE=$VLLM_CPU_KVCACHE_SPACE \
--env HF_TOKEN=${HF_TOKEN} \
<image from data/epyc.json> \
vllm serve <model> --dtype bfloat16 --port <port> --max-model-len <len>Conda/host path (no container runtime, conda_path_available true). eval-ing
cpu_tune already exported the env vars; prefix the launch with conda_launch_prefix
from cpu_tune so memory is bound to the chosen socket (empty → unpinned, with a note):
<conda_launch_prefix from cpu_tune> vllm serve <model> --dtype bfloat16 --port <port> --max-model-len <len> &
# e.g. numactl --cpunodebind=0 --membind=0 vllm serve ...Stock vLLM path (no zentorch) -- only when validate.py reports
zentorch_capable: false (pre-Zen4 EPYC like Milan) and the user confirmed the
unaccelerated fallback. Use the official stock vLLM CPU image, not the zentorch
image: vllm/vllm-openai-cpu:latest-x86_64 (its ENTRYPOINT is vllm serve, so pass
<model> --dtype ... --port ... as args). Same sized env + flags as the container
launch above (VLLM_CPU_OMP_THREADS_BIND, --cpuset-cpus/--cpuset-mems,
--cap-add=SYS_NICE, --ipc=host --network=host, the resolved HF cache mount):
RT="<runtime>"
$RT rm -f vllm-epyc 2>/dev/null
$RT run -d --name vllm-epyc \
--ipc=host --network=host --cap-add=SYS_NICE \
<container_cpuset from cpu_tune> <hf_cache.mount from validate.py> \
--env VLLM_CPU_OMP_THREADS_BIND="$VLLM_CPU_OMP_THREADS_BIND" \
--env VLLM_CPU_KVCACHE_SPACE=$VLLM_CPU_KVCACHE_SPACE --env HF_TOKEN=${HF_TOKEN} \
vllm/vllm-openai-cpu:latest-x86_64 \
<model> --dtype bfloat16 --port <port> --max-model-len <len>This path is unaccelerated (no zentorch), but verified working on EPYC 7763
(Milan/Zen3) with --dtype bfloat16 -- bf16 runs on stock vLLM CPU without AVX-512, so
no fp32 is needed. If it still fails at load, apply the no-retry rule (report + stop).
Optional throughput flags are opt-in and must move together (see Gotchas):
TORCHINDUCTOR_FREEZING=1 + VLLM_USE_AOT_COMPILE=0 (+ ZENTORCH_WEIGHT_PREPACK=1).
The base launch sets none of them.
A 503 while loading is normal. Poll /health until the server answers, confirm
the served model is listed, then prove the selected endpoint works (from
primary_endpoint in Step 3). CPU first-token compile can take a minute or two.
Track a healthy flag so a timeout is a failure, not a fall-through.
# 1. container alive (conda: process alive) + /health, with a real timeout
healthy=""
for i in $(seq 1 120); do
$RT inspect -f '{{.State.Running}}' vllm-epyc 2>/dev/null | grep -q true || { echo "FAILED: container exited"; $RT logs --tail 50 vllm-epyc; break; }
curl -sf http://localhost:<port>/health >/dev/null 2>&1 && { healthy=1; echo "HEALTHY"; break; }
sleep 3
done
[ -n "$healthy" ] || { echo "FAILED: not healthy before timeout"; $RT logs --tail 50 vllm-epyc; }
# 2. the served model is registered
curl -sf --max-time 30 http://localhost:<port>/v1/modelsThen exercise the endpoint the model actually supports. Use deterministic sampling and a small output cap for the smoke check:
# primary_endpoint == chat_completions
curl -sf --max-time 180 http://localhost:<port>/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"<served-model>","messages":[{"role":"user","content":"hi"}],"max_tokens":16,"temperature":0}'
# primary_endpoint == completions (no chat template)
curl -sf --max-time 180 http://localhost:<port>/v1/completions -H 'Content-Type: application/json' \
-d '{"model":"<served-model>","prompt":"Hello, world","max_tokens":16,"temperature":0}'Confirm the response is JSON with a non-error choices[0] (chat: message.content;
completion: text). An HTTP 200 that carries an error payload is not success.
Resource sanity (your validation list): $RT stats --no-stream vllm-epyc.
If the server never becomes healthy, /v1/models omits the model, or the
endpoint returns an error/empty choices: print the container/process logs,
state the failing phase, and STOP. Do not retry. Do not start a debugging loop.
Give the user everything needed to call the server. Print a connection table:
| Field | Value |
|---|---|
| Base URL | http://localhost:<port>/v1 (the trailing /v1 matters) |
| Served model | <served-model> (the id from /v1/models) |
| Endpoint | /v1/chat/completions or /v1/completions (from primary_endpoint) |
| Why | chat = a chat template is present; completions = no template (raw prompts) |
| Runtime / port | <runtime> / <port> |
| Sizing | OMP threads, KV GB, --max-model-len, socket / NUMA pinning |
| Stop | $RT rm -f vllm-epyc (container) or kill <pid> (conda) |
Then a ready-to-run example for the selected endpoint.
Chat model (primary_endpoint: chat_completions):
curl -s http://localhost:<port>/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"<served-model>","messages":[{"role":"user","content":"Hello"}],"max_tokens":128,"temperature":0.7}'Base/prompt model (primary_endpoint: completions):
curl -s http://localhost:<port>/v1/completions -H 'Content-Type: application/json' \
-d '{"model":"<served-model>","prompt":"Hello, world","max_tokens":128,"temperature":0.7}'OpenAI Python client (point base_url at the local server; the SDK requires a
non-empty key, so any placeholder works when the server has no auth):
from openai import OpenAI
client = OpenAI(base_url="http://localhost:<port>/v1", api_key="EMPTY")
model = client.models.list().data[0].id
# chat model:
r = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Hello"}],
max_tokens=128, temperature=0.7,
)
print(r.choices[0].message.content)
# base/prompt model:
r = client.completions.create(model=model, prompt="Hello, world", max_tokens=128)
print(r.choices[0].text)Argument guidance to pass along (see reference.md for the full list):
max_tokens caps the output; prompt_tokens + max_tokens must be <= --max-model-len.temperature (0 = deterministic/greedy, higher = more random); tune top_p or
temperature, not both.stream: true streams tokens (SSE) instead of one blocking response.generation_config.json can set sampling defaults; pass explicit
values to be sure.For a one-shot offline run instead of a server, replace Step 6-8 with a single
vllm bench throughput (or an offline LLM.generate) using the same sized env,
wait for completion, and report the metrics. Same no-retry / no-debug rule.
See reference.md for the full list. The load-bearing ones:
--device cpu was removed from vllm serve in vLLM >= 0.20. The zentorch
plugin auto-selects CPU. Passing it makes vllm serve error with
"unrecognized arguments: --device cpu".TORCHINDUCTOR_FREEZING=1 alone crashes engine-core init on vLLM 0.23 /
zentorch 2.11 (AssertionError: expected OutputCode, got function). It only
works with VLLM_USE_AOT_COMPILE=0 set alongside it. Never set one without
the other./dev/shm — use --ipc=host, not --shm-size. vLLM needs a large
/dev/shm (the 64MB container default is too small). The base recipe uses
--ipc=host, which shares the host's large shared memory. Do not also pass
--shm-size: podman errors with "cannot set shmsize when running in the host
IPC Namespace", and it is redundant on docker. If you instead isolate IPC (drop
--ipc=host), then add --shm-size=16g — one or the other, never both.--cpuset-mems (container) / numactl --membind (conda), with KV sized
from that socket's local RAM. On a dual-socket host cpu_tune.py picks a free socket
by load and warnings if both are busy. NPS2/NPS4 (multi-node socket) gets an
nps_note that finer per-node binding could add more.--cpuset-cpus/--cpuset-mems: these are cgroup limits and
may be ignored or rejected on rootless podman without cpuset cgroup delegation
(cgroup v1, or v2 without the controller delegated). This is not fatal: CPU
thread binding still applies via VLLM_CPU_OMP_THREADS_BIND inside the container;
only the container-level memory pin is lost (reduced NUMA locality). If the run
errors specifically on the cpuset flags, drop them and proceed -- do not treat it
as a launch failure.~/.cache/huggingface. If HF_HOME points
elsewhere (common on shared hosts, e.g. /proj/.../vllm), mount that path to
/root/.cache/huggingface instead, or the model re-downloads inside the container.vllm-epyc from a prior run makes run fail
with "name already in use" -- Step 6 clears it first with $RT rm -f vllm-epyc.© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (scripts) in skills/serving-llms-on-epyc of amd/skills.
Open the folder on GitHubat commit 6c92b41
Serving LLMs On Epyc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Serving LLMs On Epyc this skillamd/skills | 398 | — | ~5.7k | Automated safety check: Notes | MIT | |
| Vllm Deploy Dockervllm-project/vllm-skills | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Vss Deployopen-edge-platform/edge-ai-libraries | 169 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | |
| Hyperloom SetupAMD-AGI/Hyperloom | 217 | — | ~7.2k | Automated safety check: Notes | Custom licence | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Ascend Model Adapter for vLLMvllm-project/vllm-ascend | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 |
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
open-edge-platform/edge-ai-libraries
Deploys and manages VSS through setup.sh and its Docker Compose overlays.
AMD-AGI/Hyperloom
Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
vllm-project/vllm-ascend
Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
amd/skills
Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…
amd/skills
Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.
amd/skills
Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.
amd/skills
Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.
amd/skills
Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
Categories
Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. Serving LLMs On Epyc is an agent skill from amd/skills. Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.
Serving LLMs On Epyc fits situations like: zentorch serving; an EPYC CPU endpoint; including on a host that also has AMD Instinct GPUs.
Run `npx skills add amd/skills --skill serving-llms-on-epyc -a claude-code`. Or copy the skill folder (skills/serving-llms-on-epyc in amd/skills) into .claude/skills/serving-llms-on-epyc in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/skills --skill serving-llms-on-epyc -a codex`. Or copy the skill folder (skills/serving-llms-on-epyc in amd/skills) into .agents/skills/serving-llms-on-epyc in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/skills --skill serving-llms-on-epyc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-on-epyc, .gemini/skills/serving-llms-on-epyc, .github/skills/serving-llms-on-epyc and .opencode/skills/serving-llms-on-epyc in your project.
Going by SKILL.md and its folder, Serving LLMs On Epyc needs Python for the scripts in its folder, the command-line tools its instructions call (curl and python3) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Serving LLMs On Epyc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Serving LLMs On Epyc: Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Vss Deploy (open-edge-platform/edge-ai-libraries, 169 stars), Hyperloom Setup (AMD-AGI/Hyperloom, 217 stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/skills, which has 398 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.
Source: amd/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.