Dstack Prototyping
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.
$ npx skills add amd/skills --skill serving-llms-on-instinct -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/skills serving-llms-on-instinct --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/serving-llms-on-instinct .claude/skills/serving-llms-on-instinct && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "serving-llms-on-instinct" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-instinct into .claude/skills/serving-llms-on-instinct/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-instinct", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/skills/tree/main/skills/serving-llms-on-instinctType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/skills --skill serving-llms-on-instinct -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/skills serving-llms-on-instinct --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/serving-llms-on-instinct .agents/skills/serving-llms-on-instinct && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "serving-llms-on-instinct" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-instinct into .agents/skills/serving-llms-on-instinct/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-instinct", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/skills --skill serving-llms-on-instinct -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/skills serving-llms-on-instinct --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/serving-llms-on-instinct .cursor/skills/serving-llms-on-instinct && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "serving-llms-on-instinct" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-instinct into .cursor/skills/serving-llms-on-instinct/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-instinct", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/skills.git --path skills/serving-llms-on-instinct--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/skills --skill serving-llms-on-instinct -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/skills serving-llms-on-instinct --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/serving-llms-on-instinct .gemini/skills/serving-llms-on-instinct && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "serving-llms-on-instinct" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-instinct into .gemini/skills/serving-llms-on-instinct/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-instinct", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/skills serving-llms-on-instinctInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/skills --skill serving-llms-on-instinct -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/serving-llms-on-instinct .github/skills/serving-llms-on-instinct && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-on-instinct" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-instinct into .github/skills/serving-llms-on-instinct/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-instinct", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/skills --skill serving-llms-on-instinct -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/skills serving-llms-on-instinct --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/serving-llms-on-instinct .opencode/skills/serving-llms-on-instinct && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-on-instinct" agent skill from https://github.com/amd/skills/tree/main/skills/serving-llms-on-instinct into .opencode/skills/serving-llms-on-instinct/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-on-instinct", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
serving-llms-on-instinctServes AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.
Serving LLMs On Instinct is an agent skill from amd/skills. Serves AI models on AMD Instinct GPU hardware using vLLM. Use this skill whenever the user wants to run, serve, deploy, start, host, or launch a language model on an AMD GPU, AMD Instinct, MI300X, MI325X, MI350X, or MI355X. Also use when the user mentions vLLM on ROCm, vLLM on AMD, serving on HBM, or asks how to get a model running on AMD data center hardware. Use when the user asks "run Qwen3", "serve DeepSeek", "start a vLLM endpoint", "get a model running on my AMD machine", or any similar phrasing. Handles…
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts (for example `data/blacklist.json`, `data/gpu_overrides.json` and `data/recipes_cache.json`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, Qwen, DeepSeek and NVIDIA AI Platform. The repository describes itself as: Official AMD catalog of AI agent skills. Empower your AI agents with AMD's optimized SW stack. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6c92b41. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadFrom allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3dockercurlsshFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Serving LLMs On Instinct loads about 4k tokens when it runs. Until then it costs about 187 tokens; SKILL.md has 1,812 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
group. Fix: `sudo usermod -aG video,render $USER` (requires re-login).allowed-tools: Bash, ReadAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from amd/skills at commit 6c92b41, republished under its MIT licence (© amd). 1,812 words, ~3,972 tokens.
.claude/skills/serving-llms-on-instinct/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.Get a vLLM endpoint running on AMD Instinct GPU hardware.
amd-smi installed on the GPU hostdocker ps)/dev/kfd and /dev/dri present on the GPU hostHF_TOKEN env var (required for gated models; not
required for Qwen3 or Gemma). For gated models (Llama 3.2, Gemma, etc.),
the HF token must belong to an account that has accepted the model's license
at huggingface.co/<model_id>. A valid token without license acceptance will
fail with an opaque "Engine core initialization failed" error.ssh <user>@<host> must work
without a password prompt). If only password access is available, set up
keys first: ssh-copy-id <user>@<host>Read these files directly to get model and GPU configuration:
data/recipes_cache.json -- model configs synced from
vllm-project/recipes. Each entry
under models.<HF_ID>.recipe contains the full recipe with model.base_args,
model.base_env, features.tool_calling.args, features.reasoning.args,
hardware_overrides.amd.extra_args, hardware_overrides.amd.extra_env.
The top-level docker_image field has the latest resolved vLLM ROCm image.
data/gpu_overrides.json -- GPU-specific configuration. Contains
docker_flags (mandatory for all AMD Instinct), gpu_configs keyed by
gfx_version with env_defaults and workarounds, and legacy_models for
models not yet in vLLM recipes.
data/blacklist.json -- models in vLLM recipes that cannot be served
as LLM endpoints. Includes diffusion/image/audio generation models, embedding
models, rerankers, ASR models needing audio pipelines, and models requiring
unreleased vLLM nightly builds. Check this before attempting to serve a model.
If the user requests a blacklisted model, explain why it won't work and
suggest an alternative.
If the user doesn't specify a model, default to Qwen/Qwen3.5-9B: dense multimodal with MTP, Apache 2.0 license (no HF token needed), fits on a single GPU, strong reasoning and tool-calling.
python3 scripts/detect.py
# Remote:
python3 scripts/detect.py --host user@hostnameReturns JSON with gfx_version, vram_gb, gpu_count, rocm_version.
| gfx_version | Hardware | VRAM |
|---|---|---|
| gfx950 | MI350X / MI355X | 288 GB HBM3E |
| gfx942 | MI300X (192 GB) / MI325X (256 GB) / MI300A (128 GB) | varies |
If gfx_version is unknown: amd-smi ran but found no GPU. Check
lsmod | grep amdgpu.
python3 scripts/validate.py --auto-fix
# Remote:
python3 scripts/validate.py --auto-fix --host user@hostnameReturns JSON with ready (bool), errors, warnings, fixes_applied.
Do not proceed if ready is false.
Check fetched_at in data/recipes_cache.json. If older than 24 hours or
the file is missing, refresh:
python3 scripts/sync_recipes.pyThis shallow-clones vllm-project/recipes from GitHub and fetches the latest Docker tag from Docker Hub. Takes ~10 seconds. If it fails, the existing cache still works.
Read data/recipes_cache.json and data/gpu_overrides.json directly.
Build the Docker command by combining:
gpu_overrides.json > docker_flags (mandatory for all AMD GPUs)-v ~/.cache/huggingface:/root/.cache/huggingface
(if a shared model cache directory exists on the host, check whether
models--* directories are at the cache root or inside a hub/
subdirectory -- mount accordingly to /root/.cache/huggingface or
/root/.cache/huggingface/hub)-p <port>:<port> (default 8000)gpu_configs.<gfx_version>.env_defaults
with the recipe's model.base_env and hardware_overrides.amd.extra_env.
Always add --env HF_TOKEN=${HF_TOKEN}.docker_image from recipes_cache.json top level
(unless the model needs a pinned image, e.g. GLM-4.5 needs v0.15.1).
If the user specifies a Docker image version, check it against the recipe's
model.min_vllm_version. Warn if the image is older -- the model may crash
on startup with an opaque "Engine core initialization failed" error.--model <HF_ID>model.base_args +
hardware_overrides.amd.extra_args + features.tool_calling.args +
features.reasoning.args. Add --enable-auto-tool-choice if not present.
For multi-GPU, add --tensor-parallel-size N (see VRAM estimation below).
For MoE models on multi-GPU, also add --distributed-executor-backend mp.--port <port>If the exact model ID is not in recipes_cache.json, check for a base model
match by stripping date/version suffixes (e.g., Kimi-K2-Instruct matches
Kimi-K2-Instruct-0905). Use the base model's recipe if found.
If no recipe match, check legacy_models in gpu_overrides.json. If not
there either, use a generic config with
--enable-auto-tool-choice --trust-remote-code --tool-call-parser hermes.
Precision variant selection: Recipes may offer variants (default, fp8,
nvfp4). Check gpu_configs.<gfx_version>.precision.native in
gpu_overrides.json before selecting a variant. On gfx942 (MI300X), only
bf16, fp16, fp8_fnuz, and int8 are hardware-native. MXFP4 and NVFP4
compute is emulated (dequant to BF16 during matmul), but weights stay
compressed in VRAM so quantized models still fit in less memory.
On gfx950 (MI350X), MXFP4 is hardware-native.
VRAM estimation and fit check: Before constructing the Docker command, estimate whether the model fits the available hardware:
python3 scripts/estimate_vram.py --model-id <HF_ID> --vram-gb <per_gpu_vram> --tp <N>This queries the HuggingFace Hub API (no model download) and returns JSON with:
weight_memory_gb -- total weight sizekv_cache_bytes_per_token -- KV cache cost per token at BF16fit.weights_fit -- whether weights fit at the given TPfit.recommended_max_model_len -- max context the GPU can servefit.context_limited -- true if KV cache limits context below the
model's native maxfit.min_tp_required -- minimum TP needed (only if weights don't fit)Understanding the overhead: The script reserves ~4 GB for vLLM's runtime
overhead (activation profiling, HIP graph capture, internal buffers). During
startup, vLLM runs a profiling forward pass to measure peak activations, then
captures HIP graphs for optimized decode. This startup peak is higher than
steady-state. The remaining_for_kv_gb field reflects what's left after
weights and this overhead.
Use remaining_for_kv_gb to decide:
remaining_for_kv_gb >= 6: safe to run. If context_limited: true,
add --max-model-len <recommended_max_model_len> to the vLLM args.
Mention the FP8 KV cache option (--kv-cache-dtype fp8) if the user
needs longer context (fit.max_seq_len_fp8_kv shows the gain).remaining_for_kv_gb between 2 and 6: tight but worth trying. Launch
normally. If vLLM OOMs during HIP graph capture (check container logs for
"out of memory" after "capturing CUDA/HIP graphs"), retry with
--enforce-eager added to the vLLM args. This skips graph capture and
frees 1-2 GB. The only cost is slightly higher decode latency.remaining_for_kv_gb < 2: too tight. Will likely OOM during the
activation profiling step. Do not attempt.weights_fit: false with multiple GPUs: re-run with
--tp <min_tp_required> and check again.weights_fit: false, not enough GPUs: look for quantized
alternatives in this order:
a. Recipe variants: the recipe may have fp8 or mxfp4 variants
with a different model_id that points to a quantized checkpoint.
b. Same provider: many providers release quantized versions alongside
the base model (e.g. Qwen/Qwen3.5-122B-FP8 from Qwen). Search
HuggingFace for <provider>/<model-name> with FP8/GPTQ/AWQ suffixes.
c. AMD quantized: AMD's Quark team publishes quantized models under
the amd/ org on HuggingFace (e.g. amd/Kimi-K2-Instruct-w-mxfp4-a-fp8).
Search for amd/<model-name> variants.
Run estimate_vram.py on the quantized model ID to verify it fits,
then use that model ID instead.Docker command template:
docker run -d --name vllm-<model-slug> \
<docker_flags> \
-v <hf_cache_mount> \
-p <port>:<port> \
--env <key>=<value> (for each env var) \
--env HF_TOKEN=${HF_TOKEN} \
<docker_image> \
--model <model_id> \
<vllm_args> \
--port <port>Before launching, present a summary and ask the user to confirm:
Qwen/Qwen3.5-122B-Instruct)If a quantized alternative was selected (Step 4 fit check), explain that the original model doesn't fit and which alternative is being used.
Wait for the user's confirmation before proceeding.
Before launching, check for port conflicts:
ss -tlnp 2>/dev/null | grep ':<port> 'If a Docker container is on that port, stop it with docker rm -f <name>.
Run the Docker command. Then poll health using this loop:
while docker inspect --format='{{.State.Running}}' <container_name> 2>/dev/null | grep -q true; do
curl -sf http://localhost:<port>/health && echo "READY" && exit 0
sleep 60
done
echo "FAILED -- container exited"A 503 during loading is normal. Choose the polling strategy based on model size (weight memory from hf-mem):
timeout set to 600000 (10 minutes). Most cached
models are ready within 2-5 minutes.run_in_background set to true. Then use TaskOutput with
block: true and timeout: 600000 to wait up to 10 minutes per check.
If the task is still running after that, call TaskOutput again with
the same parameters. This uses only 1 turn per 10-minute wait instead
of burning a turn every check. The background loop runs until the
container is healthy or dies.After health returns 200, send a warmup request (triggers HIP kernel compilation, ~40-45 seconds on gfx942):
curl -s http://localhost:<port>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"<model_id>","messages":[{"role":"user","content":"say hi"}],"max_tokens":5}'After the warmup succeeds, present a connection table so the user can call the endpoint immediately:
| Field | Value |
|---|---|
| Model | <model_id> |
| Served model name | <served-model-name or model_id> |
| Base URL | http://<host>:<port>/v1 |
| API key | none (local) |
| Port | <port> |
| Tensor parallel | <tp> |
| Max context | <context> |
| GPU | <detected GPU> |
Then give a ready-to-run example using those exact values:
curl -s http://<host>:<port>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"<model_id>","messages":[{"role":"user","content":"Hello"}]}'All scripts accept --host user@hostname. When given, they SSH to the target.
Set ROCM_SSH_HOST and ROCM_SSH_USER env vars to avoid passing --host
every time.
For remote Docker commands, run them over SSH:
ssh user@host 'docker run -d ...'Use localhost for health/warmup curl URLs (curl runs on the remote host).
CUDA_VISIBLE_DEVICES set to empty string -- ROCm maps this variable to
HIP_VISIBLE_DEVICES. Setting it to an empty string hides all GPUs.
CUDA_VISIBLE_DEVICES=0,1 works fine for restricting GPUs (same as
HIP_VISIBLE_DEVICES=0,1). If the host has it set to empty, unset it:
unset CUDA_VISIBLE_DEVICES. Do not pass --env CUDA_VISIBLE_DEVICES= (empty)
into Docker -- that also hides all GPUs inside the container.
FP4BMM crash on gfx942 (MI300X) -- If the container exits immediately
with a segfault or illegal instruction: VLLM_ROCM_USE_AITER_FP4BMM must be
0 on gfx942. This is set correctly in gpu_overrides.json for gfx942.
See vLLM issue #34641.
HIP error: no kernel image -- The Docker image has no compiled kernel
for your GPU's gfx version. Use vllm/vllm-openai-rocm:latest; it includes
gfx942 and gfx950 kernels.
MLA models need --block-size 1 -- DeepSeek-R1/V3, Kimi-K2.5.
Without it the MLA attention backend silently falls back to a slower path.
This is in the recipe args for these models.
MoE models on multi-GPU need --distributed-executor-backend mp --
Qwen3-235B, GLM-4.5, MiniMax-M2. The default distributed executor does not
work reliably with MoE on ROCm.
OOM during HIP graph capture -- If the container logs show "out of memory"
after "capturing CUDA graphs" or "capturing HIP graphs", the model fits in
VRAM but there isn't enough headroom for graph capture. Retry with
--enforce-eager added to the vLLM args. This disables graph capture and
frees 1-2 GB. Trade-off: slightly higher decode latency, but the model runs.
"Engine core initialization failed" -- This opaque error means the engine
core subprocess died. Check early container logs: docker logs <name> 2>&1 | head -50. Common causes: gated model access denied (license not accepted on
HF), unsupported architecture on this vLLM version, OOM during weight loading,
missing --trust-remote-code for custom architectures, or vLLM version too old
for the model (check min_vllm_version in the recipe).
/dev/kfd permission denied -- User is not in the video or render
group. Fix: sudo usermod -aG video,render $USER (requires re-login).
SSH key not configured -- The scripts use BatchMode=yes SSH. If SSH
fails with Permission denied (publickey), configure key-based access first.
Restricting GPUs on shared hosts -- Use --env HIP_VISIBLE_DEVICES=0,1
or --env CUDA_VISIBLE_DEVICES=0,1 to target specific GPUs by index.
HIP_VISIBLE_DEVICES is the canonical AMD variable; CUDA_VISIBLE_DEVICES
also works (ROCm maps it). Never set either to an empty string.
Precision compatibility, VRAM estimation, Docker flags, and known quirks: reference.md
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts) in skills/serving-llms-on-instinct of amd/skills.
Open the folder on GitHubat commit 6c92b41
Serving LLMs On Instinct next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Serving LLMs On Instinct this skillamd/skills | 395 | — | ~4k | Automated safety check: Notes | MIT | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Add Modelguoqingbao/xinfer | 333 | — | ~4.2k | Automated safety check: Notes | MIT | |
| LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 900 | — | ~3.9k | Automated safety check: Pass | None | |
| Vllm Deploy Dockervllm-project/vllm-skills | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Jetson PackageNVIDIA/skills | 3.5k | 1 repos | ~1.8k | Automated safety check: Pass | Apache-2.0 |
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
BBuf/AI-Infra-Auto-Driven-SKILLS
Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
NVIDIA/skills
Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.
ascend-ai-coding/awesome-ascend-skills
vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.
amd/skills
Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…
amd/skills
Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.
amd/skills
Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.
amd/skills
Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.
amd/skills
Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
Works with
Categories
Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. Serving LLMs On Instinct is an agent skill from amd/skills. Serves AI models on AMD Instinct GPU hardware using vLLM.
Serving LLMs On Instinct fits situations like: the user wants to run; launch a language model on an AMD GPU; the user mentions vLLM on ROCm; asks how to get a model running on AMD data center hardware.
Run `npx skills add amd/skills --skill serving-llms-on-instinct -a claude-code`. Or copy the skill folder (skills/serving-llms-on-instinct in amd/skills) into .claude/skills/serving-llms-on-instinct in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/skills --skill serving-llms-on-instinct -a codex`. Or copy the skill folder (skills/serving-llms-on-instinct in amd/skills) into .agents/skills/serving-llms-on-instinct in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/skills --skill serving-llms-on-instinct -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-on-instinct, .gemini/skills/serving-llms-on-instinct, .github/skills/serving-llms-on-instinct and .opencode/skills/serving-llms-on-instinct in your project.
Going by SKILL.md and its folder, Serving LLMs On Instinct needs Python for the scripts in its folder, the command-line tools its instructions call (python3, docker, curl and ssh) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Serving LLMs On Instinct is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Serving LLMs On Instinct: Dstack Prototyping (dstackai/dstack, 2.3k stars), Add Model (guoqingbao/xinfer, 333 stars), LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars) and Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/skills, which has 395 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.
Source: amd/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.