Agent skill

Serving LLMs On Instinct

by amd in amd/skills

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

MITAuto-check: notesAI & LLM Engineering

Install Serving LLMs On Instinct

skills CLI
$ npx skills add amd/skills --skill serving-llms-on-instinct -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/skills serving-llms-on-instinct --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/serving-llms-on-instinct .claude/skills/serving-llms-on-instinct && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
serving-llms-on-instinct
GitHub stars
395
Token cost
~4k tokens
SKILL.md length
1,812 words
Files
13 (incl. scripts)
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

  • Works in 6 steps: Detect the GPU → Validate the environment → Refresh recipes (if stale) → …
  • The user wants to run
  • SKILL.md covers Prerequisites, Data files, Step 1: Detect the GPU and Step 2: Validate the environment, plus 7 more sections
  • Runs Python scripts from its folder; calls python3, docker and curl; needs HF_TOKEN

What it does

Serving LLMs On Instinct is an agent skill from amd/skills. Serves AI models on AMD Instinct GPU hardware using vLLM. Use this skill whenever the user wants to run, serve, deploy, start, host, or launch a language model on an AMD GPU, AMD Instinct, MI300X, MI325X, MI350X, or MI355X. Also use when the user mentions vLLM on ROCm, vLLM on AMD, serving on HBM, or asks how to get a model running on AMD data center hardware. Use when the user asks "run Qwen3", "serve DeepSeek", "start a vLLM endpoint", "get a model running on my AMD machine", or any similar phrasing. Handles…

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts (for example `data/blacklist.json`, `data/gpu_overrides.json` and `data/recipes_cache.json`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, Qwen, DeepSeek and NVIDIA AI Platform. The repository describes itself as: Official AMD catalog of AI agent skills. Empower your AI agents with AMD's optimized SW stack. The licence is MIT.

When your agent uses it

  • The user wants to run
  • Launch a language model on an AMD GPU
  • The user mentions vLLM on ROCm
  • Asks how to get a model running on AMD data center hardware

Example prompts

  • “run Qwen3”
  • “serve DeepSeek”
  • “start a vLLM endpoint”
  • “/serving-llms-on-instinct”

Requirements

  • Python 3
  • Docker
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Detect the GPU
  2. Validate the environment
  3. Refresh recipes (if stale)
  4. Construct the Docker command
  5. Confirm with the user
  6. Launch and verify

What it can do on your machine

Read from SKILL.md and the folder at commit 6c92b41. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • docker
    • curl
    • ssh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Serving LLMs On Instinct loads about 4k tokens when it runs. Until then it costs about 187 tokens; SKILL.md has 1,812 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~187
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:344
    group. Fix: `sudo usermod -aG video,render $USER` (requires re-login).
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from amd/skills at commit 6c92b41, republished under its MIT licence (© amd). 1,812 words, ~3,972 tokens.

Download SKILL.mdSave it as .claude/skills/serving-llms-on-instinct/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
serving-llms-on-instinct
description
Serves AI models on AMD Instinct GPU hardware using vLLM. Use this skill whenever the user wants to run, serve, deploy, start, host, or launch a language model on an AMD GPU, AMD Instinct, MI300X, MI325X, MI350X, or MI355X. Also use when the user mentions vLLM on ROCm, vLLM on AMD, serving on HBM, or asks how to get a model running on AMD data center hardware. Use when the user asks "run Qwen3", "serve DeepSeek", "start a vLLM endpoint", "get a model running on my AMD machine", or any similar phrasing. Handles the full flow: GPU detection, environment validation, vLLM configuration, launch, and health verification. Do not use for NVIDIA GPUs, consumer AMD GPUs (RX series, Radeon), Ryzen AI, NPU, MI250X, or MI100.
allowed-tools
Bash, Read

Serving LLMs on AMD Instinct

Get a vLLM endpoint running on AMD Instinct GPU hardware.

Prerequisites

  • ROCm driver and amd-smi installed on the GPU host
  • Docker running and accessible (check with docker ps)
  • /dev/kfd and /dev/dri present on the GPU host
  • HuggingFace token in HF_TOKEN env var (required for gated models; not required for Qwen3 or Gemma). For gated models (Llama 3.2, Gemma, etc.), the HF token must belong to an account that has accepted the model's license at huggingface.co/<model_id>. A valid token without license acceptance will fail with an opaque "Engine core initialization failed" error.
  • For remote GPU: SSH key access configured (ssh <user>@<host> must work without a password prompt). If only password access is available, set up keys first: ssh-copy-id <user>@<host>

Data files

Read these files directly to get model and GPU configuration:

  • data/recipes_cache.json -- model configs synced from vllm-project/recipes. Each entry under models.<HF_ID>.recipe contains the full recipe with model.base_args, model.base_env, features.tool_calling.args, features.reasoning.args, hardware_overrides.amd.extra_args, hardware_overrides.amd.extra_env. The top-level docker_image field has the latest resolved vLLM ROCm image.

  • data/gpu_overrides.json -- GPU-specific configuration. Contains docker_flags (mandatory for all AMD Instinct), gpu_configs keyed by gfx_version with env_defaults and workarounds, and legacy_models for models not yet in vLLM recipes.

  • data/blacklist.json -- models in vLLM recipes that cannot be served as LLM endpoints. Includes diffusion/image/audio generation models, embedding models, rerankers, ASR models needing audio pipelines, and models requiring unreleased vLLM nightly builds. Check this before attempting to serve a model. If the user requests a blacklisted model, explain why it won't work and suggest an alternative.

If the user doesn't specify a model, default to Qwen/Qwen3.5-9B: dense multimodal with MTP, Apache 2.0 license (no HF token needed), fits on a single GPU, strong reasoning and tool-calling.

Step 1: Detect the GPU

bash
python3 scripts/detect.py
# Remote:
python3 scripts/detect.py --host user@hostname

Returns JSON with gfx_version, vram_gb, gpu_count, rocm_version.

gfx_versionHardwareVRAM
gfx950MI350X / MI355X288 GB HBM3E
gfx942MI300X (192 GB) / MI325X (256 GB) / MI300A (128 GB)varies

If gfx_version is unknown: amd-smi ran but found no GPU. Check lsmod | grep amdgpu.

Step 2: Validate the environment

bash
python3 scripts/validate.py --auto-fix
# Remote:
python3 scripts/validate.py --auto-fix --host user@hostname

Returns JSON with ready (bool), errors, warnings, fixes_applied. Do not proceed if ready is false.

Step 3: Refresh recipes (if stale)

Check fetched_at in data/recipes_cache.json. If older than 24 hours or the file is missing, refresh:

bash
python3 scripts/sync_recipes.py

This shallow-clones vllm-project/recipes from GitHub and fetches the latest Docker tag from Docker Hub. Takes ~10 seconds. If it fails, the existing cache still works.

Step 4: Construct the Docker command

Read data/recipes_cache.json and data/gpu_overrides.json directly. Build the Docker command by combining:

  1. Docker flags from gpu_overrides.json > docker_flags (mandatory for all AMD GPUs)
  2. HF cache mount: -v ~/.cache/huggingface:/root/.cache/huggingface (if a shared model cache directory exists on the host, check whether models--* directories are at the cache root or inside a hub/ subdirectory -- mount accordingly to /root/.cache/huggingface or /root/.cache/huggingface/hub)
  3. Port: -p <port>:<port> (default 8000)
  4. Environment variables: merge gpu_configs.<gfx_version>.env_defaults with the recipe's model.base_env and hardware_overrides.amd.extra_env. Always add --env HF_TOKEN=${HF_TOKEN}.
  5. Docker image: use docker_image from recipes_cache.json top level (unless the model needs a pinned image, e.g. GLM-4.5 needs v0.15.1). If the user specifies a Docker image version, check it against the recipe's model.min_vllm_version. Warn if the image is older -- the model may crash on startup with an opaque "Engine core initialization failed" error.
  6. Model ID: --model <HF_ID>
  7. vLLM args: combine the recipe's model.base_args + hardware_overrides.amd.extra_args + features.tool_calling.args + features.reasoning.args. Add --enable-auto-tool-choice if not present. For multi-GPU, add --tensor-parallel-size N (see VRAM estimation below). For MoE models on multi-GPU, also add --distributed-executor-backend mp.
  8. Port arg: --port <port>

If the exact model ID is not in recipes_cache.json, check for a base model match by stripping date/version suffixes (e.g., Kimi-K2-Instruct matches Kimi-K2-Instruct-0905). Use the base model's recipe if found.

If no recipe match, check legacy_models in gpu_overrides.json. If not there either, use a generic config with --enable-auto-tool-choice --trust-remote-code --tool-call-parser hermes.

Precision variant selection: Recipes may offer variants (default, fp8, nvfp4). Check gpu_configs.<gfx_version>.precision.native in gpu_overrides.json before selecting a variant. On gfx942 (MI300X), only bf16, fp16, fp8_fnuz, and int8 are hardware-native. MXFP4 and NVFP4 compute is emulated (dequant to BF16 during matmul), but weights stay compressed in VRAM so quantized models still fit in less memory. On gfx950 (MI350X), MXFP4 is hardware-native.

VRAM estimation and fit check: Before constructing the Docker command, estimate whether the model fits the available hardware:

bash
python3 scripts/estimate_vram.py --model-id <HF_ID> --vram-gb <per_gpu_vram> --tp <N>

This queries the HuggingFace Hub API (no model download) and returns JSON with:

  • weight_memory_gb -- total weight size
  • kv_cache_bytes_per_token -- KV cache cost per token at BF16
  • fit.weights_fit -- whether weights fit at the given TP
  • fit.recommended_max_model_len -- max context the GPU can serve
  • fit.context_limited -- true if KV cache limits context below the model's native max
  • fit.min_tp_required -- minimum TP needed (only if weights don't fit)

Understanding the overhead: The script reserves ~4 GB for vLLM's runtime overhead (activation profiling, HIP graph capture, internal buffers). During startup, vLLM runs a profiling forward pass to measure peak activations, then captures HIP graphs for optimized decode. This startup peak is higher than steady-state. The remaining_for_kv_gb field reflects what's left after weights and this overhead.

Use remaining_for_kv_gb to decide:

  1. remaining_for_kv_gb >= 6: safe to run. If context_limited: true, add --max-model-len <recommended_max_model_len> to the vLLM args. Mention the FP8 KV cache option (--kv-cache-dtype fp8) if the user needs longer context (fit.max_seq_len_fp8_kv shows the gain).
  2. remaining_for_kv_gb between 2 and 6: tight but worth trying. Launch normally. If vLLM OOMs during HIP graph capture (check container logs for "out of memory" after "capturing CUDA/HIP graphs"), retry with --enforce-eager added to the vLLM args. This skips graph capture and frees 1-2 GB. The only cost is slightly higher decode latency.
  3. remaining_for_kv_gb < 2: too tight. Will likely OOM during the activation profiling step. Do not attempt.
  4. weights_fit: false with multiple GPUs: re-run with --tp <min_tp_required> and check again.
  5. weights_fit: false, not enough GPUs: look for quantized alternatives in this order: a. Recipe variants: the recipe may have fp8 or mxfp4 variants with a different model_id that points to a quantized checkpoint. b. Same provider: many providers release quantized versions alongside the base model (e.g. Qwen/Qwen3.5-122B-FP8 from Qwen). Search HuggingFace for <provider>/<model-name> with FP8/GPTQ/AWQ suffixes. c. AMD quantized: AMD's Quark team publishes quantized models under the amd/ org on HuggingFace (e.g. amd/Kimi-K2-Instruct-w-mxfp4-a-fp8). Search for amd/<model-name> variants. Run estimate_vram.py on the quantized model ID to verify it fits, then use that model ID instead.
  6. Still doesn't fit: tell the user the model requires more VRAM than available and suggest either a smaller model or multi-GPU hardware. Do not attempt to launch.

Docker command template:

docker run -d --name vllm-<model-slug> \
  <docker_flags> \
  -v <hf_cache_mount> \
  -p <port>:<port> \
  --env <key>=<value> (for each env var) \
  --env HF_TOKEN=${HF_TOKEN} \
  <docker_image> \
  --model <model_id> \
  <vllm_args> \
  --port <port>
Show full SKILL.md (746 more words)Show less

Step 5: Confirm with the user

Before launching, present a summary and ask the user to confirm:

  • Model: full HuggingFace ID (e.g. Qwen/Qwen3.5-122B-Instruct)
  • Precision: variant being used (e.g. BF16, FP8) and why
  • Weight memory: from estimate_vram.py
  • GPU: detected hardware and VRAM
  • TP: tensor parallelism degree (1, 2, 4, 8)
  • Context: max achievable context length (and whether it's limited)
  • Port: which port the endpoint will be on

If a quantized alternative was selected (Step 4 fit check), explain that the original model doesn't fit and which alternative is being used.

Wait for the user's confirmation before proceeding.

Step 6: Launch and verify

Before launching, check for port conflicts:

bash
ss -tlnp 2>/dev/null | grep ':<port> '

If a Docker container is on that port, stop it with docker rm -f <name>.

Run the Docker command. Then poll health using this loop:

bash
while docker inspect --format='{{.State.Running}}' <container_name> 2>/dev/null | grep -q true; do
  curl -sf http://localhost:<port>/health && echo "READY" && exit 0
  sleep 60
done
echo "FAILED -- container exited"

A 503 during loading is normal. Choose the polling strategy based on model size (weight memory from hf-mem):

  • Small models (< 100 GB weights): run the poll as a blocking command with the Bash tool's timeout set to 600000 (10 minutes). Most cached models are ready within 2-5 minutes.
  • Large models (>= 100 GB weights): run the poll with the Bash tool's run_in_background set to true. Then use TaskOutput with block: true and timeout: 600000 to wait up to 10 minutes per check. If the task is still running after that, call TaskOutput again with the same parameters. This uses only 1 turn per 10-minute wait instead of burning a turn every check. The background loop runs until the container is healthy or dies.

After health returns 200, send a warmup request (triggers HIP kernel compilation, ~40-45 seconds on gfx942):

bash
curl -s http://localhost:<port>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"<model_id>","messages":[{"role":"user","content":"say hi"}],"max_tokens":5}'

After the warmup succeeds, present a connection table so the user can call the endpoint immediately:

FieldValue
Model<model_id>
Served model name<served-model-name or model_id>
Base URLhttp://<host>:<port>/v1
API keynone (local)
Port<port>
Tensor parallel<tp>
Max context<context>
GPU<detected GPU>

Then give a ready-to-run example using those exact values:

bash
curl -s http://<host>:<port>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"<model_id>","messages":[{"role":"user","content":"Hello"}]}'

Remote vs. local

All scripts accept --host user@hostname. When given, they SSH to the target. Set ROCM_SSH_HOST and ROCM_SSH_USER env vars to avoid passing --host every time.

For remote Docker commands, run them over SSH:

bash
ssh user@host 'docker run -d ...'

Use localhost for health/warmup curl URLs (curl runs on the remote host).

Gotchas

CUDA_VISIBLE_DEVICES set to empty string -- ROCm maps this variable to HIP_VISIBLE_DEVICES. Setting it to an empty string hides all GPUs. CUDA_VISIBLE_DEVICES=0,1 works fine for restricting GPUs (same as HIP_VISIBLE_DEVICES=0,1). If the host has it set to empty, unset it: unset CUDA_VISIBLE_DEVICES. Do not pass --env CUDA_VISIBLE_DEVICES= (empty) into Docker -- that also hides all GPUs inside the container.

FP4BMM crash on gfx942 (MI300X) -- If the container exits immediately with a segfault or illegal instruction: VLLM_ROCM_USE_AITER_FP4BMM must be 0 on gfx942. This is set correctly in gpu_overrides.json for gfx942. See vLLM issue #34641.

HIP error: no kernel image -- The Docker image has no compiled kernel for your GPU's gfx version. Use vllm/vllm-openai-rocm:latest; it includes gfx942 and gfx950 kernels.

MLA models need --block-size 1 -- DeepSeek-R1/V3, Kimi-K2.5. Without it the MLA attention backend silently falls back to a slower path. This is in the recipe args for these models.

MoE models on multi-GPU need --distributed-executor-backend mp -- Qwen3-235B, GLM-4.5, MiniMax-M2. The default distributed executor does not work reliably with MoE on ROCm.

OOM during HIP graph capture -- If the container logs show "out of memory" after "capturing CUDA graphs" or "capturing HIP graphs", the model fits in VRAM but there isn't enough headroom for graph capture. Retry with --enforce-eager added to the vLLM args. This disables graph capture and frees 1-2 GB. Trade-off: slightly higher decode latency, but the model runs.

"Engine core initialization failed" -- This opaque error means the engine core subprocess died. Check early container logs: docker logs <name> 2>&1 | head -50. Common causes: gated model access denied (license not accepted on HF), unsupported architecture on this vLLM version, OOM during weight loading, missing --trust-remote-code for custom architectures, or vLLM version too old for the model (check min_vllm_version in the recipe).

/dev/kfd permission denied -- User is not in the video or render group. Fix: sudo usermod -aG video,render $USER (requires re-login).

SSH key not configured -- The scripts use BatchMode=yes SSH. If SSH fails with Permission denied (publickey), configure key-based access first.

Restricting GPUs on shared hosts -- Use --env HIP_VISIBLE_DEVICES=0,1 or --env CUDA_VISIBLE_DEVICES=0,1 to target specific GPUs by index. HIP_VISIBLE_DEVICES is the canonical AMD variable; CUDA_VISIBLE_DEVICES also works (ROCm maps it). Never set either to an empty string.


Reference

Precision compatibility, VRAM estimation, Docker flags, and known quirks: reference.md

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts) in skills/serving-llms-on-instinct of amd/skills.

  • SKILL.md
  • data/blacklist.json
  • data/gpu_overrides.json
  • data/recipes_cache.json
  • evals/evals.json
  • evals/hooks.py
  • evals/machine.yml
  • reference.md
  • scripts/detect.py
  • scripts/estimate_vram.py
  • scripts/sync_recipes.py
  • scripts/validate.py
  • skill-card.md

Open the folder on GitHubat commit 6c92b41

Compare with similar skills

Serving LLMs On Instinct next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Serving LLMs On Instinct compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Serving LLMs On Instinct this skillamd/skills395—~4kAutomated safety check: NotesMIT
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS900—~3.9kAutomated safety check: PassNone
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Jetson PackageNVIDIA/skills3.5k1 repos~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    900 GitHub stars~3.9k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Jetson Package

    NVIDIA/skills

    Official

    Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

    3.5k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Ascend

    ascend-ai-coding/awesome-ascend-skills

    vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

    174 GitHub stars~2.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from amd/skills

All 9 skills in this repo
  • Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…

    395 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.

    395 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    395 GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

    395 GitHub stars~5.7k tokensUpdated today
    Auto-check: notes
  • Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

    395 GitHub stars~1.7k tokensUpdated today
    Auto-check: notes
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    395 GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Questions about Serving LLMs On Instinct

What does Serving LLMs On Instinct do?

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. Serving LLMs On Instinct is an agent skill from amd/skills. Serves AI models on AMD Instinct GPU hardware using vLLM.

When should I use Serving LLMs On Instinct?

Serving LLMs On Instinct fits situations like: the user wants to run; launch a language model on an AMD GPU; the user mentions vLLM on ROCm; asks how to get a model running on AMD data center hardware.

How do I install Serving LLMs On Instinct in Claude Code?

Run `npx skills add amd/skills --skill serving-llms-on-instinct -a claude-code`. Or copy the skill folder (skills/serving-llms-on-instinct in amd/skills) into .claude/skills/serving-llms-on-instinct in your project. Claude Code loads it when a task matches its description.

How do I install Serving LLMs On Instinct in Codex?

Run `npx skills add amd/skills --skill serving-llms-on-instinct -a codex`. Or copy the skill folder (skills/serving-llms-on-instinct in amd/skills) into .agents/skills/serving-llms-on-instinct in your project. Codex loads it when a task matches its description.

Can I use Serving LLMs On Instinct in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/skills --skill serving-llms-on-instinct -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-on-instinct, .gemini/skills/serving-llms-on-instinct, .github/skills/serving-llms-on-instinct and .opencode/skills/serving-llms-on-instinct in your project.

What does Serving LLMs On Instinct need to run?

Going by SKILL.md and its folder, Serving LLMs On Instinct needs Python for the scripts in its folder, the command-line tools its instructions call (python3, docker, curl and ssh) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read.

Does Serving LLMs On Instinct access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Serving LLMs On Instinct safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Serving LLMs On Instinct use?

Serving LLMs On Instinct is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Serving LLMs On Instinct use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Serving LLMs On Instinct?

Skills that share tags, products or a category with Serving LLMs On Instinct: Dstack Prototyping (dstackai/dstack, 2.3k stars), Add Model (guoqingbao/xinfer, 333 stars), LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars) and Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Serving LLMs On Instinct?

amd (a GitHub organization) maintains it in amd/skills, which has 395 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.

Source: amd/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.