Agent skill

Vllm Deploy Docker

by vllm-project in vllm-project/vllm-skills

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Vllm Deploy Docker

skills CLI
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-skills vllm-deploy-docker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-docker .claude/skills/vllm-deploy-docker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-deploy-docker
GitHub stars
103
Token cost
~2.5k tokens
SKILL.md length
961 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers What this skill does, Prerequisites, Quickstart using Pre-built… and Build Docker image from source, plus 5 more sections
  • Calls docker and curl; reaches github.com; needs HF_TOKEN
  • Tasks that involve Containers

What it does

Vllm Deploy Docker is an agent skill from vllm-project/vllm-skills. Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Containers. It works with vLLM, Docker, NVIDIA AI Platform and OpenAI. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Containers

Example prompts

  • “/vllm-deploy-docker”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • docs.vllm.ai
    • hub.docker.com
    • docs.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Deploy Docker loads about 2.5k tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 961 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:190
    sudo groupadd docker
  • NoteRuns commands with sudoSKILL.md:193
    sudo usermod -aG docker $USER

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 961 words, ~2,466 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-deploy-docker/SKILL.md (or your agent's skills folder).
name
vllm-deploy-docker
description
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vLLM Docker Deployment

A Claude skill describing how to deploy vLLM with Docker using the official pre-built images or building the image from source supporting NVIDIA GPUs with CUDA. Instructions include NVIDIA CUDA support, example docker run and a minimal docker-compose snippet, recommended flags, and troubleshooting notes. For AMD, Intel, or other accelerators, please refer to the vLLM documentation for alternative deployment methods.

What this skill does

  • Deploy vLLM with docker using pre-built images (recommended for most users) or build from source for custom configurations
  • Provide example commands for running the OpenAI-compatible server with GPU access and mounted Hugging Face cache
  • Point to build-from-source instructions when a custom image or optional dependencies are needed
  • Explain common flags: --ipc=host, shared cache mounts, and HF_TOKEN handling

Prerequisites

  • Docker Engine installed (Docker 20.10+ recommended)
  • NVIDIA GPU(s) with appropriate drivers and CUDA toolkit installed
  • Optional: curl for API tests
  • A Hugging Face token if pulling private models or to avoid rate-limits: HF_TOKEN

Run a vLLM OpenAI-compatible server with GPU access, mounting the HF cache and forwarding port 8000:

bash
docker run --rm --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --env "HF_TOKEN=$HF_TOKEN" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model Qwen/Qwen2.5-1.5B-Instruct
  • --gpus all exposes all GPUs to the container. Adjust if you need specific GPUs.
  • --ipc=host or an appropriately large --shm-size is recommended so PyTorch and vLLM can share host shared memory.
  • Mounting ~/.cache/huggingface avoids re-downloading models inside the container.

Note: vLLM and this skill recommend using the latest Docker image (vllm/vllm-openai:latest). For legacy version images, you may refer to the Docker Hub image tags.

Build Docker image from source

You can build and run vLLM from source by using the provided docker/Dockerfile. First, check the hardware of the host machine and ensure you have the necessary dependencies installed (e.g., NVIDIA drivers, CUDA toolkit, Docker with BuildKit support). For ARM64/aarch64 builds, refer to the "Building for ARM64/aarch64" section.

Basic build command
bash
DOCKER_BUILDKIT=1 docker build . \
  --target vllm-openai \
  --tag vllm/vllm-openai \
  --file docker/Dockerfile

The --target vllm-openai specifies that you are building the OpenAI-compatible server image. The DOCKER_BUILDKIT=1 environment variable enables BuildKit, which provides better caching and faster builds.

Build arguments and options
  • --build-arg max_jobs=<N> — sets the number of parallel compilation jobs for building CUDA kernels. Useful for speeding up builds on multi-core systems.
  • --build-arg nvcc_threads=<N> — controls CUDA compiler threads. Recommended to use a smaller value than max_jobs to avoid excessive memory usage.
  • --build-arg torch_cuda_arch_list="" — if set to empty string, vLLM will detect and build only for the current GPU's compute capability. By default, vLLM builds for all GPU types for wider distribution.
Using precompiled wheels to speed up builds

If you have not changed any C++ or CUDA kernel code, you can use precompiled wheels to significantly reduce Docker build time:

  • Enable precompiled wheels: Add --build-arg VLLM_USE_PRECOMPILED="1" to your build command.
  • How it works: By default, vLLM automatically finds the correct precompiled wheels from the Nightly Builds by using the merge-base commit with the upstream main branch.
  • Specify a commit: To use wheels from a specific commit, add --build-arg VLLM_PRECOMPILED_WHEEL_COMMIT=<commit_hash>.

Example with precompiled wheels and options for fast compilation:

bash
DOCKER_BUILDKIT=1 docker build . \
  --target vllm-openai \
  --tag vllm/vllm-openai \
  --file docker/Dockerfile \
  --build-arg max_jobs=8 \
  --build-arg nvcc_threads=2 \
  --build-arg VLLM_USE_PRECOMPILED="1"
Building with optional dependencies (optional)

vLLM does not include optional dependencies (e.g., audio processing) in the pre-built image to avoid licensing issues. If you need optional dependencies, create a custom Dockerfile that extends the base image:

Example: adding audio optional dependencies

dockerfile
# NOTE: MAKE SURE the version of vLLM matches the base image!
FROM vllm/vllm-openai:0.11.0

# Install audio optional dependencies
RUN uv pip install --system vllm[audio]==0.11.0

Example: using development version of transformers:

dockerfile
FROM vllm/vllm-openai:latest

# Install development version of Transformers from source
RUN uv pip install --system git+https://github.com/huggingface/transformers.git

Build this custom Dockerfile with:

bash
docker build -t my-vllm-custom:latest -f Dockerfile .

Then use it like any other vLLM image:

bash
docker run --rm --gpus all \
  -p 8000:8000 \
  --ipc=host \
  my-vllm-custom:latest \
  --model Qwen/Qwen2.5-1.5B-Instruct
Show full SKILL.md (416 more words)Show less
Building for ARM64/aarch64

A Docker container can be built for ARM64 systems (e.g., NVIDIA Grace-Hopper and Grace-Blackwell). Use the flag --platform "linux/arm64":

bash
DOCKER_BUILDKIT=1 docker build . \
  --target vllm-openai \
  --tag vllm/vllm-openai \
  --file docker/Dockerfile \
  --platform "linux/arm64"

Note: Multiple modules must be compiled, so this process can take longer. Use build arguments like --build-arg max_jobs=8 --build-arg nvcc_threads=2 to speed up the process (ensure max_jobs is substantially larger than nvcc_threads). Monitor memory usage, as parallel jobs can require significant RAM.

For cross-compilation (building ARM64 on an x86_64 host), register QEMU user-static handlers first:

bash
docker run --rm --privileged multiarch/qemu-user-static --reset -p yes

Then use the --platform "linux/arm64" flag in your build command.

Running your custom-built image

After building, run your image just like the pre-built image:

bash
docker run --rm --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --env "HF_TOKEN=$HF_TOKEN" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai \
  --model Qwen/Qwen2.5-1.5B-Instruct

Replace vllm/vllm-openai with the tag you specified during the build (e.g., my-vllm-custom:latest).

Note: --runtime nvidia is deprecated for most environments. Prefer --gpus ... with NVIDIA Container Toolkit. Use --runtime nvidia only for legacy Docker configurations.

Common server flags

  • --model <MODEL_ID> — model to load (HF ID or local path)
  • --port <PORT> — server port (default 8000 for OpenAI-compatible server)
  • --log-level — adjust verbosity
  • You may pass additional engine_args after the image tag; see vLLM docs for tuning options.

Testing the API

After the container starts, make a quick test request against the OpenAI-compatible endpoint:

bash
curl -s http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"Qwen/Qwen2.5-1.5B-Instruct","messages":[{"role":"user","content":"Who are you?"}],"max_tokens":128}'

Security and operational notes

  • Keep HF_TOKEN secret; prefer passing it via environment variables or a secret manager.
  • For production, run behind a reverse proxy (Nginx) with TLS and authentication.
  • Mount only necessary host paths into the container.

Troubleshooting

  • Container can't access GPUs: ensure nvidia-container-toolkit is installed and restart Docker.
  • Model download failures: check HF_TOKEN and network; mount cache directory to persist downloads.
  • Memory / OOM errors: try a smaller model or add more GPU memory; check --shm-size.
  • If the container fails with NCCL library path issues (rare): set VLLM_NCCL_SO_PATH per upstream guidance.
  • Permission issues of current user with Docker: ensure your user is in the docker group or suggest user to add current user to docker group manually following:
bash
# 1. Create docker group if it doesn't exist (may already exist on some systems)
sudo groupadd docker

# 2. Add current user to the docker group (replace $USER with your username if needed)
sudo usermod -aG docker $USER

# 3. Apply the new group membership (you may need to log out and log back in for this to take effect)
newgrp docker

# 4. Verify that the user is in the docker group (output should include docker)
groups $USER
  • Docker pull failed due to a network timeout connecting to Docker Hub: check the network connection and try again. If the issue persists, suggest and show the user how to use a mirror for Docker Hub.
  • Error during downloading models from Hugging Face: check if the HF_TOKEN is passed to the container and is valid. Check if HTTP_PROXY and HTTPS_PROXY are passed to the container if the host is behind a proxy. Also, verify that the model ID is correct and that the model is public or accessible with the provided token.

References

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/vllm-skills/skills/vllm-deploy-docker of vllm-project/vllm-skills.

Open the folder on GitHubat commit c996234

Compare with similar skills

Vllm Deploy Docker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Deploy Docker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Deploy Docker this skillvllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Vllm Serversickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT
Jetson PackageNVIDIA/skills3.5k1 repos~1.8kAutomated safety check: PassApache-2.0
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Ascend Model Adapter for vLLMvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Vllm Server

    sickn33/agentic-awesome-skills

    Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Jetson Package

    NVIDIA/skills

    Official

    Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

    3.5k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-skills

  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    103 GitHub stars~1.5k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Vllm Prefix Cache Bench

    vllm-project/vllm-skills

    This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

    103 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Vllm Deploy Docker

What does Vllm Deploy Docker do?

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. Vllm Deploy Docker is an agent skill from vllm-project/vllm-skills. Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

When should I use Vllm Deploy Docker?

Vllm Deploy Docker fits situations like: tasks that involve LLM inference and serving; tasks that involve Containers.

How do I install Vllm Deploy Docker in Claude Code?

Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-docker in vllm-project/vllm-skills) into .claude/skills/vllm-deploy-docker in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Deploy Docker in Codex?

Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-docker in vllm-project/vllm-skills) into .agents/skills/vllm-deploy-docker in your project. Codex loads it when a task matches its description.

Can I use Vllm Deploy Docker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-deploy-docker, .gemini/skills/vllm-deploy-docker, .github/skills/vllm-deploy-docker and .opencode/skills/vllm-deploy-docker in your project.

What does Vllm Deploy Docker need to run?

Going by SKILL.md and its folder, Vllm Deploy Docker needs the command-line tools its instructions call (docker and curl) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker.

Does Vllm Deploy Docker access the network?

SKILL.md names 4 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.vllm.ai, hub.docker.com and docs.nvidia.com. This is read from the text; nothing was executed.

Is Vllm Deploy Docker safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Vllm Deploy Docker use?

Vllm Deploy Docker is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Deploy Docker use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vllm Deploy Docker?

Skills that share tags, products or a category with Vllm Deploy Docker: Vllm Server (sickn33/agentic-awesome-skills, 47k stars), Jetson Package (NVIDIA/skills, 3.5k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Deploy Docker?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 103 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.

Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.