Vllm Server
sickn33/agentic-awesome-skills
Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-docker --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-docker .claude/skills/vllm-deploy-docker && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-deploy-docker" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-docker into .claude/skills/vllm-deploy-docker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-docker", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-dockerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-docker --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-docker .agents/skills/vllm-deploy-docker && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-deploy-docker" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-docker into .agents/skills/vllm-deploy-docker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-docker", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-docker --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-docker .cursor/skills/vllm-deploy-docker && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-deploy-docker" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-docker into .cursor/skills/vllm-deploy-docker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-docker", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-skills.git --path plugins/vllm-skills/skills/vllm-deploy-docker--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-docker --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-docker .gemini/skills/vllm-deploy-docker && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-deploy-docker" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-docker into .gemini/skills/vllm-deploy-docker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-docker", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-skills vllm-deploy-dockerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-docker .github/skills/vllm-deploy-docker && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-deploy-docker" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-docker into .github/skills/vllm-deploy-docker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-docker", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-docker --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-docker .opencode/skills/vllm-deploy-docker && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-deploy-docker" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-docker into .opencode/skills/vllm-deploy-docker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-docker", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-deploy-dockerDeploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Vllm Deploy Docker is an agent skill from vllm-project/vllm-skills. Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Containers. It works with vLLM, Docker, NVIDIA AI Platform and OpenAI. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockercurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comAlso links to:
docs.vllm.aihub.docker.comdocs.nvidia.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm Deploy Docker loads about 2.5k tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 961 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
sudo groupadd dockersudo usermod -aG docker $USERAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 961 words, ~2,466 tokens.
.claude/skills/vllm-deploy-docker/SKILL.md (or your agent's skills folder).A Claude skill describing how to deploy vLLM with Docker using the official pre-built images or building the image from source supporting NVIDIA GPUs with CUDA. Instructions include NVIDIA CUDA support, example docker run and a minimal docker-compose snippet, recommended flags, and troubleshooting notes. For AMD, Intel, or other accelerators, please refer to the vLLM documentation for alternative deployment methods.
--ipc=host, shared cache mounts, and HF_TOKEN handlingcurl for API testsHF_TOKENRun a vLLM OpenAI-compatible server with GPU access, mounting the HF cache and forwarding port 8000:
docker run --rm --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:latest \
--model Qwen/Qwen2.5-1.5B-Instruct--gpus all exposes all GPUs to the container. Adjust if you need specific GPUs.--ipc=host or an appropriately large --shm-size is recommended so PyTorch and vLLM can share host shared memory.~/.cache/huggingface avoids re-downloading models inside the container.Note: vLLM and this skill recommend using the latest Docker image (
vllm/vllm-openai:latest). For legacy version images, you may refer to the Docker Hub image tags.
You can build and run vLLM from source by using the provided docker/Dockerfile. First, check the hardware of the host machine and ensure you have the necessary dependencies installed (e.g., NVIDIA drivers, CUDA toolkit, Docker with BuildKit support). For ARM64/aarch64 builds, refer to the "Building for ARM64/aarch64" section.
DOCKER_BUILDKIT=1 docker build . \
--target vllm-openai \
--tag vllm/vllm-openai \
--file docker/DockerfileThe --target vllm-openai specifies that you are building the OpenAI-compatible server image. The DOCKER_BUILDKIT=1 environment variable enables BuildKit, which provides better caching and faster builds.
--build-arg max_jobs=<N> — sets the number of parallel compilation jobs for building CUDA kernels. Useful for speeding up builds on multi-core systems.--build-arg nvcc_threads=<N> — controls CUDA compiler threads. Recommended to use a smaller value than max_jobs to avoid excessive memory usage.--build-arg torch_cuda_arch_list="" — if set to empty string, vLLM will detect and build only for the current GPU's compute capability. By default, vLLM builds for all GPU types for wider distribution.If you have not changed any C++ or CUDA kernel code, you can use precompiled wheels to significantly reduce Docker build time:
--build-arg VLLM_USE_PRECOMPILED="1" to your build command.main branch.--build-arg VLLM_PRECOMPILED_WHEEL_COMMIT=<commit_hash>.Example with precompiled wheels and options for fast compilation:
DOCKER_BUILDKIT=1 docker build . \
--target vllm-openai \
--tag vllm/vllm-openai \
--file docker/Dockerfile \
--build-arg max_jobs=8 \
--build-arg nvcc_threads=2 \
--build-arg VLLM_USE_PRECOMPILED="1"vLLM does not include optional dependencies (e.g., audio processing) in the pre-built image to avoid licensing issues. If you need optional dependencies, create a custom Dockerfile that extends the base image:
Example: adding audio optional dependencies
# NOTE: MAKE SURE the version of vLLM matches the base image!
FROM vllm/vllm-openai:0.11.0
# Install audio optional dependencies
RUN uv pip install --system vllm[audio]==0.11.0Example: using development version of transformers:
FROM vllm/vllm-openai:latest
# Install development version of Transformers from source
RUN uv pip install --system git+https://github.com/huggingface/transformers.gitBuild this custom Dockerfile with:
docker build -t my-vllm-custom:latest -f Dockerfile .Then use it like any other vLLM image:
docker run --rm --gpus all \
-p 8000:8000 \
--ipc=host \
my-vllm-custom:latest \
--model Qwen/Qwen2.5-1.5B-InstructA Docker container can be built for ARM64 systems (e.g., NVIDIA Grace-Hopper and Grace-Blackwell). Use the flag --platform "linux/arm64":
DOCKER_BUILDKIT=1 docker build . \
--target vllm-openai \
--tag vllm/vllm-openai \
--file docker/Dockerfile \
--platform "linux/arm64"Note: Multiple modules must be compiled, so this process can take longer. Use build arguments like --build-arg max_jobs=8 --build-arg nvcc_threads=2 to speed up the process (ensure max_jobs is substantially larger than nvcc_threads). Monitor memory usage, as parallel jobs can require significant RAM.
For cross-compilation (building ARM64 on an x86_64 host), register QEMU user-static handlers first:
docker run --rm --privileged multiarch/qemu-user-static --reset -p yesThen use the --platform "linux/arm64" flag in your build command.
After building, run your image just like the pre-built image:
docker run --rm --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=$HF_TOKEN" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai \
--model Qwen/Qwen2.5-1.5B-InstructReplace vllm/vllm-openai with the tag you specified during the build (e.g., my-vllm-custom:latest).
Note:
--runtime nvidiais deprecated for most environments. Prefer--gpus ...with NVIDIA Container Toolkit. Use--runtime nvidiaonly for legacy Docker configurations.
--model <MODEL_ID> — model to load (HF ID or local path)--port <PORT> — server port (default 8000 for OpenAI-compatible server)--log-level — adjust verbosityengine_args after the image tag; see vLLM docs for tuning options.After the container starts, make a quick test request against the OpenAI-compatible endpoint:
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen2.5-1.5B-Instruct","messages":[{"role":"user","content":"Who are you?"}],"max_tokens":128}'HF_TOKEN secret; prefer passing it via environment variables or a secret manager.nvidia-container-toolkit is installed and restart Docker.HF_TOKEN and network; mount cache directory to persist downloads.--shm-size.VLLM_NCCL_SO_PATH per upstream guidance.docker group or suggest user to add current user to docker group manually following:# 1. Create docker group if it doesn't exist (may already exist on some systems)
sudo groupadd docker
# 2. Add current user to the docker group (replace $USER with your username if needed)
sudo usermod -aG docker $USER
# 3. Apply the new group membership (you may need to log out and log back in for this to take effect)
newgrp docker
# 4. Verify that the user is in the docker group (output should include docker)
groups $USERHF_TOKEN is passed to the container and is valid. Check if HTTP_PROXY and HTTPS_PROXY are passed to the container if the host is behind a proxy. Also, verify that the model ID is correct and that the model is public or accessible with the provided token.© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/vllm-skills/skills/vllm-deploy-docker of vllm-project/vllm-skills.
Open the folder on GitHubat commit c996234
Vllm Deploy Docker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm Deploy Docker this skillvllm-project/vllm-skills | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Vllm Serversickn33/agentic-awesome-skills | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Jetson PackageNVIDIA/skills | 3.5k | 1 repos | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Ascend Model Adapter for vLLMvllm-project/vllm-ascend | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 |
sickn33/agentic-awesome-skills
Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.
NVIDIA/skills
Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
vllm-project/vllm-ascend
Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
vllm-project/vllm-skills
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.
vllm-project/vllm-skills
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
vllm-project/vllm-skills
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
vllm-project/vllm-skills
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.
Categories
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. Vllm Deploy Docker is an agent skill from vllm-project/vllm-skills. Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Vllm Deploy Docker fits situations like: tasks that involve LLM inference and serving; tasks that involve Containers.
Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-docker in vllm-project/vllm-skills) into .claude/skills/vllm-deploy-docker in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-docker in vllm-project/vllm-skills) into .agents/skills/vllm-deploy-docker in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-docker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-deploy-docker, .gemini/skills/vllm-deploy-docker, .github/skills/vllm-deploy-docker and .opencode/skills/vllm-deploy-docker in your project.
Going by SKILL.md and its folder, Vllm Deploy Docker needs the command-line tools its instructions call (docker and curl) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker.
SKILL.md names 4 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.vllm.ai, hub.docker.com and docs.nvidia.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Vllm Deploy Docker is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vllm Deploy Docker: Vllm Server (sickn33/agentic-awesome-skills, 47k stars), Jetson Package (NVIDIA/skills, 3.5k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 103 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.
Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.