Vllm Deploy Docker
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.
$ npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sickn33/agentic-awesome-skills vllm-server --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vllm-server .claude/skills/vllm-server && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-server" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/vllm-server into .claude/skills/vllm-server/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-server", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/vllm-serverType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sickn33/agentic-awesome-skills vllm-server --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/vllm-server .agents/skills/vllm-server && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-server" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/vllm-server into .agents/skills/vllm-server/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-server", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sickn33/agentic-awesome-skills vllm-server --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/vllm-server .cursor/skills/vllm-server && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-server" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/vllm-server into .cursor/skills/vllm-server/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-server", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sickn33/agentic-awesome-skills.git --path skills/vllm-server--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sickn33/agentic-awesome-skills vllm-server --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/vllm-server .gemini/skills/vllm-server && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-server" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/vllm-server into .gemini/skills/vllm-server/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-server", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sickn33/agentic-awesome-skills vllm-serverInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/vllm-server .github/skills/vllm-server && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-server" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/vllm-server into .github/skills/vllm-server/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-server", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sickn33/agentic-awesome-skills vllm-server --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/vllm-server .opencode/skills/vllm-server && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-server" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/vllm-server into .opencode/skills/vllm-server/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-server", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-serverDeploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.
Vllm Server is an agent skill from sickn33/agentic-awesome-skills. Deploy and manage vLLM for high-throughput LLM inference. Configure continuous batching, tensor parallelism, quantization, and OpenAI-compatible API endpoints for production LLM serving.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, OpenAI, Docker and Kubernetes. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.
Read from SKILL.md and the folder at commit ec02547. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipcurldockerpythonhuggingface-cliFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, curl and docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HUGGING_FACE_HUB_TOKENHF_TOKENVLLM_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
From compatibility in the SKILL.md frontmatter.
Vllm Server loads about 1.7k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 285 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sickn33/agentic-awesome-skills at commit ec02547, republished under its MIT licence (© sickn33). 285 words, ~1,696 tokens.
.claude/skills/vllm-server/SKILL.md (or your agent's skills folder).Deploy production-grade LLM inference servers with vLLM — the fastest open-source LLM serving engine with PagedAttention and continuous batching.
Use this skill when:
nvidia-container-toolkit for Docker GPU passthrough# Install vLLM
pip install vllm
# Serve a model (OpenAI-compatible API)
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 8000 \
--api-key your-secret-key
# Test the endpoint
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-secret-key" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [{"role": "user", "content": "Hello!"}]
}'docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:latest \
--model meta-llama/Llama-3.1-8B-Instruct \
--api-key your-secret-keyservices:
vllm:
image: vllm/vllm-openai:latest
runtime: nvidia
environment:
- NVIDIA_VISIBLE_DEVICES=all
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
volumes:
- model-cache:/root/.cache/huggingface
ports:
- "8000:8000"
ipc: host
command: >
--model meta-llama/Llama-3.1-70B-Instruct
--tensor-parallel-size 2
--max-model-len 32768
--gpu-memory-utilization 0.90
--api-key ${VLLM_API_KEY}
restart: unless-stopped
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
volumes:
model-cache:# Split one model across 4 GPUs
vllm serve meta-llama/Llama-3.1-70B-Instruct \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.90# AWQ quantization (70B on 2x A100 40GB)
vllm serve casperhansen/llama-3-70b-instruct-awq \
--quantization awq \
--tensor-parallel-size 2
# GPTQ quantization
vllm serve TheBloke/Llama-2-70B-Chat-GPTQ \
--quantization gptq
# FP8 (H100 NVL native)
vllm serve meta-llama/Llama-3.1-405B-Instruct \
--quantization fp8 \
--tensor-parallel-size 8vllm serve meta-llama/Llama-3.1-8B-Instruct \
--enable-auto-tool-choice \
--tool-call-parser llama3_json \
--guided-decoding-backend outlinesvllm serve meta-llama/Llama-3.1-8B-Instruct \
--enable-lora \
--lora-modules sql-lora=/path/to/sql-lora \
code-lora=/path/to/code-lora \
--max-lora-rank 64# Maximize throughput for batch workloads
vllm serve <model> \
--max-num-seqs 256 \ # max concurrent sequences
--max-num-batched-tokens 8192 \ # tokens per batch
--gpu-memory-utilization 0.95 \ # use 95% VRAM
--swap-space 4 # CPU swap (GiB)
# Minimize latency for interactive use
vllm serve <model> \
--max-num-seqs 32 \
--enforce-eager # disable CUDA graph capture# Install benchmark tool
pip install vllm
# Run throughput benchmark
python -m vllm.entrypoints.openai.run_batch \
--model meta-llama/Llama-3.1-8B-Instruct \
--input-file prompts.jsonl \
--output-file results.jsonl
# Benchmark with vllm bench
vllm bench throughput \
--model meta-llama/Llama-3.1-8B-Instruct \
--num-prompts 1000 \
--input-len 512 \
--output-len 128# Check running server stats
curl http://localhost:8000/metrics # Prometheus metrics
# Key metrics to watch:
# vllm:num_requests_running - active requests
# vllm:gpu_cache_usage_perc - KV cache utilization
# vllm:generation_tokens_per_s - throughput
# vllm:time_to_first_token_ms - TTFT latency
# vllm:e2e_request_latency_seconds - end-to-end latency| Issue | Cause | Fix |
|---|---|---|
CUDA out of memory | Model too large for VRAM | Add --quantization awq or reduce --gpu-memory-utilization |
| Slow cold start | Model not cached | Pre-pull with huggingface-cli download <model> |
| Low throughput | Too few concurrent requests | Increase --max-num-seqs |
| KV cache full errors | Context length too long | Set --max-model-len lower |
tokenizer error | Tokenizer mismatch | Use --tokenizer to specify correct tokenizer |
--gpu-memory-utilization 0.90 to leave headroom for CUDA kernels.--revision for reproducible deployments.HF_HUB_OFFLINE=1 in production to prevent unexpected downloads.--enable-chunked-prefill for long-context workloads.gpu_cache_usage_perc — above 95% causes queuing.llm-inference-scaling) - Auto-scaling vLLM deploymentsgpu-server-management) - GPU driver setupllm-gateway) - Load balancing across vLLM instancesllm-cost-optimization) - Cost managementmodel-serving-kubernetes) - K8s deployment© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/vllm-server of sickn33/agentic-awesome-skills.
Open the folder on GitHubat commit ec02547
We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.
Vllm Server next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm Server this skillsickn33/agentic-awesome-skills | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Vllm Deploy Dockervllm-project/vllm-skills | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Vllmmagnus919/agent-skills | 113 | — | ~4.1k | Automated safety check: Notes | MIT | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT |
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
magnus919/agent-skills
Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
sickn33/agentic-awesome-skills
Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.
sickn33/agentic-awesome-skills
Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.
sickn33/agentic-awesome-skills
Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.
sickn33/agentic-awesome-skills
Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.
sickn33/agentic-awesome-skills
Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.
sickn33/agentic-awesome-skills
Drafts and reviews audience-specific content from supplied brand examples, with local scripts for brand voice and SEO diagnostics, channel templates and a content calendar.
Works with
Categories
Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. Vllm Server is an agent skill from sickn33/agentic-awesome-skills. Deploy and manage vLLM for high-throughput LLM inference.
Vllm Server fits situations like: tasks that involve LLM inference and serving.
Run `npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a claude-code`. Or copy the skill folder (skills/vllm-server in sickn33/agentic-awesome-skills) into .claude/skills/vllm-server in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a codex`. Or copy the skill folder (skills/vllm-server in sickn33/agentic-awesome-skills) into .agents/skills/vllm-server in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-server, .gemini/skills/vllm-server, .github/skills/vllm-server and .opencode/skills/vllm-server in your project.
Going by SKILL.md and its folder, Vllm Server needs the command-line tools its instructions call (pip, curl, docker, python and huggingface-cli) and credentials named HUGGING_FACE_HUB_TOKEN, HF_TOKEN and VLLM_API_KEY. Our summary lists: Python 3; Docker; A credential in HUGGING_FACE_HUB_TOKEN; A credential in VLLM_API_KEY. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..
SKILL.md contains no URLs. Its commands use pip, curl and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vllm Server is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vllm Server: Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Vllm (magnus919/agent-skills, 113 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,343 GitHub stars. The repository holds 1,354 skills in this directory. The repository was last updated on October 7, 2026.
Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.