Dstack Prototyping
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-simple --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-simple .claude/skills/vllm-deploy-simple && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-deploy-simple" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-simple into .claude/skills/vllm-deploy-simple/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-simple", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-simpleType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-simple --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-simple .agents/skills/vllm-deploy-simple && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-deploy-simple" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-simple into .agents/skills/vllm-deploy-simple/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-simple", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-simple --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-simple .cursor/skills/vllm-deploy-simple && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-deploy-simple" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-simple into .cursor/skills/vllm-deploy-simple/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-simple", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-skills.git --path plugins/vllm-skills/skills/vllm-deploy-simple--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-simple --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-simple .gemini/skills/vllm-deploy-simple && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-deploy-simple" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-simple into .gemini/skills/vllm-deploy-simple/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-simple", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-skills vllm-deploy-simpleInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-simple .github/skills/vllm-deploy-simple && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-deploy-simple" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-simple into .github/skills/vllm-deploy-simple/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-simple", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-skills vllm-deploy-simple --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/vllm-skills/skills/vllm-deploy-simple .opencode/skills/vllm-deploy-simple && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-deploy-simple" agent skill from https://github.com/vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-deploy-simple into .opencode/skills/vllm-deploy-simple/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-deploy-simple", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-deploy-simpleQuick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
Vllm Deploy Simple is an agent skill from vllm-project/vllm-skills. Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/quickstart.sh`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, OpenAI, Python and NVIDIA AI Platform. The repository describes itself as: Agent skills for vLLM. The licence is Apache-2.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c996234. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
curluvpython3pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl and uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm Deploy Simple loads about 1.6k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 555 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from vllm-project/vllm-skills at commit c996234, republished under its Apache-2.0 licence (© vllm-project). 555 words, ~1,585 tokens.
.claude/skills/vllm-deploy-simple/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.A simple skill to quickly install vLLM, start a server, and validate the OpenAI-compatible API.
This skill provides a streamlined workflow to:
If user did not specify the venv path or asked to deploy in the current environment, create a venv using uv with python 3.12 in the current folder. If uv not found, make a folder in this path and use python to create a virtual environment.
If user did not specify the venv path, model, or port, use default options:
# Default deployment options (--venv "." --model "Qwen/Qwen2.5-1.5B-Instruct" --port 8000 --gpu_memory_utilization 0.8)
scripts/quickstart.shOr with custom options:
# Use custom virtual environment
scripts/quickstart.sh --venv /path/to/venv
# Use custom model and port
scripts/quickstart.sh --model "Qwen/Qwen2.5-1.5B-Instruct" --port 8000
# Use custom GPU memory utilization
scripts/quickstart.sh --gpu_memory_utilization 0.6
# Combine all options
scripts/quickstart.sh --venv /path/to/venv --model "Qwen/Qwen2.5-1.5B-Instruct" --port 8000 --gpu_memory_utilization 0.8This will:
Install vLLM:
scripts/quickstart.sh install
# Or with virtual environment
scripts/quickstart.sh install --venv /path/to/venvStart the server:
scripts/quickstart.sh start
# Or with custom options
scripts/quickstart.sh start --venv /path/to/venv --model "Qwen/Qwen2.5-1.5B-Instruct" --port 8000 --gpu_memory_utilization 0.8Test the API:
scripts/quickstart.sh test
# Or with custom port
scripts/quickstart.sh test --port 8000Stop the server:
scripts/quickstart.sh stop
# Or with virtual environment
scripts/quickstart.sh stop --venv /path/to/venvCheck server status:
scripts/quickstart.sh statusRestart the server:
scripts/quickstart.sh restart
# Or with custom options
scripts/quickstart.sh restart --venv /path/to/venv --port 8000 --gpu_memory_utilization 0.8The script supports the following command-line options:
scripts/quickstart.sh [command] [OPTIONS]
Commands:
install - Install vLLM and dependencies
start - Start the vLLM server
stop - Stop the vLLM server
test - Test the OpenAI-compatible API
status - Show server status
restart - Restart the server
all - Run complete workflow (default)
Options:
--model MODEL Model to use (default: Qwen/Qwen2.5-1.5B-Instruct)
--port PORT Port to run server on (default: 8000)
--venv VENV_PATH Virtual environment path (default: .)
--gpu_memory_utilization VRAM GPU memory utilization (default: 0.8)The script automatically detects your hardware and installs the appropriate vLLM version:
nvidia-smi command/dev/kfd and /dev/dri devicesTPU_NAME environment variable or gcloud commandFor Google TPU, the script installs vllm-tpu instead of the standard vllm package.
The test script sends a simple chat completion request:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [{"role": "user", "content": "Say hello!"}],
"max_tokens": 50
}'Virtual environment not found:
--venv exists and is a valid virtual environmentbin/activate on Linux/macOS or Scripts/activate on Windows)uv venv /path/to/venv (suggested); or with pip: python3 -m venv /path/to/venvServer won't start:
lsof -i :8000nvidia-smi (for NVIDIA) or rocm-smi (for AMD)python -c "import vllm; print(vllm.__version__)"$VENV_PATH/tmp/vllm-server.logAPI returns errors:
cat $VENV_PATH/tmp/vllm-server.logscripts/quickstart.sh statusOut of memory:
--gpu-memory-utilization parameterWrong backend detected:
nvidia-smi is in your PATHTPU_NAME environment variable or install gcloud$VENV_PATH/tmp/vllm-server.log$VENV_PATH/tmp/vllm-server.pid for easy managementuv if available, otherwise falls back to pipscripts/quickstart.sh --port 8080 start --venv /path/to/venv)© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in plugins/vllm-skills/skills/vllm-deploy-simple of vllm-project/vllm-skills.
Open the folder on GitHubat commit c996234
Vllm Deploy Simple next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm Deploy Simple this skillvllm-project/vllm-skills | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~2.8k | Automated safety check: Pass | None | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT |
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
vllm-project/vllm-skills
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.
vllm-project/vllm-skills
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
vllm-project/vllm-skills
This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Works with
Categories
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. Vllm Deploy Simple is an agent skill from vllm-project/vllm-skills. Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
Vllm Deploy Simple fits situations like: tasks that involve LLM inference and serving.
Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a claude-code`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-simple in vllm-project/vllm-skills) into .claude/skills/vllm-deploy-simple in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a codex`. Or copy the skill folder (plugins/vllm-skills/skills/vllm-deploy-simple in vllm-project/vllm-skills) into .agents/skills/vllm-deploy-simple in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-skills --skill vllm-deploy-simple -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-deploy-simple, .gemini/skills/vllm-deploy-simple, .github/skills/vllm-deploy-simple and .opencode/skills/vllm-deploy-simple in your project.
Going by SKILL.md and its folder, Vllm Deploy Simple needs a shell for the scripts in its folder and the command-line tools its instructions call (curl, uv, python3 and python). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use curl and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Vllm Deploy Simple is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vllm Deploy Simple: Dstack Prototyping (dstackai/dstack, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-skills, which has 103 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 3, 2026.
Source: vllm-project/vllm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.