Agent skill

Vllm Server

by sickn33 in sickn33/agentic-awesome-skills

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

MITAuto-check passedAI & LLM Engineering

Install Vllm Server

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills vllm-server --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vllm-server .claude/skills/vllm-server && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-server
GitHub stars
47k
Used in
2 other repos
Token cost
~1.7k tokens
SKILL.md length
285 words
Files
1
Skills in repo
1,354
Repo updated
First seen
Licence
MIT

At a glance

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers When to Use This Skill, Prerequisites, Quick Start and Docker Deployment, plus 9 more sections
  • Calls pip, curl and docker; needs HUGGING_FACE_HUB_TOKEN and HF_TOKEN

What it does

Vllm Server is an agent skill from sickn33/agentic-awesome-skills. Deploy and manage vLLM for high-throughput LLM inference. Configure continuous batching, tensor parallelism, quantization, and OpenAI-compatible API endpoints for production LLM serving.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, OpenAI, Docker and Kubernetes. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/vllm-server”

Requirements

  • Python 3
  • Docker
  • A credential in HUGGING_FACE_HUB_TOKEN
  • A credential in VLLM_API_KEY
  • Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit ec02547. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • curl
    • docker
    • python
    • huggingface-cli

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, curl and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HUGGING_FACE_HUB_TOKEN
    • HF_TOKEN
    • VLLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

Vllm Server loads about 1.7k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 285 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit ec02547, republished under its MIT licence (© sickn33). 285 words, ~1,696 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-server/SKILL.md (or your agent's skills folder).
name
vllm-server
description
Deploy and manage vLLM for high-throughput LLM inference. Configure continuous batching, tensor parallelism, quantization, and OpenAI-compatible API endpoints for production LLM serving.
compatibility
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

vLLM Server Management

Deploy production-grade LLM inference servers with vLLM — the fastest open-source LLM serving engine with PagedAttention and continuous batching.

When to Use This Skill

Use this skill when:

  • Serving open-source LLMs (Llama, Mistral, Qwen, Gemma) at scale
  • Building an OpenAI-compatible API endpoint for self-hosted models
  • Optimizing LLM throughput and latency for production traffic
  • Running multi-GPU inference with tensor or pipeline parallelism
  • Deploying quantized models to reduce GPU memory requirements

Prerequisites

  • NVIDIA GPU(s) with CUDA 12.1+ (A100/H100 recommended for production)
  • Docker or Python 3.9+ with pip
  • 40GB+ VRAM for 70B models; 8GB+ for 7B models
  • nvidia-container-toolkit for Docker GPU passthrough

Quick Start

bash
# Install vLLM
pip install vllm

# Serve a model (OpenAI-compatible API)
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --host 0.0.0.0 \
  --port 8000 \
  --api-key your-secret-key

# Test the endpoint
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-secret-key" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Docker Deployment

bash
docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --api-key your-secret-key

Docker Compose (Production)

yaml
services:
  vllm:
    image: vllm/vllm-openai:latest
    runtime: nvidia
    environment:
      - NVIDIA_VISIBLE_DEVICES=all
      - HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
    volumes:
      - model-cache:/root/.cache/huggingface
    ports:
      - "8000:8000"
    ipc: host
    command: >
      --model meta-llama/Llama-3.1-70B-Instruct
      --tensor-parallel-size 2
      --max-model-len 32768
      --gpu-memory-utilization 0.90
      --api-key ${VLLM_API_KEY}
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      timeout: 10s
      retries: 3

volumes:
  model-cache:

Key Configuration Options

Multi-GPU Tensor Parallelism
bash
# Split one model across 4 GPUs
vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.90
Quantization (Lower VRAM)
bash
# AWQ quantization (70B on 2x A100 40GB)
vllm serve casperhansen/llama-3-70b-instruct-awq \
  --quantization awq \
  --tensor-parallel-size 2

# GPTQ quantization
vllm serve TheBloke/Llama-2-70B-Chat-GPTQ \
  --quantization gptq

# FP8 (H100 NVL native)
vllm serve meta-llama/Llama-3.1-405B-Instruct \
  --quantization fp8 \
  --tensor-parallel-size 8
Structured Output & Tools
bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-auto-tool-choice \
  --tool-call-parser llama3_json \
  --guided-decoding-backend outlines
LoRA Adapters
bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-lora \
  --lora-modules sql-lora=/path/to/sql-lora \
                 code-lora=/path/to/code-lora \
  --max-lora-rank 64

Performance Tuning

bash
# Maximize throughput for batch workloads
vllm serve <model> \
  --max-num-seqs 256 \          # max concurrent sequences
  --max-num-batched-tokens 8192 \ # tokens per batch
  --gpu-memory-utilization 0.95 \ # use 95% VRAM
  --swap-space 4                  # CPU swap (GiB)

# Minimize latency for interactive use
vllm serve <model> \
  --max-num-seqs 32 \
  --enforce-eager              # disable CUDA graph capture

Benchmarking

bash
# Install benchmark tool
pip install vllm

# Run throughput benchmark
python -m vllm.entrypoints.openai.run_batch \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --input-file prompts.jsonl \
  --output-file results.jsonl

# Benchmark with vllm bench
vllm bench throughput \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --num-prompts 1000 \
  --input-len 512 \
  --output-len 128

Monitoring

bash
# Check running server stats
curl http://localhost:8000/metrics  # Prometheus metrics

# Key metrics to watch:
# vllm:num_requests_running       - active requests
# vllm:gpu_cache_usage_perc       - KV cache utilization
# vllm:generation_tokens_per_s    - throughput
# vllm:time_to_first_token_ms     - TTFT latency
# vllm:e2e_request_latency_seconds - end-to-end latency

Common Issues

IssueCauseFix
CUDA out of memoryModel too large for VRAMAdd --quantization awq or reduce --gpu-memory-utilization
Slow cold startModel not cachedPre-pull with huggingface-cli download <model>
Low throughputToo few concurrent requestsIncrease --max-num-seqs
KV cache full errorsContext length too longSet --max-model-len lower
tokenizer errorTokenizer mismatchUse --tokenizer to specify correct tokenizer

Best Practices

  • Use --gpu-memory-utilization 0.90 to leave headroom for CUDA kernels.
  • Pin model versions with --revision for reproducible deployments.
  • Set HF_HUB_OFFLINE=1 in production to prevent unexpected downloads.
  • Use AWQ or GPTQ quantization before tensor parallelism — lower VRAM first.
  • Enable --enable-chunked-prefill for long-context workloads.
  • Monitor gpu_cache_usage_perc — above 95% causes queuing.
  • llm-inference-scaling (llm-inference-scaling) - Auto-scaling vLLM deployments
  • gpu-server-management (gpu-server-management) - GPU driver setup
  • llm-gateway (llm-gateway) - Load balancing across vLLM instances
  • llm-cost-optimization (llm-cost-optimization) - Cost management
  • model-serving-kubernetes (model-serving-kubernetes) - K8s deployment

Limitations

  • Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
  • Docs-only import: upstream scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/vllm-server of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit ec02547

Used in 2 other repositories

We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Vllm Server next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Server compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Server this skillsickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Vllmmagnus919/agent-skills113—~4.1kAutomated safety check: NotesMIT
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Vllm

    magnus919/agent-skills

    Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

    113 GitHub stars~4.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,354 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Content Creator

    sickn33/agentic-awesome-skills

    Drafts and reviews audience-specific content from supplied brand examples, with local scripts for brand voice and SEO diagnostics, channel templates and a content calendar.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed

Questions about Vllm Server

What does Vllm Server do?

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. Vllm Server is an agent skill from sickn33/agentic-awesome-skills. Deploy and manage vLLM for high-throughput LLM inference.

When should I use Vllm Server?

Vllm Server fits situations like: tasks that involve LLM inference and serving.

How do I install Vllm Server in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a claude-code`. Or copy the skill folder (skills/vllm-server in sickn33/agentic-awesome-skills) into .claude/skills/vllm-server in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Server in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a codex`. Or copy the skill folder (skills/vllm-server in sickn33/agentic-awesome-skills) into .agents/skills/vllm-server in your project. Codex loads it when a task matches its description.

Can I use Vllm Server in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill vllm-server -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-server, .gemini/skills/vllm-server, .github/skills/vllm-server and .opencode/skills/vllm-server in your project.

What does Vllm Server need to run?

Going by SKILL.md and its folder, Vllm Server needs the command-line tools its instructions call (pip, curl, docker, python and huggingface-cli) and credentials named HUGGING_FACE_HUB_TOKEN, HF_TOKEN and VLLM_API_KEY. Our summary lists: Python 3; Docker; A credential in HUGGING_FACE_HUB_TOKEN; A credential in VLLM_API_KEY. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..

Does Vllm Server access the network?

SKILL.md contains no URLs. Its commands use pip, curl and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vllm Server safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Server use?

Vllm Server is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Server use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vllm Server?

Skills that share tags, products or a category with Vllm Server: Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Vllm (magnus919/agent-skills, 113 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Server?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,343 GitHub stars. The repository holds 1,354 skills in this directory. The repository was last updated on October 7, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.