Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

MITAuto-check passedAI & LLM Engineering

Install vLLM Model Serving

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs serving-llms-vllm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/12-inference-serving/vllm .claude/skills/serving-llms-vllm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
serving-llms-vllm
GitHub stars
13k
Used in
5 other repos
Token cost
~2.3k tokens
SKILL.md length
491 words
Files
5 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

  • Standing up a production LLM API behind an OpenAI-style endpoint
  • SKILL.md covers Quick start, Common workflows, When to use vs alternatives and Common issues, plus 3 more sections
  • Calls pip, python and curl
  • Tuning inference latency and throughput for a served model

What it does

vLLM serves language models using PagedAttention, a block-based KV cache, together with continuous batching. The skill covers installing it, running offline inference with LLM and SamplingParams, and starting an OpenAI-compatible server with vllm serve that you query through the OpenAI SDK. The first workflow is a production API rollout: choose server settings by model size, load test with Locust, watch Prometheus metrics such as time to first token and KV cache usage, deploy in Docker, then verify latency, throughput, GPU utilization and the absence of out-of-memory errors.

A second workflow handles offline batch inference over a file of prompts without running a server. The description adds quantization with GPTQ, AWQ and FP8 plus tensor parallelism, and reference files cover optimization, quantization, server deployment and troubleshooting. The excerpt stops partway through the batch workflow.

When your agent uses it

  • Standing up a production LLM API behind an OpenAI-style endpoint
  • Tuning inference latency and throughput for a served model
  • Running a model on a GPU with limited memory using quantization
  • Processing a large file of prompts offline without a server

Example prompts

  • “Serve Llama 3 8B Instruct with vLLM and show how to call it from the OpenAI Python SDK.”
  • “Write a Docker command to run vLLM with the OpenAI-compatible server on one GPU.”
  • “Set up Prometheus monitoring for our vLLM deployment and name the metrics to watch.”
  • “Run prompts.txt through vLLM in offline batch mode and save the outputs.”

Requirements

  • Python with the `vllm` package
  • A GPU for serving models
  • Docker for the container deployment step

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python
    • curl
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.vllm.ai
    • github.com
    • discuss.vllm.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

vLLM Model Serving loads about 2.3k tokens when it runs, and up to ~9k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 491 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 491 words, ~2,257 tokens.

Download SKILL.mdSave it as .claude/skills/serving-llms-vllm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
serving-llms-vllm
description
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
version
1.0.0
author
Orchestra Research
license
MIT
tags
vLLM, Inference Serving, PagedAttention, Continuous Batching, High Throughput, Production, OpenAI API, Quantization, Tensor Parallelism
dependencies
vllm, torch, transformers

vLLM - High-Performance LLM Serving

Quick start

vLLM achieves 24x higher throughput than standard transformers through PagedAttention (block-based KV cache) and continuous batching (mixing prefill/decode requests).

Installation:

bash
pip install vllm

Basic offline inference:

python
from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-3-8B-Instruct")
sampling = SamplingParams(temperature=0.7, max_tokens=256)

outputs = llm.generate(["Explain quantum computing"], sampling)
print(outputs[0].outputs[0].text)

OpenAI-compatible server:

bash
vllm serve meta-llama/Llama-3-8B-Instruct

# Query with OpenAI SDK
python -c "
from openai import OpenAI
client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY')
print(client.chat.completions.create(
    model='meta-llama/Llama-3-8B-Instruct',
    messages=[{'role': 'user', 'content': 'Hello!'}]
).choices[0].message.content)
"

Common workflows

Workflow 1: Production API deployment

Copy this checklist and track progress:

Deployment Progress:
- [ ] Step 1: Configure server settings
- [ ] Step 2: Test with limited traffic
- [ ] Step 3: Enable monitoring
- [ ] Step 4: Deploy to production
- [ ] Step 5: Verify performance metrics

Step 1: Configure server settings

Choose configuration based on your model size:

bash
# For 7B-13B models on single GPU
vllm serve meta-llama/Llama-3-8B-Instruct \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192 \
  --port 8000

# For 30B-70B models with tensor parallelism
vllm serve meta-llama/Llama-2-70b-hf \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.9 \
  --quantization awq \
  --port 8000

# For production with caching and metrics
vllm serve meta-llama/Llama-3-8B-Instruct \
  --gpu-memory-utilization 0.9 \
  --enable-prefix-caching \
  --enable-metrics \
  --metrics-port 9090 \
  --port 8000 \
  --host 0.0.0.0

Step 2: Test with limited traffic

Run load test before production:

bash
# Install load testing tool
pip install locust

# Create test_load.py with sample requests
# Run: locust -f test_load.py --host http://localhost:8000

Verify TTFT (time to first token) < 500ms and throughput > 100 req/sec.

Step 3: Enable monitoring

vLLM exposes Prometheus metrics on port 9090:

bash
curl http://localhost:9090/metrics | grep vllm

Key metrics to monitor:

  • vllm:time_to_first_token_seconds - Latency
  • vllm:num_requests_running - Active requests
  • vllm:gpu_cache_usage_perc - KV cache utilization

Step 4: Deploy to production

Use Docker for consistent deployment:

bash
# Run vLLM in Docker
docker run --gpus all -p 8000:8000 \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3-8B-Instruct \
  --gpu-memory-utilization 0.9 \
  --enable-prefix-caching

Step 5: Verify performance metrics

Check that deployment meets targets:

  • TTFT < 500ms (for short prompts)
  • Throughput > target req/sec
  • GPU utilization > 80%
  • No OOM errors in logs
Workflow 2: Offline batch inference

For processing large datasets without server overhead.

Copy this checklist:

Batch Processing:
- [ ] Step 1: Prepare input data
- [ ] Step 2: Configure LLM engine
- [ ] Step 3: Run batch inference
- [ ] Step 4: Process results

Step 1: Prepare input data

python
# Load prompts from file
prompts = []
with open("prompts.txt") as f:
    prompts = [line.strip() for line in f]

print(f"Loaded {len(prompts)} prompts")

Step 2: Configure LLM engine

python
from vllm import LLM, SamplingParams

llm = LLM(
    model="meta-llama/Llama-3-8B-Instruct",
    tensor_parallel_size=2,  # Use 2 GPUs
    gpu_memory_utilization=0.9,
    max_model_len=4096
)

sampling = SamplingParams(
    temperature=0.7,
    top_p=0.95,
    max_tokens=512,
    stop=["</s>", "\n\n"]
)

Step 3: Run batch inference

vLLM automatically batches requests for efficiency:

python
# Process all prompts in one call
outputs = llm.generate(prompts, sampling)

# vLLM handles batching internally
# No need to manually chunk prompts

Step 4: Process results

python
# Extract generated text
results = []
for output in outputs:
    prompt = output.prompt
    generated = output.outputs[0].text
    results.append({
        "prompt": prompt,
        "generated": generated,
        "tokens": len(output.outputs[0].token_ids)
    })

# Save to file
import json
with open("results.jsonl", "w") as f:
    for result in results:
        f.write(json.dumps(result) + "\n")

print(f"Processed {len(results)} prompts")
Workflow 3: Quantized model serving

Fit large models in limited GPU memory.

Quantization Setup:
- [ ] Step 1: Choose quantization method
- [ ] Step 2: Find or create quantized model
- [ ] Step 3: Launch with quantization flag
- [ ] Step 4: Verify accuracy

Step 1: Choose quantization method

  • AWQ: Best for 70B models, minimal accuracy loss
  • GPTQ: Wide model support, good compression
  • FP8: Fastest on H100 GPUs

Step 2: Find or create quantized model

Use pre-quantized models from HuggingFace:

bash
# Search for AWQ models
# Example: TheBloke/Llama-2-70B-AWQ

Step 3: Launch with quantization flag

bash
# Using pre-quantized model
vllm serve TheBloke/Llama-2-70B-AWQ \
  --quantization awq \
  --tensor-parallel-size 1 \
  --gpu-memory-utilization 0.95

# Results: 70B model in ~40GB VRAM

Step 4: Verify accuracy

Test outputs match expected quality:

python
# Compare quantized vs non-quantized responses
# Verify task-specific performance unchanged

When to use vs alternatives

Use vLLM when:

  • Deploying production LLM APIs (100+ req/sec)
  • Serving OpenAI-compatible endpoints
  • Limited GPU memory but need large models
  • Multi-user applications (chatbots, assistants)
  • Need low latency with high throughput

Use alternatives instead:

  • llama.cpp: CPU/edge inference, single-user
  • HuggingFace transformers: Research, prototyping, one-off generation
  • TensorRT-LLM: NVIDIA-only, need absolute maximum performance
  • Text-Generation-Inference: Already in HuggingFace ecosystem
Show full SKILL.md (192 more words)Show less

Common issues

Issue: Out of memory during model loading

Reduce memory usage:

bash
vllm serve MODEL \
  --gpu-memory-utilization 0.7 \
  --max-model-len 4096

Or use quantization:

bash
vllm serve MODEL --quantization awq

Issue: Slow first token (TTFT > 1 second)

Enable prefix caching for repeated prompts:

bash
vllm serve MODEL --enable-prefix-caching

For long prompts, enable chunked prefill:

bash
vllm serve MODEL --enable-chunked-prefill

Issue: Model not found error

Use --trust-remote-code for custom models:

bash
vllm serve MODEL --trust-remote-code

Issue: Low throughput (<50 req/sec)

Increase concurrent sequences:

bash
vllm serve MODEL --max-num-seqs 512

Check GPU utilization with nvidia-smi - should be >80%.

Issue: Inference slower than expected

Verify tensor parallelism uses power of 2 GPUs:

bash
vllm serve MODEL --tensor-parallel-size 4  # Not 3

Enable speculative decoding for faster generation:

bash
vllm serve MODEL --speculative-model DRAFT_MODEL

Advanced topics

Server deployment patterns: See references/server-deployment.md for Docker, Kubernetes, and load balancing configurations.

Performance optimization: See references/optimization.md for PagedAttention tuning, continuous batching details, and benchmark results.

Quantization guide: See references/quantization.md for AWQ/GPTQ/FP8 setup, model preparation, and accuracy comparisons.

Troubleshooting: See references/troubleshooting.md for detailed error messages, debugging steps, and performance diagnostics.

Hardware requirements

  • Small models (7B-13B): 1x A10 (24GB) or A100 (40GB)
  • Medium models (30B-40B): 2x A100 (40GB) with tensor parallelism
  • Large models (70B+): 4x A100 (40GB) or 2x A100 (80GB), use AWQ/GPTQ

Supported platforms: NVIDIA (primary), AMD ROCm, Intel GPUs, TPUs

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in 12-inference-serving/vllm of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/optimization.md
  • references/quantization.md
  • references/server-deployment.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 773a529

Used in 5 other repositories

We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

vLLM Model Serving next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

vLLM Model Serving compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
vLLM Model Serving this skillOrchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT
Aqua Metricsoracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.0
ML AIgrafana/skills282—~1.3kAutomated safety check: PassApache-2.0
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0
Qdrant Advisorqdrant/skills254—~1.7kAutomated safety check: PassApache-2.0
Vllm Deploy K8svllm-project/vllm-skills102—~2kAutomated safety check: PassApache-2.0

Similar skills

  • Aqua Metrics

    oracle/accelerated-data-science

    Official

    Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

    125 GitHub stars~1.5k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • ML AI

    grafana/skills

    Official

    Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

    282 GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • Qdrant Advisor

    qdrant/skills

    Official

    Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.

    254 GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    102 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about vLLM Model Serving

What does vLLM Model Serving do?

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. vLLM serves language models using PagedAttention, a block-based KV cache, together with continuous batching. The skill covers installing it, running offline inference with LLM and SamplingParams, and starting an OpenAI-compatible server with vllm serve that you query through the OpenAI SDK.

When should I use vLLM Model Serving?

vLLM Model Serving fits situations like: standing up a production LLM API behind an OpenAI-style endpoint; tuning inference latency and throughput for a served model; running a model on a GPU with limited memory using quantization; processing a large file of prompts offline without a server.

How do I install vLLM Model Serving in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a claude-code`. Or copy the skill folder (12-inference-serving/vllm in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/serving-llms-vllm in your project. Claude Code loads it when a task matches its description.

How do I install vLLM Model Serving in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a codex`. Or copy the skill folder (12-inference-serving/vllm in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/serving-llms-vllm in your project. Codex loads it when a task matches its description.

Can I use vLLM Model Serving in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-vllm, .gemini/skills/serving-llms-vllm, .github/skills/serving-llms-vllm and .opencode/skills/serving-llms-vllm in your project.

What does vLLM Model Serving need to run?

Going by SKILL.md and its folder, vLLM Model Serving needs the command-line tools its instructions call (pip, python, curl and docker). Our summary lists: Python with the `vllm` package; A GPU for serving models; Docker for the container deployment step.

Does vLLM Model Serving access the network?

SKILL.md names 3 domains. As links in the text: docs.vllm.ai, github.com and discuss.vllm.ai. This is read from the text; nothing was executed.

Is vLLM Model Serving safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does vLLM Model Serving use?

vLLM Model Serving is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does vLLM Model Serving use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.

What are the alternatives to vLLM Model Serving?

Skills that share tags, products or a category with vLLM Model Serving: Aqua Metrics (oracle/accelerated-data-science, 125 stars), ML AI (grafana/skills, 282 stars), SageMaker Production Defaults (huggingface/skills, 11k stars) and Qdrant Advisor (qdrant/skills, 254 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains vLLM Model Serving?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.