Aqua Metrics
oracle/accelerated-data-science
Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs serving-llms-vllm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/12-inference-serving/vllm .claude/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/vllm into .claude/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/vllmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs serving-llms-vllm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/12-inference-serving/vllm .agents/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/vllm into .agents/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs serving-llms-vllm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/12-inference-serving/vllm .cursor/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/vllm into .cursor/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 12-inference-serving/vllm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs serving-llms-vllm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/12-inference-serving/vllm .gemini/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/vllm into .gemini/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs serving-llms-vllmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/12-inference-serving/vllm .github/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/vllm into .github/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs serving-llms-vllm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/12-inference-serving/vllm .opencode/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/12-inference-serving/vllm into .opencode/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
serving-llms-vllmDeploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
vLLM serves language models using PagedAttention, a block-based KV cache, together with continuous batching. The skill covers installing it, running offline inference with LLM and SamplingParams, and starting an OpenAI-compatible server with vllm serve that you query through the OpenAI SDK. The first workflow is a production API rollout: choose server settings by model size, load test with Locust, watch Prometheus metrics such as time to first token and KV cache usage, deploy in Docker, then verify latency, throughput, GPU utilization and the absence of out-of-memory errors.
A second workflow handles offline batch inference over a file of prompts without running a server. The description adds quantization with GPTQ, AWQ and FP8 plus tensor parallelism, and reference files cover optimization, quantization, server deployment and troubleshooting. The excerpt stops partway through the batch workflow.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pippythoncurldockerFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.vllm.aigithub.comdiscuss.vllm.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
vLLM Model Serving loads about 2.3k tokens when it runs, and up to ~9k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 491 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 491 words, ~2,257 tokens.
.claude/skills/serving-llms-vllm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.vLLM achieves 24x higher throughput than standard transformers through PagedAttention (block-based KV cache) and continuous batching (mixing prefill/decode requests).
Installation:
pip install vllmBasic offline inference:
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Llama-3-8B-Instruct")
sampling = SamplingParams(temperature=0.7, max_tokens=256)
outputs = llm.generate(["Explain quantum computing"], sampling)
print(outputs[0].outputs[0].text)OpenAI-compatible server:
vllm serve meta-llama/Llama-3-8B-Instruct
# Query with OpenAI SDK
python -c "
from openai import OpenAI
client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY')
print(client.chat.completions.create(
model='meta-llama/Llama-3-8B-Instruct',
messages=[{'role': 'user', 'content': 'Hello!'}]
).choices[0].message.content)
"Copy this checklist and track progress:
Deployment Progress:
- [ ] Step 1: Configure server settings
- [ ] Step 2: Test with limited traffic
- [ ] Step 3: Enable monitoring
- [ ] Step 4: Deploy to production
- [ ] Step 5: Verify performance metricsStep 1: Configure server settings
Choose configuration based on your model size:
# For 7B-13B models on single GPU
vllm serve meta-llama/Llama-3-8B-Instruct \
--gpu-memory-utilization 0.9 \
--max-model-len 8192 \
--port 8000
# For 30B-70B models with tensor parallelism
vllm serve meta-llama/Llama-2-70b-hf \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.9 \
--quantization awq \
--port 8000
# For production with caching and metrics
vllm serve meta-llama/Llama-3-8B-Instruct \
--gpu-memory-utilization 0.9 \
--enable-prefix-caching \
--enable-metrics \
--metrics-port 9090 \
--port 8000 \
--host 0.0.0.0Step 2: Test with limited traffic
Run load test before production:
# Install load testing tool
pip install locust
# Create test_load.py with sample requests
# Run: locust -f test_load.py --host http://localhost:8000Verify TTFT (time to first token) < 500ms and throughput > 100 req/sec.
Step 3: Enable monitoring
vLLM exposes Prometheus metrics on port 9090:
curl http://localhost:9090/metrics | grep vllmKey metrics to monitor:
vllm:time_to_first_token_seconds - Latencyvllm:num_requests_running - Active requestsvllm:gpu_cache_usage_perc - KV cache utilizationStep 4: Deploy to production
Use Docker for consistent deployment:
# Run vLLM in Docker
docker run --gpus all -p 8000:8000 \
vllm/vllm-openai:latest \
--model meta-llama/Llama-3-8B-Instruct \
--gpu-memory-utilization 0.9 \
--enable-prefix-cachingStep 5: Verify performance metrics
Check that deployment meets targets:
For processing large datasets without server overhead.
Copy this checklist:
Batch Processing:
- [ ] Step 1: Prepare input data
- [ ] Step 2: Configure LLM engine
- [ ] Step 3: Run batch inference
- [ ] Step 4: Process resultsStep 1: Prepare input data
# Load prompts from file
prompts = []
with open("prompts.txt") as f:
prompts = [line.strip() for line in f]
print(f"Loaded {len(prompts)} prompts")Step 2: Configure LLM engine
from vllm import LLM, SamplingParams
llm = LLM(
model="meta-llama/Llama-3-8B-Instruct",
tensor_parallel_size=2, # Use 2 GPUs
gpu_memory_utilization=0.9,
max_model_len=4096
)
sampling = SamplingParams(
temperature=0.7,
top_p=0.95,
max_tokens=512,
stop=["</s>", "\n\n"]
)Step 3: Run batch inference
vLLM automatically batches requests for efficiency:
# Process all prompts in one call
outputs = llm.generate(prompts, sampling)
# vLLM handles batching internally
# No need to manually chunk promptsStep 4: Process results
# Extract generated text
results = []
for output in outputs:
prompt = output.prompt
generated = output.outputs[0].text
results.append({
"prompt": prompt,
"generated": generated,
"tokens": len(output.outputs[0].token_ids)
})
# Save to file
import json
with open("results.jsonl", "w") as f:
for result in results:
f.write(json.dumps(result) + "\n")
print(f"Processed {len(results)} prompts")Fit large models in limited GPU memory.
Quantization Setup:
- [ ] Step 1: Choose quantization method
- [ ] Step 2: Find or create quantized model
- [ ] Step 3: Launch with quantization flag
- [ ] Step 4: Verify accuracyStep 1: Choose quantization method
Step 2: Find or create quantized model
Use pre-quantized models from HuggingFace:
# Search for AWQ models
# Example: TheBloke/Llama-2-70B-AWQStep 3: Launch with quantization flag
# Using pre-quantized model
vllm serve TheBloke/Llama-2-70B-AWQ \
--quantization awq \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.95
# Results: 70B model in ~40GB VRAMStep 4: Verify accuracy
Test outputs match expected quality:
# Compare quantized vs non-quantized responses
# Verify task-specific performance unchangedUse vLLM when:
Use alternatives instead:
Issue: Out of memory during model loading
Reduce memory usage:
vllm serve MODEL \
--gpu-memory-utilization 0.7 \
--max-model-len 4096Or use quantization:
vllm serve MODEL --quantization awqIssue: Slow first token (TTFT > 1 second)
Enable prefix caching for repeated prompts:
vllm serve MODEL --enable-prefix-cachingFor long prompts, enable chunked prefill:
vllm serve MODEL --enable-chunked-prefillIssue: Model not found error
Use --trust-remote-code for custom models:
vllm serve MODEL --trust-remote-codeIssue: Low throughput (<50 req/sec)
Increase concurrent sequences:
vllm serve MODEL --max-num-seqs 512Check GPU utilization with nvidia-smi - should be >80%.
Issue: Inference slower than expected
Verify tensor parallelism uses power of 2 GPUs:
vllm serve MODEL --tensor-parallel-size 4 # Not 3Enable speculative decoding for faster generation:
vllm serve MODEL --speculative-model DRAFT_MODELServer deployment patterns: See references/server-deployment.md for Docker, Kubernetes, and load balancing configurations.
Performance optimization: See references/optimization.md for PagedAttention tuning, continuous batching details, and benchmark results.
Quantization guide: See references/quantization.md for AWQ/GPTQ/FP8 setup, model preparation, and accuracy comparisons.
Troubleshooting: See references/troubleshooting.md for detailed error messages, debugging steps, and performance diagnostics.
Supported platforms: NVIDIA (primary), AMD ROCm, Intel GPUs, TPUs
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in 12-inference-serving/vllm of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
vLLM Model Serving next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| vLLM Model Serving this skillOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Aqua Metricsoracle/accelerated-data-science | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | |
| ML AIgrafana/skills | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Production Defaultshuggingface/skills | 11k | 1 repos | ~6.9k | Automated safety check: Pass | Apache-2.0 | |
| Qdrant Advisorqdrant/skills | 254 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Vllm Deploy K8svllm-project/vllm-skills | 102 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
oracle/accelerated-data-science
Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.
grafana/skills
Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…
huggingface/skills
Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.
qdrant/skills
Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
Works with
Categories
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. vLLM serves language models using PagedAttention, a block-based KV cache, together with continuous batching. The skill covers installing it, running offline inference with LLM and SamplingParams, and starting an OpenAI-compatible server with vllm serve that you query through the OpenAI SDK.
vLLM Model Serving fits situations like: standing up a production LLM API behind an OpenAI-style endpoint; tuning inference latency and throughput for a served model; running a model on a GPU with limited memory using quantization; processing a large file of prompts offline without a server.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a claude-code`. Or copy the skill folder (12-inference-serving/vllm in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/serving-llms-vllm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a codex`. Or copy the skill folder (12-inference-serving/vllm in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/serving-llms-vllm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill serving-llms-vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-vllm, .gemini/skills/serving-llms-vllm, .github/skills/serving-llms-vllm and .opencode/skills/serving-llms-vllm in your project.
Going by SKILL.md and its folder, vLLM Model Serving needs the command-line tools its instructions call (pip, python, curl and docker). Our summary lists: Python with the `vllm` package; A GPU for serving models; Docker for the container deployment step.
SKILL.md names 3 domains. As links in the text: docs.vllm.ai, github.com and discuss.vllm.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
vLLM Model Serving is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with vLLM Model Serving: Aqua Metrics (oracle/accelerated-data-science, 125 stars), ML AI (grafana/skills, 282 stars), SageMaker Production Defaults (huggingface/skills, 11k stars) and Qdrant Advisor (qdrant/skills, 254 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.