Agent skill

Serving LLMs Vllm

by Luciole-Studio in Luciole-Studio/Misaka-Agent

vLLM: high-throughput LLM serving, OpenAI API, quantization.

MITAuto-check passedAI & LLM Engineering

Install Serving LLMs Vllm

skills CLI
$ npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Luciole-Studio/Misaka-Agent serving-llms-vllm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm .claude/skills/serving-llms-vllm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
serving-llms-vllm
GitHub stars
158
Used in
2 other repos
Token cost
~2.3k tokens
SKILL.md length
534 words
Files
5 (incl. references)
Skills in repo
77
Repo updated
First seen
Licence
MIT

At a glance

vLLM: high-throughput LLM serving, OpenAI API, quantization.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers When to use, Quick start, Common workflows and When to use vs alternatives, plus 4 more sections
  • Calls pip, python and curl

What it does

Serving LLMs Vllm is an agent skill from Luciole-Studio/Misaka-Agent. vLLM: high-throughput LLM serving, OpenAI API, quantization.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/optimization.md`, `references/quantization.md` and `references/server-deployment.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and OpenAI. The repository describes itself as: A multi-agent research system for the humanities and social sciences. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/serving-llms-vllm”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 3bcf7a3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python
    • curl
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.vllm.ai
    • github.com
    • discuss.vllm.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Serving LLMs Vllm loads about 2.3k tokens when it runs, and up to ~9k if it reads all its reference files. Until then it costs about 20 tokens; SKILL.md has 534 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Luciole-Studio/Misaka-Agent at commit 3bcf7a3, republished under its MIT licence (© Luciole-Studio). 534 words, ~2,334 tokens.

Download SKILL.mdSave it as .claude/skills/serving-llms-vllm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
serving-llms-vllm
description
vLLM: high-throughput LLM serving, OpenAI API, quantization.
version
1.0.1
author
Orchestra Research
license
MIT
dependencies
vllm, torch, transformers
platforms
linux, macos

vLLM - High-Performance LLM Serving

When to use

Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

Quick start

vLLM achieves 24x higher throughput than standard transformers through PagedAttention (block-based KV cache) and continuous batching (mixing prefill/decode requests).

Installation:

bash
pip install vllm

Basic offline inference:

python
from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct")
sampling = SamplingParams(temperature=0.7, max_tokens=256)

outputs = llm.generate(["Explain quantum computing"], sampling)
print(outputs[0].outputs[0].text)

OpenAI-compatible server:

bash
vllm serve meta-llama/Meta-Llama-3-8B-Instruct

# Query with OpenAI SDK
python -c "
from openai import OpenAI
client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY')
print(client.chat.completions.create(
    model='meta-llama/Meta-Llama-3-8B-Instruct',
    messages=[{'role': 'user', 'content': 'Hello!'}]
).choices[0].message.content)
"

Common workflows

Workflow 1: Production API deployment

Copy this checklist and track progress:

Deployment Progress:
- [ ] Step 1: Configure server settings
- [ ] Step 2: Test with limited traffic
- [ ] Step 3: Enable monitoring
- [ ] Step 4: Deploy to production
- [ ] Step 5: Verify performance metrics

Step 1: Configure server settings

Choose configuration based on your model size:

bash
# For 7B-13B models on single GPU
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192 \
  --port 8000

# For 30B-70B models with tensor parallelism
vllm serve meta-llama/Meta-Llama-3-70B-Instruct \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.9 \
  --quantization awq \
  --port 8000

# For production with caching (Prometheus metrics are exposed
# automatically at /metrics on the API port)
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
  --gpu-memory-utilization 0.9 \
  --enable-prefix-caching \
  --port 8000 \
  --host 0.0.0.0

Step 2: Test with limited traffic

Run load test before production:

bash
# Install load testing tool
pip install locust

# Create test_load.py with sample requests
# Run: locust -f test_load.py --host http://localhost:8000

Verify TTFT (time to first token) < 500ms and throughput > 100 req/sec.

Step 3: Enable monitoring

vLLM exposes Prometheus metrics at /metrics on the API port (default 8000):

bash
curl http://localhost:8000/metrics | grep vllm

Key metrics to monitor:

  • vllm:time_to_first_token_seconds - Latency
  • vllm:num_requests_running - Active requests
  • vllm:gpu_cache_usage_perc - KV cache utilization

Step 4: Deploy to production

Use Docker for consistent deployment:

bash
# Run vLLM in Docker
docker run --gpus all -p 8000:8000 \
  vllm/vllm-openai:latest \
  --model meta-llama/Meta-Llama-3-8B-Instruct \
  --gpu-memory-utilization 0.9 \
  --enable-prefix-caching

Step 5: Verify performance metrics

Check that deployment meets targets:

  • TTFT < 500ms (for short prompts)
  • Throughput > target req/sec
  • GPU utilization > 80%
  • No OOM errors in logs
Workflow 2: Offline batch inference

For processing large datasets without server overhead.

Copy this checklist:

Batch Processing:
- [ ] Step 1: Prepare input data
- [ ] Step 2: Configure LLM engine
- [ ] Step 3: Run batch inference
- [ ] Step 4: Process results

Step 1: Prepare input data

python
# Load prompts from file
prompts = []
with open("prompts.txt") as f:
    prompts = [line.strip() for line in f]

print(f"Loaded {len(prompts)} prompts")

Step 2: Configure LLM engine

python
from vllm import LLM, SamplingParams

llm = LLM(
    model="meta-llama/Meta-Llama-3-8B-Instruct",
    tensor_parallel_size=2,  # Use 2 GPUs
    gpu_memory_utilization=0.9,
    max_model_len=4096
)

sampling = SamplingParams(
    temperature=0.7,
    top_p=0.95,
    max_tokens=512,
    stop=["</s>", "\n\n"]
)

Step 3: Run batch inference

vLLM automatically batches requests for efficiency:

python
# Process all prompts in one call
outputs = llm.generate(prompts, sampling)

# vLLM handles batching internally
# No need to manually chunk prompts

Step 4: Process results

python
# Extract generated text
results = []
for output in outputs:
    prompt = output.prompt
    generated = output.outputs[0].text
    results.append({
        "prompt": prompt,
        "generated": generated,
        "tokens": len(output.outputs[0].token_ids)
    })

# Save to file
import json
with open("results.jsonl", "w") as f:
    for result in results:
        f.write(json.dumps(result) + "\n")

print(f"Processed {len(results)} prompts")
Workflow 3: Quantized model serving

Fit large models in limited GPU memory.

Quantization Setup:
- [ ] Step 1: Choose quantization method
- [ ] Step 2: Find or create quantized model
- [ ] Step 3: Launch with quantization flag
- [ ] Step 4: Verify accuracy

Step 1: Choose quantization method

  • AWQ: Best for 70B models, minimal accuracy loss
  • GPTQ: Wide model support, good compression
  • FP8: Fastest on H100 GPUs

Step 2: Find or create quantized model

Use pre-quantized models from HuggingFace:

bash
# Search for AWQ models
# Example: TheBloke/Llama-2-70B-AWQ

Step 3: Launch with quantization flag

bash
# Using pre-quantized model
vllm serve TheBloke/Llama-2-70B-AWQ \
  --quantization awq \
  --tensor-parallel-size 1 \
  --gpu-memory-utilization 0.95

# Results: 70B model in ~40GB VRAM

Step 4: Verify accuracy

Test outputs match expected quality:

python
# Compare quantized vs non-quantized responses
# Verify task-specific performance unchanged

When to use vs alternatives

Use vLLM when:

  • Deploying production LLM APIs (100+ req/sec)
  • Serving OpenAI-compatible endpoints
  • Limited GPU memory but need large models
  • Multi-user applications (chatbots, assistants)
  • Need low latency with high throughput

Use alternatives instead:

  • llama.cpp: CPU/edge inference, single-user
  • HuggingFace transformers: Research, prototyping, one-off generation
  • TensorRT-LLM: NVIDIA-only, need absolute maximum performance
  • Text-Generation-Inference: Already in HuggingFace ecosystem
Show full SKILL.md (203 more words)Show less

Common issues

Issue: Out of memory during model loading

Reduce memory usage:

bash
vllm serve MODEL \
  --gpu-memory-utilization 0.7 \
  --max-model-len 4096

Or use quantization:

bash
vllm serve MODEL --quantization awq

Issue: Slow first token (TTFT > 1 second)

Enable prefix caching for repeated prompts:

bash
vllm serve MODEL --enable-prefix-caching

For long prompts, enable chunked prefill:

bash
vllm serve MODEL --enable-chunked-prefill

Issue: Model not found error

Use --trust-remote-code for custom models:

bash
vllm serve MODEL --trust-remote-code

Issue: Low throughput (<50 req/sec)

Increase concurrent sequences:

bash
vllm serve MODEL --max-num-seqs 512

Check GPU utilization with nvidia-smi - should be >80%.

Issue: Inference slower than expected

Verify tensor parallelism uses power of 2 GPUs:

bash
vllm serve MODEL --tensor-parallel-size 4  # Not 3

Enable speculative decoding for faster generation (pass config as JSON; --speculative-model was removed in favor of --speculative-config):

bash
vllm serve MODEL \
  --speculative-config '{"model": "DRAFT_MODEL", "num_speculative_tokens": 5, "method": "draft_model"}'

Advanced topics

Server deployment patterns: See references/server-deployment.md for Docker, Kubernetes, and load balancing configurations.

Performance optimization: See references/optimization.md for PagedAttention tuning, continuous batching details, and benchmark results.

Quantization guide: See references/quantization.md for AWQ/GPTQ/FP8 setup, model preparation, and accuracy comparisons.

Troubleshooting: See references/troubleshooting.md for detailed error messages, debugging steps, and performance diagnostics.

Hardware requirements

  • Small models (7B-13B): 1x A10 (24GB) or A100 (40GB)
  • Medium models (30B-40B): 2x A100 (40GB) with tensor parallelism
  • Large models (70B+): 4x A100 (40GB) or 2x A100 (80GB), use AWQ/GPTQ

Supported platforms: NVIDIA (primary), AMD ROCm, Intel GPUs, TPUs

Resources

© Luciole-Studio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm of Luciole-Studio/Misaka-Agent.

  • SKILL.md
  • references/optimization.md
  • references/quantization.md
  • references/server-deployment.md
  • references/troubleshooting.md

Open the folder on GitHubat commit 3bcf7a3

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Luciole-Studio/Misaka-Agent, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Serving LLMs Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Serving LLMs Vllm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Serving LLMs Vllm this skillLuciole-Studio/Misaka-Agent1582 repos~2.3kAutomated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT
Vllm Bench Random Syntheticvllm-project/vllm-skills103—~1.5kAutomated safety check: PassApache-2.0
Vllm Bench Servevllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    103 GitHub stars~1.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from Luciole-Studio/Misaka-Agent

All 77 skills in this repo
  • Kanban Video Orchestrator

    Luciole-Studio/Misaka-Agent

    Plan and run multi-agent video production pipelines. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check: notes
  • Ast Grep

    Luciole-Studio/Misaka-Agent

    AST-aware structural code search and rewrite via ast-grep. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Drug Discovery

    Luciole-Studio/Misaka-Agent

    Drug discovery: ChEMBL search, drug-likeness, interactions. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Fitness Nutrition

    Luciole-Studio/Misaka-Agent

    Workout planning, macros, and body metrics via wger/USDA. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Hyperframes

    Luciole-Studio/Misaka-Agent

    Render MP4/WebM videos from HTML compositions. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Osint Investigation

    Luciole-Studio/Misaka-Agent

    Follow the money via public records and sanctions data. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check passed

Works with

Questions about Serving LLMs Vllm

What does Serving LLMs Vllm do?

vLLM: high-throughput LLM serving, OpenAI API, quantization. Serving LLMs Vllm is an agent skill from Luciole-Studio/Misaka-Agent. vLLM: high-throughput LLM serving, OpenAI API, quantization.

When should I use Serving LLMs Vllm?

Serving LLMs Vllm fits situations like: tasks that involve LLM inference and serving.

How do I install Serving LLMs Vllm in Claude Code?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a claude-code`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm in Luciole-Studio/Misaka-Agent) into .claude/skills/serving-llms-vllm in your project. Claude Code loads it when a task matches its description.

How do I install Serving LLMs Vllm in Codex?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a codex`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm in Luciole-Studio/Misaka-Agent) into .agents/skills/serving-llms-vllm in your project. Codex loads it when a task matches its description.

Can I use Serving LLMs Vllm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-vllm, .gemini/skills/serving-llms-vllm, .github/skills/serving-llms-vllm and .opencode/skills/serving-llms-vllm in your project.

What does Serving LLMs Vllm need to run?

Going by SKILL.md and its folder, Serving LLMs Vllm needs the command-line tools its instructions call (pip, python, curl and docker). Our summary lists: Python 3; Docker.

Does Serving LLMs Vllm access the network?

SKILL.md names 3 domains. As links in the text: docs.vllm.ai, github.com and discuss.vllm.ai. This is read from the text; nothing was executed.

Is Serving LLMs Vllm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Serving LLMs Vllm use?

Serving LLMs Vllm is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Serving LLMs Vllm use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.

What are the alternatives to Serving LLMs Vllm?

Skills that share tags, products or a category with Serving LLMs Vllm: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Bench Random Synthetic (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Serving LLMs Vllm?

Luciole-Studio (a GitHub organization) maintains it in Luciole-Studio/Misaka-Agent, which has 158 GitHub stars. The repository holds 77 skills in this directory. The repository was last updated on October 8, 2026.

Source: Luciole-Studio/Misaka-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.