Agent skill

Vllm

by AlexAI-MCP in AlexAI-MCP/hermes-CCC

Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

MITAuto-check passedAI & LLM Engineering

Install Vllm

skills CLI
$ npx skills add AlexAI-MCP/hermes-CCC --skill vllm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexAI-MCP/hermes-CCC vllm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vllm .claude/skills/vllm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm
GitHub stars
135
Token cost
~2.3k tokens
SKILL.md length
741 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers Purpose, Install, What vLLM Gives You and Common Models, plus 15 more sections
  • Calls python, curl and pip

What it does

Vllm is an agent skill from AlexAI-MCP/hermes-CCC. Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and OpenAI. The repository describes itself as: Hermes Agent ported to Claude Code Channel — 46 native skills, no OAuth, no external process. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/vllm”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 8107e89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • curl
    • pip
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, pip and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm loads about 2.3k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 741 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlexAI-MCP/hermes-CCC at commit 8107e89, republished under its MIT licence (© AlexAI-MCP). 741 words, ~2,277 tokens.

Download SKILL.mdSave it as .claude/skills/vllm/SKILL.md (or your agent's skills folder).
name
vllm
description
Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.
version
1.0.0
author
hermes-CCC (ported from Hermes Agent by NousResearch)
license
MIT

vLLM

Purpose

  • Use this skill to deploy local or remote LLM inference with vllm.
  • Prefer it when you need OpenAI-compatible serving, high throughput, and modern GPU utilization.
  • vLLM is strongest for decoder-only chat and completion models.
  • It is a good default for production inference when latency and token throughput matter.

Install

  • Install from PyPI:
bash
pip install vllm
  • Verify the install:
bash
python -c "import vllm; print(vllm.__version__)"
  • Match your CUDA, NVIDIA driver, and PyTorch stack before deploying on GPUs.
  • If deployment is containerized, prefer pinning a known-good image or package version.

What vLLM Gives You

  • OpenAI-compatible HTTP API for chat and completion workflows
  • PagedAttention for efficient KV cache memory usage
  • Continuous batching for higher aggregate throughput
  • Tensor parallelism for multi-GPU serving
  • Quantization support for lower memory footprints
  • Streaming token responses
  • Good support for major open-weight model families

Common Models

  • meta-llama/Llama-3.1-8B-Instruct
  • meta-llama/Llama-3.1-70B-Instruct
  • Qwen/Qwen2.5-7B-Instruct
  • Qwen/Qwen2.5-14B-Instruct
  • mistralai/Mistral-7B-Instruct-v0.3
  • microsoft/Phi-3-medium-4k-instruct
  • google/gemma-2-9b-it

Basic Serve

  • Minimal OpenAI-compatible server:
bash
python -m vllm.entrypoints.openai.api_server --model meta-llama/Llama-3.1-8B-Instruct
  • By default this serves on port 8000.
  • The API base path is http://localhost:8000/v1.
  • Health check endpoint:
bash
curl http://localhost:8000/health

Key Flags

  • --port: bind a non-default server port
  • --tensor-parallel-size: split one model across multiple GPUs
  • --gpu-memory-utilization: cap how much GPU RAM vLLM should attempt to use
  • --max-model-len: set maximum sequence length for inference
  • --quantization: load quantized weights when supported

Serve Examples

  • Serve on port 8080:
bash
python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --port 8080
  • Raise the context window and tune memory usage:
bash
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-7B-Instruct \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.92
  • Use four GPUs with tensor parallelism:
bash
python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-70B-Instruct \
  --tensor-parallel-size 4

Quantization

  • vLLM commonly works with these quantization paths:

  • awq

  • gptq

  • bitsandbytes

  • fp8

  • Example with AWQ:

bash
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-7B-Instruct-AWQ \
  --quantization awq
  • Example with GPTQ:
bash
python -m vllm.entrypoints.openai.api_server \
  --model TheBloke/Mistral-7B-Instruct-v0.2-GPTQ \
  --quantization gptq
  • Example with FP8:
bash
python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --quantization fp8
  • AWQ and GPTQ reduce VRAM pressure at some accuracy or compatibility cost.
  • bitsandbytes is useful when using 4-bit or 8-bit HF-compatible flows.
  • FP8 is hardware-sensitive and best validated on your actual deployment target.

Call Through the OpenAI Client

  • vLLM exposes an OpenAI-style API, so the standard OpenAI Python client is a common fit.
  • Point the client at the local base URL:
python
from openai import OpenAI

client = OpenAI(
    api_key="dummy",
    base_url="http://localhost:8000/v1",
)

response = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Summarize why continuous batching matters."},
    ],
    temperature=0.2,
    max_tokens=256,
)

print(response.choices[0].message.content)
  • base_url='http://localhost:8000/v1' is the key integration setting.
  • Many SDK-based applications can switch from OpenAI-hosted inference to vLLM with only the base URL and model name changed.

Curl Chat Example

bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [
      {"role": "system", "content": "You are concise."},
      {"role": "user", "content": "Explain PagedAttention in two sentences."}
    ],
    "temperature": 0.2,
    "max_tokens": 128
  }'

Async Engine Usage in Python

  • Use the async engine when embedding vLLM inside a Python application instead of only exposing the HTTP server.
  • This pattern is useful for custom services, pipelines, or batched internal inference.
python
import asyncio

from vllm import AsyncEngineArgs, AsyncLLMEngine, SamplingParams


async def main() -> None:
    engine_args = AsyncEngineArgs(
        model="meta-llama/Llama-3.1-8B-Instruct",
        gpu_memory_utilization=0.9,
        max_model_len=8192,
    )
    engine = AsyncLLMEngine.from_engine_args(engine_args)
    sampling_params = SamplingParams(temperature=0.2, max_tokens=128)

    request_id = "req-1"
    prompt = "List three use cases for OpenAI-compatible local inference."

    async for output in engine.generate(prompt, sampling_params, request_id):
        if output.finished:
            print(output.outputs[0].text)


asyncio.run(main())
  • Use unique request IDs for concurrent work.
  • Reuse one engine per process rather than creating a new engine for every request.
  • Keep sampling params explicit for reproducibility in evaluation workflows.

Multi-GPU Notes

  • Scale large models with --tensor-parallel-size 4 or another GPU count that matches the host.
  • Ensure all GPUs are visible and comparable in capability.
  • Cross-GPU communication performance matters, so NVLink or high-bandwidth PCIe topology helps.
  • Validate memory headroom with the exact context size and batch profile you plan to serve.
Show full SKILL.md (299 more words)Show less

Benchmarking

  • Benchmark throughput before promoting a serving config to production.
  • Example benchmark entry point:
bash
python benchmarks/benchmark_throughput.py
  • Track at least:
  • prompt tokens per second
  • generated tokens per second
  • p50 and p95 latency
  • GPU memory usage
  • concurrency behavior under steady load

Docker Deployment

  • Container deployment is common for consistent drivers, dependencies, and rollout processes.
  • Example run command:
bash
docker run --gpus all --rm -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3.1-8B-Instruct
  • Mount the Hugging Face cache to avoid repeated model downloads.
  • Pin the image tag in production instead of relying on latest.
  • Pass the same serving flags in Docker that you would use on bare metal.

Health and Readiness

  • Basic health check:
bash
curl http://localhost:8000/health
  • Add a model list probe for smoke testing:
bash
curl http://localhost:8000/v1/models
  • In production, combine health checks with a real inference probe before shifting traffic.

Operational Guidance

  • Start with an 8B instruct model before scaling to larger checkpoints.
  • Keep --gpu-memory-utilization conservative during initial rollout.
  • Set --max-model-len only as high as the workload requires.
  • Prefer quantization only after validating quality, tool calling, and long-context behavior.
  • Profile with real prompts, not only synthetic benchmarks.

Common Failure Modes

  • Out-of-memory on startup:

  • lower --max-model-len

  • lower concurrency expectations

  • use quantized weights

  • reduce model size

  • Poor throughput:

  • increase batch pressure

  • verify GPU utilization

  • benchmark with realistic prompt and completion lengths

  • Client compatibility issues:

  • confirm the app is targeting http://localhost:8000/v1

  • confirm the requested model name exactly matches the served model

  • Quantized model load failures:

  • confirm the checkpoint format matches the selected --quantization mode

  • test a non-quantized baseline first

When To Use This Skill

  • You need an OpenAI-compatible endpoint for self-hosted LLMs.
  • You want high-throughput local or cluster inference.
  • You need tensor parallel serving for models larger than one GPU can hold.
  • You are comparing quantization tradeoffs in a production-style serving stack.

Quick Reference

  • Install: pip install vllm
  • Serve: python -m vllm.entrypoints.openai.api_server --model meta-llama/Llama-3.1-8B-Instruct
  • Health: curl http://localhost:8000/health
  • API base URL: http://localhost:8000/v1
  • Multi-GPU: --tensor-parallel-size 4
  • Benchmark: python benchmarks/benchmark_throughput.py

© AlexAI-MCP, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/vllm of AlexAI-MCP/hermes-CCC.

Open the folder on GitHubat commit 8107e89

Compare with similar skills

Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm this skillAlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT
Vllm Bench Random Syntheticvllm-project/vllm-skills102—~1.5kAutomated safety check: PassApache-2.0
Vllm Bench Servevllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    102 GitHub stars~1.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    102 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    102 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from AlexAI-MCP/hermes-CCC

All 44 skills in this repo
  • GitHub Code Review

    AlexAI-MCP/hermes-CCC

    Review GitHub pull requests with a findings-first engineering mindset.

    135 GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • GitHub PR Workflow

    AlexAI-MCP/hermes-CCC

    Run a disciplined GitHub pull request workflow from branch creation through merge.

    135 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Memory

    AlexAI-MCP/hermes-CCC

    Manage durable project memory for Claude Code. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Route

    AlexAI-MCP/hermes-CCC

    Route Claude Code work by complexity, risk, and tool needs. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Skill

    AlexAI-MCP/hermes-CCC

    Create, improve, inventory, and audit Claude Code skills. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Traj

    AlexAI-MCP/hermes-CCC

    Capture Claude Code interaction trajectories in training-friendly formats.

    135 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Vllm

What does Vllm do?

Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support. Vllm is an agent skill from AlexAI-MCP/hermes-CCC. Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

When should I use Vllm?

Vllm fits situations like: tasks that involve LLM inference and serving.

How do I install Vllm in Claude Code?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill vllm -a claude-code`. Or copy the skill folder (skills/vllm in AlexAI-MCP/hermes-CCC) into .claude/skills/vllm in your project. Claude Code loads it when a task matches its description.

How do I install Vllm in Codex?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill vllm -a codex`. Or copy the skill folder (skills/vllm in AlexAI-MCP/hermes-CCC) into .agents/skills/vllm in your project. Codex loads it when a task matches its description.

Can I use Vllm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexAI-MCP/hermes-CCC --skill vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm, .gemini/skills/vllm, .github/skills/vllm and .opencode/skills/vllm in your project.

What does Vllm need to run?

Going by SKILL.md and its folder, Vllm needs the command-line tools its instructions call (python, curl, pip and docker). Our summary lists: Python 3; Docker.

Does Vllm access the network?

SKILL.md contains no URLs. Its commands use curl, pip and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vllm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm use?

Vllm is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vllm?

Skills that share tags, products or a category with Vllm: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Bench Random Synthetic (vllm-project/vllm-skills, 102 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm?

AlexAI-MCP (a GitHub user) maintains it in AlexAI-MCP/hermes-CCC, which has 135 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on April 8, 2026.

Source: AlexAI-MCP/hermes-CCC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.