Agent skill

Vllm

by magnus919 in magnus919/agent-skills

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

MITAuto-check: notesAI & LLM Engineering

Install Vllm

skills CLI
$ npx skills add magnus919/agent-skills --skill vllm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/agent-skills vllm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/vllm .claude/skills/vllm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm
GitHub stars
115
Token cost
~4.1k tokens
SKILL.md length
1,878 words
Files
14 (incl. scripts, references)
Skills in repo
131
Repo updated
First seen
Licence
MIT

At a glance

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

  • Works in 5 steps: Record the deployment before tuning it.… → Confirm the target, scope, and rollback… → A server that responds is not a server… → …
  • Running a vLLM server (vllm serve
  • SKILL.md covers Operating contract, The vllm-health script, Operating loop and Deployment: Docker and…, plus 11 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Vllm is an agent skill from magnus919/agent-skills. Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API serving, throughput and latency benchmarking, continuous batching tuning, GPU operation, and upgrade/rollback. Use when deploying or running a vLLM server (vllm serve, vllm/vllm-openai), sizing a model and its KV cache for GPUs, selecting quantization and parallelism, serving via /v1 endpoints, measuring serving…

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `references/00-source-index.md`). Compatibility notes: Requires a vLLM release (v0.26.0 or a pinned older release), an NVIDIA CUDA, AMD ROCm, or Intel XPU GPU with the matching driver, or a supported CPU build…

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, llama.cpp, OpenAI and Docker. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.

When your agent uses it

  • Running a vLLM server (vllm serve
  • Vllm/vllm-openai)
  • Sizing a model and its KV cache for GPUs
  • Selecting quantization and parallelism

Example prompts

  • “/vllm”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires a vLLM release (v0.26.0 or a pinned older release), an NVIDIA CUDA, AMD ROCm, or Intel XPU GPU with the matching driver, or a supported CPU build. The bundled vllm-health script runs on Python 3.9+ and needs no vLLM server for --help; live probes require HTTP(S) access to a running vLLM server.

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Record the deployment before tuning it. Capture the vLLM version or image digest, model and revision, quantization, parallelism…
  2. Confirm the target, scope, and rollback path before acting. Read-only discovery (health probes, /metrics, nvidia-smi) may proceed without…
  3. A server that responds is not a server that serves. /health returning 200 proves liveness, not that the model loaded or that inference…
  4. Benchmark before and after every change. vLLM flags, defaults, and behavior change between releases; an unmeasured tuning change is a…
  5. Keep evidence bounded. Summarize logs, configs, and metrics; never dump full server logs, .env files, or HF tokens into chat…

What it can do on your machine

Read from SKILL.md and the folder at commit 22b4723. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires a vLLM release (v0.26.0 or a pinned older release), an NVIDIA CUDA, AMD ROCm, or Intel XPU GPU with the matching driver, or a supported CPU build. The bundled vllm-health script runs on Python 3.9+ and needs no vLLM server for --help; live probes require HTTP(S) access to a running vLLM server.

    From compatibility in the SKILL.md frontmatter.

Context cost

Vllm loads about 4.1k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 222 tokens; SKILL.md has 1,878 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~222
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:38
    d metrics; never dump full server logs, `.env` files, or HF tokens into chat. `--enable-log-requests` with debug logging
  • NoteMentions a .env fileSKILL.md:142
    - Never print or commit HF tokens, `.env` contents, or full server logs; summarize evidence instead.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from magnus919/agent-skills at commit 22b4723, republished under its MIT licence (© magnus919). 1,878 words, ~4,139 tokens.

Download SKILL.mdSave it as .claude/skills/vllm/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
vllm
description
Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API serving, throughput and latency benchmarking, continuous batching tuning, GPU operation, and upgrade/rollback. Use when deploying or running a vLLM server (vllm serve, vllm/vllm-openai), sizing a model and its KV cache for GPUs, selecting quantization and parallelism, serving via /v1 endpoints, measuring serving throughput or latency, tuning batching, or diagnosing GPU, OOM, or startup failures in a vLLM deployment. Do not use for model training, fine-tuning, evaluation-set design, or engine-selection methodology (that is ml-engineering), or for operating the llama.cpp stack with GGUF models (that is llama-cpp); other inference engines (TGI, Ollama, Triton) are out of scope.
compatibility
Requires a vLLM release (v0.26.0 or a pinned older release), an NVIDIA CUDA, AMD ROCm, or Intel XPU GPU with the matching driver, or a supported CPU build. The bundled vllm-health script runs on Python 3.9+ and needs no vLLM server for --help; live probes require HTTP(S) access to a running vLLM server.
license
MIT
metadata.source
https://docs.vllm.ai/en/latest/
metadata.source_index
references/00-source-index.md
metadata.research_checked
2026-08-03

vLLM Inference Serving

Use this skill to operate vLLM as a production inference server: deploy it with Docker or Kubernetes, configure the model and engine (quantization, tensor parallelism, KV cache, context length), serve the OpenAI-compatible API surface, benchmark throughput and latency with comparable evidence, tune continuous batching, operate the GPUs underneath, and upgrade or roll back safely. This is a tool skill for one named engine. Serving methodology — engine selection, quantization trade-offs, deployment plans, regression triage — belongs to ml-engineering; local single-node GGUF serving with the llama.cpp stack belongs to llama-cpp. This skill owns the day-to-day operation of vLLM itself.

Operating contract

  1. Record the deployment before tuning it. Capture the vLLM version or image digest, model and revision, quantization, parallelism, max-model-len, KV cache settings, batching limits, GPU inventory, and workload. The serving config template exists for exactly this.
  2. Confirm the target, scope, and rollback path before acting. Read-only discovery (health probes, /metrics, nvidia-smi) may proceed without confirmation. Mutations — restarting a server, changing serving args, scaling replicas, upgrading the image — require an explicit human directive naming the deployment.
  3. A server that responds is not a server that serves. /health returning 200 proves liveness, not that the model loaded or that inference works. Verify at the delivery boundary: /v1/models reports the served model and a representative request returns generated tokens.
  4. Benchmark before and after every change. vLLM flags, defaults, and behavior change between releases; an unmeasured tuning change is a guess. Compare only matched conditions (version, model, GPU, context, batch, workload) and record the evidence in the benchmark run record.
  5. Keep evidence bounded. Summarize logs, configs, and metrics; never dump full server logs, .env files, or HF tokens into chat. --enable-log-requests with debug logging can leak prompt content; keep request logging off or redacted in shared sessions.

The vllm-health script

scripts/vllm-health is an agent-first, read-only probe for a running vLLM server. It issues GET requests only, never mutates, and emits bounded JSON.

bash
scripts/vllm-health --help                    # no server needed
scripts/vllm-health --url http://127.0.0.1:8000 --json
scripts/vllm-health --check health --check models --json
scripts/vllm-health --check metrics --timeout 10 --json

Exit codes: 0 all checks passed, 1 issues found or a fatal error, 2 usage error, 124 timeout. Checks: health (/health), version (/version), models (/v1/models), load (/load), and metrics (a bounded prefix of /metrics). The script never sends data anywhere and never writes files.

Operating loop

  1. Identify the deployment: vLLM version or image digest, model and revision, served model name, parallelism, and how it is deployed (bare vllm serve, Docker, Kubernetes).
  2. Collect evidence: run vllm-health --json for health, version, models, and load; check /metrics counters (vllm:num_requests_running, vllm:num_requests_waiting, vllm:gpu_cache_usage_perc); inspect GPU state with nvidia-smi.
  3. Triage against the symptom: map the problem to the evidence (OOM → KV cache or gpu_memory_utilization; high latency → batching, TTFT vs TPOT; model not found → served name or chat template; slow start → model download or compile cache).
  4. Act with confirmation: bounded, scoped changes after a human directive, with a rollback path named first.
  5. Verify: re-run the probe and the representative request at the delivery boundary, and re-benchmark if the change affects performance.

Deployment: Docker and Kubernetes

  • Docker: the official image is vllm/vllm-openai (Docker Hub). Run with GPU access, the Hugging Face cache mounted, the HF token for gated models, port 8000 published, and --ipc=host (or a --shm-size) for the shared memory tensor parallelism relies on. See references/01-deployment.md.
  • Kubernetes: a Deployment with nvidia.com/gpu (or amd.com/gpu) resources, a PVC for the model cache, an emptyDir backed by Memory at /dev/shm, liveness/readiness probes on /health port 8000, and a Service. Raise probe failureThreshold for large models that take minutes to load — a premature kill shows up as KeyboardInterrupt: terminated in the container log.
  • Pin image tags to a release (for example vllm/vllm-openai:v0.26.0) instead of latest, and persist the compile cache (default ~/.cache/vllm) across restarts so torch.compile artifacts are reused.

Model configuration

  • Model identity: --model is the HF repo or local path; --revision pins the exact weights. --served-model-name sets the name clients must use in /v1 requests and in the model field of responses. --trust-remote-code is required for some model repos and should be reviewed before use.
  • Context length: --max-model-len bounds prompt plus output per request. Unset, it derives from the model config; -1/auto picks the largest length that fits GPU memory. It is the single biggest driver of KV cache size.
  • Quantization-aware serving: pass --quantization (or -q) only when the model weights require it (GPTQ/AWQ/GGUF checkpoints load their scheme from config). Weight types and activation dtypes must match what the kernels support; a quantized model served at the wrong dtype fails to load or silently degrades. Hardware support varies by method (see references/02-model-configuration.md).
  • Tensor parallelism: --tensor-parallel-size N shards one model across N GPUs in the same node; --pipeline-parallel-size splits layers across nodes. TP requires NVLink/fast interconnect and equal per-GPU memory; startup logs the memory profiling result, which is the evidence that the model fits.
  • KV cache: --gpu-memory-utilization (default 0.92) caps the fraction of GPU memory the model plus KV cache may use. --kv-cache-dtype fp8 shrinks the cache for long contexts on supported GPUs. The engine logs GPU KV cache size: N tokens and the implied max concurrency — record both; they tell you how many concurrent requests of a given length the box can hold.

OpenAI-compatible API surface

  • Basic endpoints: /health (liveness), /version, /v1/models (served models), /load (load metrics), /metrics (Prometheus). Inference: /v1/completions and /v1/chat/completions (chat requires the model to ship a chat template, or pass --chat-template); /v1/embeddings for pooling models; /v1/responses for the Responses API.
  • Streaming, tool calling (--enable-auto-tool-choice --tool-call-parser openai), structured outputs, and parallel sampling are server-side options that change request/response behavior — verify each against the installed release rather than assuming parity.
  • Exposing the server beyond loopback requires an explicit decision about bind address, API keys, TLS or a trusted reverse proxy, and firewall rules. Development-only endpoints (/reset_prefix_cache, weight transfer, profiling) must not be exposed in production.

Benchmarking: throughput and latency

  • Online serving benchmark: run vllm bench serve against a live server with a representative dataset (ShareGPT, a local custom JSONL, or your own prompts) and fixed --num-prompts, --request-rate, and --max-concurrency. It reports request throughput (req/s), output token throughput (tok/s), total token throughput, and TTFT/TPOT/ITL percentiles.
  • Offline throughput: vllm bench throughput measures raw engine throughput without the HTTP path; use it for engine-only comparisons, not end-to-end user latency.
  • Comparable evidence: the benchmark run record template freezes version, model, quantization, parallelism, context, batching, GPU, dataset, and load pattern. Never compare numbers across different conditions as if one variable changed. TTFT is a latency metric; token throughput is a throughput metric — an optimization that helps one can hurt the other.
  • For production capacity testing, vLLM's docs recommend the separate GuideLLM framework; this skill's scope is the bundled vllm bench tools.
Show full SKILL.md (781 more words)Show less

Continuous batching tuning

  • vLLM batches continuously by default: the scheduler admits sequences as capacity frees up, mixing prefill and decode. --max-num-seqs caps sequences per iteration, --max-num-batched-tokens caps tokens per iteration, and --enable-chunked-prefill lets prefill share an iteration with decode.
  • Start from defaults and change one knob at a time against the frozen benchmark: raising --max-num-seqs raises throughput at the cost of per-request latency and KV cache pressure; lowering it improves latency stability at the cost of utilization.
  • --enable-prefix-caching reuses KV blocks across requests with shared prefixes (chat system prompts, RAG contexts); the hit rate is visible in /metrics and in the benchmark's input token accounting. --performance-mode trades between interactivity (latency) and throughput at the kernel level.

GPU operation

  • Verify GPUs with nvidia-smi (or rocm-smi on AMD): device list, memory, utilization, temperature, and ECC errors before and after changes. CUDA_VISIBLE_DEVICES selects which GPUs a vllm serve process sees; tensor parallel ranks map to the visible devices in order.
  • Watch /metrics for vllm:gpu_cache_usage_perc (KV cache pressure), vllm:num_requests_running/waiting, and vllm:generation_tokens_total. A cache-usage signal near 1.0 with requests waiting means the deployment is at capacity — scale out or reduce max-model-len/concurrency rather than overcommitting.
  • OOM during startup usually means the model + KV cache did not fit: lower --gpu-memory-utilization does not help if weights alone exceed memory — reduce --max-model-len, switch quantization, or add GPUs. OOM mid-run means KV cache pressure: shrink context, concurrency, or batch limits.

Upgrade and rollback

  • Pin everything: image tag or pip install vllm==<version>, model revision, and the full serving command. latest images and unpinned revisions make rollback impossible and upgrades unreproducible.
  • Upgrade path: read the release notes for the full version span, review changed/removed flags (--engine-args change frequently), validate the new version on a scratch instance with the real model and workload, re-run the frozen benchmark, then swap with a rollback plan: previous image tag and previous serving config ready to reapply.
  • Rollback: because the config is versioned, rollback is a redeploy of the previous pinned image + config. KV cache layout, defaults, and flag names change between releases — do not assume a config that ran on v0.25.x behaves identically on v0.26.x without re-validating and re-benchmarking.

Reference routing

Load whenReference
Sources, version observations, refresh procedurereferences/00-source-index.md
Docker and Kubernetes deployment, image pinning, probes, storagereferences/01-deployment.md
Model config: quantization, tensor parallelism, KV cache, memory budgetingreferences/02-model-configuration.md
OpenAI-compatible API surface, chat templates, tools, authreferences/03-openai-api.md
Benchmarking methodology and vllm bench commandsreferences/04-benchmarking.md
Continuous batching, chunked prefill, prefix caching, performance modereferences/05-batching-and-tuning.md
GPU operation, observability, upgrade/rollback, troubleshootingreferences/06-gpu-ops-and-lifecycle.md

Included artifacts

  • scripts/vllm-health: read-only health/version/models/load/metrics probe (stdlib-only, --json, --check subsets, --help without a server).
  • tests/test_vllm_health.py: deterministic tests against a local stub HTTP server, including the read-only contract.
  • templates/serving-config.md and templates/benchmark-run-record.md: fillable records that make deployments reproducible and benchmark evidence comparable.
  • references/: seven dated, source-indexed references covering the operational topics above.
  • evals/evals.json: six output-quality evaluation cases for agent runs.

Verification boundary

ClaimMinimum evidence
The server is alivevllm-health --check health reports /health 200
The right model is served/v1/models lists the expected served model name
Inference worksA representative /v1/chat/completions or /v1/completions request returns generated tokens with a finish_reason
The model fitsStartup log shows memory profiling completed and GPU KV cache size: N tokens for the configured parallelism
A tuning change helpedThe frozen benchmark shows the declared metric improving with matched conditions, variance reported
The deployment is upgradablePrevious pinned image + serving config are recorded and the upgrade was rehearsed on a scratch instance
A diagnosis is soundEvidence was collected before the claim, and the fix was verified by re-running the probe and the benchmark

Hard boundaries

  • Never restart, redeploy, scale, or upgrade a vLLM deployment without an explicit human directive naming the target and a stated rollback path. Read-only discovery may proceed freely.
  • Never expose an unauthenticated server beyond loopback by accident; development-only endpoints and profiling routes must stay off production ingress.
  • Never print or commit HF tokens, .env contents, or full server logs; summarize evidence instead.
  • Never compare benchmark numbers from different versions, models, quants, parallelism, contexts, batches, or workloads as if one variable changed.
  • Never run vllm-health as anything but what it is — read-only. It has no mutation surface.

When not to use

  • Model training, fine-tuning, evaluation-set design, quantization decisions, and serving methodology — that is ml-engineering.
  • The llama.cpp stack (llama-cli, llama-server, GGUF conversion and quantization, local Metal/CUDA builds) — that is llama-cpp.
  • Other inference engines (TGI, Ollama, Triton, vLLM's embedding/rerank-only workloads are in scope, but engine selection among them is not) — engine-selection trade-offs belong to ml-engineering.
  • Kubernetes and Docker fundamentals (manifests, RBAC, image registries, GPU device plugins) — that is kubernetes and docker-compose.
  • GPU infrastructure provisioning (drivers, cluster scheduling, capacity planning) — that is platform-engineering; this skill operates the GPUs a vLLM server already targets.

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references) in vllm of magnus919/agent-skills.

  • SKILL.md
  • README.md
  • evals/evals.json
  • references/00-source-index.md
  • references/01-deployment.md
  • references/02-model-configuration.md
  • references/03-openai-api.md
  • references/04-benchmarking.md
  • references/05-batching-and-tuning.md
  • references/06-gpu-ops-and-lifecycle.md
  • scripts/vllm-health
  • templates/benchmark-run-record.md
  • templates/serving-config.md
  • tests/test_vllm_health.py

Open the folder on GitHubat commit 22b4723

Compare with similar skills

Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm this skillmagnus919/agent-skills115—~4.1kAutomated safety check: NotesMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Vllm Serversickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Server

    sickn33/agentic-awesome-skills

    Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from magnus919/agent-skills

All 131 skills in this repo
  • Artifact Pyramids

    magnus919/agent-skills

    Organize durable agent research outputs as summaries, analysis, and evidence dossiers.

    115 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Ascii City Engine

    magnus919/agent-skills

    Build portable, first-person colored ASCII city engines and small GIS-derived city packs.

    115 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Color Management

    magnus919/agent-skills

    Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.

    115 GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    115 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Docker Compose

    magnus919/agent-skills

    Use Docker Compose to define, run, debug, and harden multi-container applications.

    115 GitHub stars~2k tokensUpdated today
    Auto-check: notes
  • Fpga Development

    magnus919/agent-skills

    Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.

    115 GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Questions about Vllm

What does Vllm do?

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…. Vllm is an agent skill from magnus919/agent-skills. Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API serving, throughput and latency benchmarking, continuous batching tuning, GPU operation, and upgrade/rollback.

When should I use Vllm?

Vllm fits situations like: running a vLLM server (vllm serve; vllm/vllm-openai); sizing a model and its KV cache for GPUs; selecting quantization and parallelism.

How do I install Vllm in Claude Code?

Run `npx skills add magnus919/agent-skills --skill vllm -a claude-code`. Or copy the skill folder (vllm in magnus919/agent-skills) into .claude/skills/vllm in your project. Claude Code loads it when a task matches its description.

How do I install Vllm in Codex?

Run `npx skills add magnus919/agent-skills --skill vllm -a codex`. Or copy the skill folder (vllm in magnus919/agent-skills) into .agents/skills/vllm in your project. Codex loads it when a task matches its description.

Can I use Vllm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm, .gemini/skills/vllm, .github/skills/vllm and .opencode/skills/vllm in your project.

What does Vllm need to run?

Going by SKILL.md and its folder, Vllm needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; Docker. Compatibility (from SKILL.md): Requires a vLLM release (v0.26.0 or a pinned older release), an NVIDIA CUDA, AMD ROCm, or Intel XPU GPU with the matching driver, or a supported CPU build. The bundled vllm-health script runs on Python 3.9+ and needs no vLLM server for --help; live probes require HTTP(S) access to a running vLLM server..

Does Vllm access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vllm safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Vllm use?

Vllm is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm use?

About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.6k tokens, read only when the agent opens those files.

What are the alternatives to Vllm?

Skills that share tags, products or a category with Vllm: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Vllm Server (sickn33/agentic-awesome-skills, 47k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm?

magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 115 GitHub stars. The repository holds 131 skills in this directory. The repository was last updated on October 10, 2026.

Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.