Agent skill

Vllm

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

MITAuto-check passedAI & LLM Engineering

Install Vllm

skills CLI
$ npx skills add ericrisco/rsc-harness --skill vllm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness vllm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vllm .claude/skills/vllm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm
GitHub stars
174
Token cost
~3.6k tokens
SKILL.md length
1,600 words
Files
5 (incl. references)
Skills in repo
233
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

  • Self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint
  • SKILL.md covers Version & setup reality…, vllm serve → an…, Parallelism: fit the model,… and Quantization: smaller weights…, plus 8 more sections
  • Calls curl and pip; needs VLLM_API_KEY
  • Splitting a model across GPUs with tensor

What it does

Vllm is an agent skill from ericrisco/rsc-harness. Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, loading quantized weights, serving one or many LoRA adapters, and debugging KV-cache OOM from memory-utilisation and context-length flags. NOT renting or provisioning the GPU box (that is runpod or modal), NOT single-user laptop inference (that is ollama), NOT a hosted inference API you do not operate (that is…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/flags-and-endpoints.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Fine-tuning. It works with vLLM, OpenAI, Ollama and Hugging Face. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint
  • Splitting a model across GPUs with tensor
  • Pipeline parallelism
  • Loading quantized weights

Example prompts

  • “/vllm”

Requirements

  • Python 3
  • A credential in VLLM_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.vllm.ai
    • pypi.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VLLM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm loads about 3.6k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 1,600 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,600 words, ~3,566 tokens.

Download SKILL.mdSave it as .claude/skills/vllm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
vllm
description
Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, loading quantized weights, serving one or many LoRA adapters, and debugging KV-cache OOM from memory-utilisation and context-length flags. NOT renting or provisioning the GPU box (that is `runpod` or `modal`), NOT single-user laptop inference (that is `ollama`), NOT a hosted inference API you do not operate (that is `together-fireworks` or `huggingface`).
tags
vllm, llm-serving, inference-server, tensor-parallel, quantization
recommends
open-weights, finetuning, runpod, modal, ollama
origin
risco

vLLM — high-throughput serving of open-weight models

vLLM is the inference engine you put in front of an open-weight model when many requests hit it at once. Its job — and this skill's — is throughput under concurrency: keep the GPU busy across dozens of simultaneous requests, not squeeze one prompt out fast. You own the vllm serve flags; the box those flags run on is a runpod/modal concern.

Why not just loop a transformers generate()? Naive serving runs one request at a time and pads every batch to the longest sequence, so the GPU idles. vLLM fixes both:

  • PagedAttention stores the KV cache in non-contiguous fixed-size blocks (like OS virtual-memory paging), so there is almost no padding/reservation waste and long contexts pack tightly.
  • Continuous batching admits and retires requests token-by-token instead of per-batch, so a new request joins the running batch immediately rather than waiting for the slowest one to finish.

Net effect: an order-of-magnitude more concurrent throughput than single-request serving. If you only ever have one user on a laptop, that machinery is wasted — that is ollama, not this.

Version & setup reality (verify at author time — vLLM ships ~weekly)

  • Latest stable is ~0.25.x (July 2026) — releases land roughly weekly. Check PyPI / releases; do not pin to a number you read here.
  • The V1 engine is the default since v0.8.0 and recent releases have removed the legacy V0 path, so treat V1 as the only engine. It runs the scheduler + core loop in a separate process and turns on chunked prefill and prefix caching out of the box, so you rarely tune the scheduler by hand (V1 guide, accessed 2026-07). VLLM_USE_V1 historically toggled it — if you see it referenced, it is legacy.
  • Install: pip install vllm (default build is CUDA/NVIDIA). ROCm, CPU, TPU and other backends have separate install paths — see the docs' installation matrix. Needs Python + a supported GPU.
  • Auth is OFF by default. A bare vllm serve is an open endpoint on 0.0.0.0:8000. To require a bearer token, pass --api-key <KEY> (or set VLLM_API_KEY); clients then send Authorization: Bearer <KEY>. Do not expose an unauthenticated server publicly.

vllm serve → an OpenAI-compatible server

bash
vllm serve Qwen/Qwen3-8B                 # download from HF (or a local path) and serve on :8000
vllm serve /models/qwen3-8b --api-key sk-local-xyz --port 8000 --host 0.0.0.0

The server speaks the OpenAI wire protocol, so any OpenAI client works unchanged — just repoint base_url and use a dummy (or your --api-key) key. Endpoints (online serving docs, accessed 2026-07):

EndpointPurpose
POST /v1/chat/completionschat protocol (messages array) — the usual path
POST /v1/completionsraw text completion (single prompt string)
GET /v1/modelslist the served model + any loaded LoRA adapters
POST /v1/embeddingsonly for an embedding model (--task embed)
GET /healthliveness — 200 when ready, no auth, no body
GET /metricsPrometheus metrics (queue depth, throughput, cache usage)
python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-local-xyz")  # key = your --api-key
r = client.chat.completions.create(
    model="Qwen/Qwen3-8B",               # the id `/v1/models` reports (or a LoRA adapter name)
    messages=[{"role": "user", "content": "Name three primes."}],
)
print(r.choices[0].message.content)

The model field must match what /v1/models returns — the served model id or a LoRA adapter name (below), not an arbitrary string.

Parallelism: fit the model, then scale out

Two orthogonal knobs. Reach for parallelism only when the model does not fit one GPU — a model that fits should stay on a single GPU (no split), because every split adds communication overhead.

  • --tensor-parallel-size N — shard each layer's weights across N GPUs on one node. Use this first: it needs fast intra-node links (NVLink / PCIe) because GPUs sync every layer. N must divide the model's attention-head count. This is how you serve a model too big for one GPU but fitting the node.
  • --pipeline-parallel-size M — split the model by layer stages across M nodes. Tolerates slower inter-node network. Use it when the model does not fit even a full node.

Rule of thumb from the docs (parallelism & scaling, accessed 2026-07): set tensor-parallel = GPUs per node, pipeline-parallel = number of nodes.

bash
vllm serve meta-llama/Llama-3.3-70B-Instruct --tensor-parallel-size 4          # 1 node, 4 GPUs
vllm serve <huge-model> --tensor-parallel-size 8 --pipeline-parallel-size 2    # 2 nodes × 8 GPUs

Multi-node needs a Ray cluster wired up first — that orchestration is a runpod/modal concern, not a vLLM flag.

Quantization: smaller weights ≠ automatic speedup

Quantization shrinks the weights so a model fits fewer/smaller GPUs and frees VRAM for KV cache. Select a method with --quantization (often auto-detected from the checkpoint's config). vLLM supports AWQ, GPTQ/GPTQModel, FP8 (W8A8), compressed-tensors (LLM Compressor), INT4/INT8, bitsandbytes and more (quantization docs, accessed 2026-07).

bash
vllm serve TheModel/Qwen3-8B-AWQ --quantization awq
vllm serve neuralmagic/Model-FP8   --quantization compressed-tensors   # FP8 needs Ada/Hopper+

The honest tradeoffs:

  • Support is GPU-arch-specific — FP8 W8A8 wants Ada/Hopper-class (or AMD) hardware; a method that is fast on one GPU may be unsupported or emulated on another. Check the compatibility table.
  • Quant is a memory win, not a guaranteed throughput win. At low batch / low concurrency the workload is memory-bandwidth-bound and a quant can help; but dequant overhead can cost latency, and at high batch you may be compute-bound where a weight-only quant does little. Benchmark your quant on your hardware at your real concurrency before assuming it is faster.
  • Quality drops — usually small for 8-bit/well-calibrated 4-bit, larger for aggressive 4-bit. Serve a quant to save VRAM, not as a free lunch. Picking which quantized checkpoint to pull is an open-weights decision.

LoRA adapters: serve the fine-tuning output

vLLM serves LoRA adapters on top of one loaded base model, so your finetuning/unsloth output goes live without merging or a second server (LoRA docs, accessed 2026-07).

bash
vllm serve meta-llama/Llama-3.2-3B-Instruct \
  --enable-lora \
  --lora-modules sql=/adapters/sql-lora legal=/adapters/legal-lora \
  --max-loras 2 \        # how many adapters resident at once
  --max-lora-rank 16     # must be >= the rank the adapter was trained at

Route to an adapter by naming it in the request model field: "model": "sql" hits the SQL adapter, "model": "meta-llama/Llama-3.2-3B-Instruct" hits the untuned base — same server, no reload.

Load/unload at runtime with VLLM_ALLOW_RUNTIME_LORA_UPDATING=True, then POST /v1/load_lora_adapter / POST /v1/unload_lora_adapter. Base + adapter must match (same family and dims), and --max-lora-rank must be ≥ the trained rank or load fails.

Memory & throughput: where OOM comes from

The two knobs that cause (and cure) most OOM (engine args, accessed 2026-07):

  • --gpu-memory-utilization (default ~0.92) — fraction of each GPU vLLM may claim. Weights are loaded, then the rest of this budget becomes the KV-cache pool. Raising it toward 1.0 buys more concurrent sequences but risks OOM from activation/CUDA-graph spikes; lower it if you get OOM at load or under burst.
  • --max-model-len — max context (prompt + output) per request; auto-derived from the model config if unset. This is the single biggest OOM lever: KV cache scales with max-model-len × concurrent sequences. A model whose weights fit will still OOM if you leave the full 128K context on and let many long requests batch. Cap --max-model-len to what you actually need.

The mental model: KV_pool = gpu_memory_utilization × VRAM − weights, and concurrent_sequences ≈ KV_pool ÷ (bytes_per_token × context_len). Too-long context or too-high utilization eats the pool. --max-num-seqs and --max-num-batched-tokens cap the batch to trade latency vs throughput. Full KV math + an OOM playbook: references/memory-and-throughput.md.

Show full SKILL.md (562 more words)Show less

Self-host health check (no credential needed)

/health requires no API key, so it is the safe first probe on any endpoint you did not just start:

bash
export VLLM_BASE_URL=http://localhost:8000
curl -fsS "$VLLM_BASE_URL/health" && echo "  up"     # 200, empty body when ready
curl -fsS "$VLLM_BASE_URL/v1/models" \
  -H "Authorization: Bearer ${VLLM_API_KEY:-sk-local}"  # confirms WHICH model/adapters are served

/health says "the server is alive"; /v1/models says "and it is serving the model you expect" (plus any LoRA adapters). If --api-key is set, /v1/models needs the bearer header but /health never does. During load a slow first /health is normal — big models take a while to page in.

Honest alternatives — when to prefer another engine

  • Text Generation Inference (TGI) — Hugging Face's server; reach for it if you are all-in on the HF ecosystem/tooling and want their supported stack.
  • SGLang — competitive throughput with strong RadixAttention prefix-cache reuse; prefer it for heavy shared-prefix workloads (agents, long system prompts, structured generation).
  • TensorRT-LLM — squeezes maximum latency/throughput on NVIDIA GPUs via compiled engines, at the cost of a heavier build/ahead-of-time compile step; prefer it when you must wring out every last ms on NVIDIA hardware and can pay the ops complexity.
  • Ollama / llama.cpp — single-box / laptop / CPU-or-Metal, one or few users, GGUF quants, zero-ceremony. Prefer it when there is no concurrency to exploit — that is the ollama skill.

Guardrails

  • Cap --max-model-len. Leaving the full context window on is the top OOM cause; KV cache scales with context × concurrency (engine args).
  • Tune --gpu-memory-utilization for OOM, not throughput first. Lower it if OOM at load/burst; raise it (carefully) for more concurrency. It is a fraction of one GPU's memory.
  • A model that fits one GPU should stay on one GPU. Tensor parallelism adds sync overhead — use it to fit, not to speed up an already-fitting model.
  • Quant is a VRAM win, not free speed, and its quality cost is real — verify throughput at your concurrency and quality on your task (quantization).
  • Never expose an unauthenticated server. No --api-key = open endpoint on 0.0.0.0:8000.
  • --max-lora-rank must be ≥ the adapter's trained rank, and base + adapter must match, or the LoRA fails to load.
  • This is not a laptop tool. No GPU / one user → ollama. The GPU box + autoscaling → runpod / modal.
  • open-weights — choose the model/size/license/quant to serve. vLLM runs it; it does not pick it.
  • finetuning (and single-GPU unsloth) — produce the LoRA adapter or merged weights that --enable-lora/vllm serve then hosts. They train; vLLM serves.
  • runpod — rent/provision the GPU box vLLM runs on (bring-your-own-container GPU rental).
  • modal — serverless GPU containers + autoscaling around a vLLM process. The box and scaling, vs vLLM the engine.
  • ollama — the single-user / laptop counterpart. No concurrency to exploit → use it, not vLLM.
  • Hosted inference you do not operate → together-fireworks / huggingface (you call an API; here you run the server).

Checklist

  • Chose vLLM because there is concurrency to exploit (else ollama / a hosted API).
  • vllm serve <model> up; /health returns 200 and /v1/models shows the expected model.
  • --api-key (or VLLM_API_KEY) set if the endpoint is reachable beyond localhost.
  • Parallelism only if the model does not fit one GPU: TP = GPUs/node, PP = nodes.
  • --max-model-len capped to real need; --gpu-memory-utilization set with OOM headroom.
  • Quant chosen for VRAM fit and benchmarked at real concurrency (not assumed faster).
  • LoRA: --max-lora-rank ≥ trained rank; adapter reachable by name via the model field.
  • Client points at base_url=<server>/v1; model matches a /v1/models id.

References

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/vllm of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/flags-and-endpoints.md
  • references/memory-and-throughput.md

Open the folder on GitHubat commit e3d5b33

Compare with similar skills

Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm this skillericrisco/rsc-harness174—~3.6kAutomated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Mesh APImr-tbot/mesh-api180—~1.8kAutomated safety check: PassGPL-3.0
Hqq QuantizationOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
Vllm Deploy K8svllm-project/vllm-skills103—~2kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Mesh API

    mr-tbot/mesh-api

    Interact with a Meshtastic LoRa mesh network through MESH-API — list nodes, read messages, send texts, and check connection status.

    180 GitHub stars~1.8k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hqq Quantization

    Orchestra-Research/AI-Research-SKILLs

    Half-Quadratic Quantization for LLMs without calibration data.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub stars~1.7k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed

More from ericrisco/rsc-harness

All 233 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    174 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    174 GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    174 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    174 GitHub stars~3.2k tokensUpdated 2 days ago
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    174 GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    174 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed

Questions about Vllm

What does Vllm do?

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…. Vllm is an agent skill from ericrisco/rsc-harness. Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, loading quantized weights, serving one or many LoRA adapters, and debugging KV-cache OOM from memory-utilisation and context-length flags.

When should I use Vllm?

Vllm fits situations like: self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint; splitting a model across GPUs with tensor; pipeline parallelism; loading quantized weights.

How do I install Vllm in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill vllm -a claude-code`. Or copy the skill folder (skills/vllm in ericrisco/rsc-harness) into .claude/skills/vllm in your project. Claude Code loads it when a task matches its description.

How do I install Vllm in Codex?

Run `npx skills add ericrisco/rsc-harness --skill vllm -a codex`. Or copy the skill folder (skills/vllm in ericrisco/rsc-harness) into .agents/skills/vllm in your project. Codex loads it when a task matches its description.

Can I use Vllm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm, .gemini/skills/vllm, .github/skills/vllm and .opencode/skills/vllm in your project.

What does Vllm need to run?

Going by SKILL.md and its folder, Vllm needs the command-line tools its instructions call (curl and pip) and credentials named VLLM_API_KEY. Our summary lists: Python 3; A credential in VLLM_API_KEY.

Does Vllm access the network?

SKILL.md names 3 domains. As links in the text: docs.vllm.ai, pypi.org and github.com. This is read from the text; nothing was executed.

Is Vllm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm use?

Vllm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to Vllm?

Skills that share tags, products or a category with Vllm: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Resolve (alexziskind1/model-shelf, 130 stars), Mesh API (mr-tbot/mesh-api, 180 stars) and Hqq Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 174 GitHub stars. The repository holds 233 skills in this directory. The repository was last updated on October 7, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.