Aider Delegate
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…
$ npx skills add ericrisco/rsc-harness --skill vllm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness vllm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vllm .claude/skills/vllm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vllm into .claude/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/vllmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill vllm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness vllm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/vllm .agents/skills/vllm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vllm into .agents/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill vllm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness vllm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/vllm .cursor/skills/vllm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vllm into .cursor/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/vllm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill vllm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness vllm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/vllm .gemini/skills/vllm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vllm into .gemini/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness vllmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill vllm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/vllm .github/skills/vllm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vllm into .github/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill vllm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness vllm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/vllm .opencode/skills/vllm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/vllm into .opencode/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllmA skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…
Vllm is an agent skill from ericrisco/rsc-harness. Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, loading quantized weights, serving one or many LoRA adapters, and debugging KV-cache OOM from memory-utilisation and context-length flags. NOT renting or provisioning the GPU box (that is runpod or modal), NOT single-user laptop inference (that is ollama), NOT a hosted inference API you do not operate (that is…
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/flags-and-endpoints.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving and Fine-tuning. It works with vLLM, OpenAI, Ollama and Hugging Face. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.vllm.aipypi.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VLLM_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm loads about 3.6k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 1,600 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,600 words, ~3,566 tokens.
.claude/skills/vllm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.vLLM is the inference engine you put in front of an open-weight model when many requests hit it at
once. Its job — and this skill's — is throughput under concurrency: keep the GPU busy across dozens of
simultaneous requests, not squeeze one prompt out fast. You own the vllm serve flags; the box those
flags run on is a runpod/modal concern.
Why not just loop a transformers generate()? Naive serving runs one request at a time and pads
every batch to the longest sequence, so the GPU idles. vLLM fixes both:
Net effect: an order-of-magnitude more concurrent throughput than single-request serving. If you only
ever have one user on a laptop, that machinery is wasted — that is ollama, not this.
VLLM_USE_V1
historically toggled it — if you see it referenced, it is legacy.pip install vllm (default build is CUDA/NVIDIA). ROCm, CPU, TPU and other backends
have separate install paths — see the docs' installation matrix. Needs Python + a supported GPU.vllm serve is an open endpoint on 0.0.0.0:8000. To require a
bearer token, pass --api-key <KEY> (or set VLLM_API_KEY); clients then send
Authorization: Bearer <KEY>. Do not expose an unauthenticated server publicly.vllm serve → an OpenAI-compatible servervllm serve Qwen/Qwen3-8B # download from HF (or a local path) and serve on :8000
vllm serve /models/qwen3-8b --api-key sk-local-xyz --port 8000 --host 0.0.0.0The server speaks the OpenAI wire protocol, so any OpenAI client works unchanged — just repoint
base_url and use a dummy (or your --api-key) key. Endpoints
(online serving docs, accessed 2026-07):
| Endpoint | Purpose |
|---|---|
POST /v1/chat/completions | chat protocol (messages array) — the usual path |
POST /v1/completions | raw text completion (single prompt string) |
GET /v1/models | list the served model + any loaded LoRA adapters |
POST /v1/embeddings | only for an embedding model (--task embed) |
GET /health | liveness — 200 when ready, no auth, no body |
GET /metrics | Prometheus metrics (queue depth, throughput, cache usage) |
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-local-xyz") # key = your --api-key
r = client.chat.completions.create(
model="Qwen/Qwen3-8B", # the id `/v1/models` reports (or a LoRA adapter name)
messages=[{"role": "user", "content": "Name three primes."}],
)
print(r.choices[0].message.content)The model field must match what /v1/models returns — the served model id or a LoRA adapter name
(below), not an arbitrary string.
Two orthogonal knobs. Reach for parallelism only when the model does not fit one GPU — a model that fits should stay on a single GPU (no split), because every split adds communication overhead.
--tensor-parallel-size N — shard each layer's weights across N GPUs on one node. Use this
first: it needs fast intra-node links (NVLink / PCIe) because GPUs sync every layer. N must divide
the model's attention-head count. This is how you serve a model too big for one GPU but fitting the
node.--pipeline-parallel-size M — split the model by layer stages across M nodes. Tolerates slower
inter-node network. Use it when the model does not fit even a full node.Rule of thumb from the docs (parallelism & scaling, accessed 2026-07): set tensor-parallel = GPUs per node, pipeline-parallel = number of nodes.
vllm serve meta-llama/Llama-3.3-70B-Instruct --tensor-parallel-size 4 # 1 node, 4 GPUs
vllm serve <huge-model> --tensor-parallel-size 8 --pipeline-parallel-size 2 # 2 nodes × 8 GPUsMulti-node needs a Ray cluster wired up first — that orchestration is a runpod/modal concern, not a
vLLM flag.
Quantization shrinks the weights so a model fits fewer/smaller GPUs and frees VRAM for KV cache. Select
a method with --quantization (often auto-detected from the checkpoint's config). vLLM supports
AWQ, GPTQ/GPTQModel, FP8 (W8A8), compressed-tensors (LLM Compressor), INT4/INT8, bitsandbytes and
more (quantization docs, accessed 2026-07).
vllm serve TheModel/Qwen3-8B-AWQ --quantization awq
vllm serve neuralmagic/Model-FP8 --quantization compressed-tensors # FP8 needs Ada/Hopper+The honest tradeoffs:
open-weights decision.vLLM serves LoRA adapters on top of one loaded base model, so your finetuning/unsloth output goes
live without merging or a second server
(LoRA docs, accessed 2026-07).
vllm serve meta-llama/Llama-3.2-3B-Instruct \
--enable-lora \
--lora-modules sql=/adapters/sql-lora legal=/adapters/legal-lora \
--max-loras 2 \ # how many adapters resident at once
--max-lora-rank 16 # must be >= the rank the adapter was trained atRoute to an adapter by naming it in the request model field: "model": "sql" hits the SQL adapter,
"model": "meta-llama/Llama-3.2-3B-Instruct" hits the untuned base — same server, no reload.
Load/unload at runtime with VLLM_ALLOW_RUNTIME_LORA_UPDATING=True, then
POST /v1/load_lora_adapter / POST /v1/unload_lora_adapter. Base + adapter must match (same family
and dims), and --max-lora-rank must be ≥ the trained rank or load fails.
The two knobs that cause (and cure) most OOM (engine args, accessed 2026-07):
--gpu-memory-utilization (default ~0.92) — fraction of each GPU vLLM may claim. Weights are
loaded, then the rest of this budget becomes the KV-cache pool. Raising it toward 1.0 buys more
concurrent sequences but risks OOM from activation/CUDA-graph spikes; lower it if you get OOM at load
or under burst.--max-model-len — max context (prompt + output) per request; auto-derived from the model config
if unset. This is the single biggest OOM lever: KV cache scales with max-model-len × concurrent sequences. A model whose weights fit will still OOM if you leave the full 128K context on and let
many long requests batch. Cap --max-model-len to what you actually need.The mental model: KV_pool = gpu_memory_utilization × VRAM − weights, and
concurrent_sequences ≈ KV_pool ÷ (bytes_per_token × context_len). Too-long context or too-high
utilization eats the pool. --max-num-seqs and --max-num-batched-tokens cap the batch to trade
latency vs throughput. Full KV math + an OOM playbook: references/memory-and-throughput.md.
/health requires no API key, so it is the safe first probe on any endpoint you did not just start:
export VLLM_BASE_URL=http://localhost:8000
curl -fsS "$VLLM_BASE_URL/health" && echo " up" # 200, empty body when ready
curl -fsS "$VLLM_BASE_URL/v1/models" \
-H "Authorization: Bearer ${VLLM_API_KEY:-sk-local}" # confirms WHICH model/adapters are served/health says "the server is alive"; /v1/models says "and it is serving the model you expect" (plus
any LoRA adapters). If --api-key is set, /v1/models needs the bearer header but /health never
does. During load a slow first /health is normal — big models take a while to page in.
ollama skill.--max-model-len. Leaving the full context window on is the top OOM cause; KV cache scales
with context × concurrency (engine args).--gpu-memory-utilization for OOM, not throughput first. Lower it if OOM at load/burst;
raise it (carefully) for more concurrency. It is a fraction of one GPU's memory.--api-key = open endpoint on 0.0.0.0:8000.--max-lora-rank must be ≥ the adapter's trained rank, and base + adapter must match, or the LoRA
fails to load.ollama. The GPU box + autoscaling → runpod /
modal.open-weights — choose the model/size/license/quant to serve. vLLM runs it; it does not pick it.finetuning (and single-GPU unsloth) — produce the LoRA adapter or merged weights that
--enable-lora/vllm serve then hosts. They train; vLLM serves.runpod — rent/provision the GPU box vLLM runs on (bring-your-own-container GPU rental).modal — serverless GPU containers + autoscaling around a vLLM process. The box and scaling,
vs vLLM the engine.ollama — the single-user / laptop counterpart. No concurrency to exploit → use it, not vLLM.together-fireworks / huggingface (you call an API; here you
run the server).ollama / a hosted API).vllm serve <model> up; /health returns 200 and /v1/models shows the expected model.--api-key (or VLLM_API_KEY) set if the endpoint is reachable beyond localhost.--max-model-len capped to real need; --gpu-memory-utilization set with OOM headroom.--max-lora-rank ≥ trained rank; adapter reachable by name via the model field.base_url=<server>/v1; model matches a /v1/models id.vllm serve
flags, the full endpoint catalog with example requests, parallelism sizing, and the LoRA runtime API.© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/vllm of ericrisco/rsc-harness.
Open the folder on GitHubat commit e3d5b33
Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm this skillericrisco/rsc-harness | 174 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Aider DelegateamElnagdy/delegate-skills | 2.3k | 2 repos | ~3k | Automated safety check: Pass | MIT | |
| Resolvealexziskind1/model-shelf | 130 | — | ~792 | Automated safety check: Pass | MIT | |
| Mesh APImr-tbot/mesh-api | 180 | — | ~1.8k | Automated safety check: Pass | GPL-3.0 | |
| Hqq QuantizationOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Vllm Deploy K8svllm-project/vllm-skills | 103 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
mr-tbot/mesh-api
Interact with a Meshtastic LoRa mesh network through MESH-API — list nodes, read messages, send texts, and check connection status.
Orchestra-Research/AI-Research-SKILLs
Half-Quadratic Quantization for LLMs without calibration data.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…. Vllm is an agent skill from ericrisco/rsc-harness. Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, loading quantized weights, serving one or many LoRA adapters, and debugging KV-cache OOM from memory-utilisation and context-length flags.
Vllm fits situations like: self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint; splitting a model across GPUs with tensor; pipeline parallelism; loading quantized weights.
Run `npx skills add ericrisco/rsc-harness --skill vllm -a claude-code`. Or copy the skill folder (skills/vllm in ericrisco/rsc-harness) into .claude/skills/vllm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill vllm -a codex`. Or copy the skill folder (skills/vllm in ericrisco/rsc-harness) into .agents/skills/vllm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm, .gemini/skills/vllm, .github/skills/vllm and .opencode/skills/vllm in your project.
Going by SKILL.md and its folder, Vllm needs the command-line tools its instructions call (curl and pip) and credentials named VLLM_API_KEY. Our summary lists: Python 3; A credential in VLLM_API_KEY.
SKILL.md names 3 domains. As links in the text: docs.vllm.ai, pypi.org and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vllm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Vllm: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Resolve (alexziskind1/model-shelf, 130 stars), Mesh API (mr-tbot/mesh-api, 180 stars) and Hqq Quantization (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 174 GitHub stars. The repository holds 233 skills in this directory. The repository was last updated on October 7, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.