Aider Delegate
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
vLLM: high-throughput LLM serving, OpenAI API, quantization.
$ npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent serving-llms-vllm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm .claude/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm into .claude/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent serving-llms-vllm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm .agents/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm into .agents/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent serving-llms-vllm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm .cursor/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm into .cursor/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Luciole-Studio/Misaka-Agent.git --path misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent serving-llms-vllm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm .gemini/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm into .gemini/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Luciole-Studio/Misaka-Agent serving-llms-vllmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm .github/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm into .github/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Luciole-Studio/Misaka-Agent serving-llms-vllm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm .opencode/skills/serving-llms-vllm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "serving-llms-vllm" agent skill from https://github.com/Luciole-Studio/Misaka-Agent/tree/main/misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm into .opencode/skills/serving-llms-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "serving-llms-vllm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
serving-llms-vllmvLLM: high-throughput LLM serving, OpenAI API, quantization.
Serving LLMs Vllm is an agent skill from Luciole-Studio/Misaka-Agent. vLLM: high-throughput LLM serving, OpenAI API, quantization.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/optimization.md`, `references/quantization.md` and `references/server-deployment.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and OpenAI. The repository describes itself as: A multi-agent research system for the humanities and social sciences. The licence is MIT.
Read from SKILL.md and the folder at commit 3bcf7a3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pippythoncurldockerFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.vllm.aigithub.comdiscuss.vllm.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Serving LLMs Vllm loads about 2.3k tokens when it runs, and up to ~9k if it reads all its reference files. Until then it costs about 20 tokens; SKILL.md has 534 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Luciole-Studio/Misaka-Agent at commit 3bcf7a3, republished under its MIT licence (© Luciole-Studio). 534 words, ~2,334 tokens.
.claude/skills/serving-llms-vllm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
vLLM achieves 24x higher throughput than standard transformers through PagedAttention (block-based KV cache) and continuous batching (mixing prefill/decode requests).
Installation:
pip install vllmBasic offline inference:
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct")
sampling = SamplingParams(temperature=0.7, max_tokens=256)
outputs = llm.generate(["Explain quantum computing"], sampling)
print(outputs[0].outputs[0].text)OpenAI-compatible server:
vllm serve meta-llama/Meta-Llama-3-8B-Instruct
# Query with OpenAI SDK
python -c "
from openai import OpenAI
client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY')
print(client.chat.completions.create(
model='meta-llama/Meta-Llama-3-8B-Instruct',
messages=[{'role': 'user', 'content': 'Hello!'}]
).choices[0].message.content)
"Copy this checklist and track progress:
Deployment Progress:
- [ ] Step 1: Configure server settings
- [ ] Step 2: Test with limited traffic
- [ ] Step 3: Enable monitoring
- [ ] Step 4: Deploy to production
- [ ] Step 5: Verify performance metricsStep 1: Configure server settings
Choose configuration based on your model size:
# For 7B-13B models on single GPU
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
--gpu-memory-utilization 0.9 \
--max-model-len 8192 \
--port 8000
# For 30B-70B models with tensor parallelism
vllm serve meta-llama/Meta-Llama-3-70B-Instruct \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.9 \
--quantization awq \
--port 8000
# For production with caching (Prometheus metrics are exposed
# automatically at /metrics on the API port)
vllm serve meta-llama/Meta-Llama-3-8B-Instruct \
--gpu-memory-utilization 0.9 \
--enable-prefix-caching \
--port 8000 \
--host 0.0.0.0Step 2: Test with limited traffic
Run load test before production:
# Install load testing tool
pip install locust
# Create test_load.py with sample requests
# Run: locust -f test_load.py --host http://localhost:8000Verify TTFT (time to first token) < 500ms and throughput > 100 req/sec.
Step 3: Enable monitoring
vLLM exposes Prometheus metrics at /metrics on the API port (default 8000):
curl http://localhost:8000/metrics | grep vllmKey metrics to monitor:
vllm:time_to_first_token_seconds - Latencyvllm:num_requests_running - Active requestsvllm:gpu_cache_usage_perc - KV cache utilizationStep 4: Deploy to production
Use Docker for consistent deployment:
# Run vLLM in Docker
docker run --gpus all -p 8000:8000 \
vllm/vllm-openai:latest \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--gpu-memory-utilization 0.9 \
--enable-prefix-cachingStep 5: Verify performance metrics
Check that deployment meets targets:
For processing large datasets without server overhead.
Copy this checklist:
Batch Processing:
- [ ] Step 1: Prepare input data
- [ ] Step 2: Configure LLM engine
- [ ] Step 3: Run batch inference
- [ ] Step 4: Process resultsStep 1: Prepare input data
# Load prompts from file
prompts = []
with open("prompts.txt") as f:
prompts = [line.strip() for line in f]
print(f"Loaded {len(prompts)} prompts")Step 2: Configure LLM engine
from vllm import LLM, SamplingParams
llm = LLM(
model="meta-llama/Meta-Llama-3-8B-Instruct",
tensor_parallel_size=2, # Use 2 GPUs
gpu_memory_utilization=0.9,
max_model_len=4096
)
sampling = SamplingParams(
temperature=0.7,
top_p=0.95,
max_tokens=512,
stop=["</s>", "\n\n"]
)Step 3: Run batch inference
vLLM automatically batches requests for efficiency:
# Process all prompts in one call
outputs = llm.generate(prompts, sampling)
# vLLM handles batching internally
# No need to manually chunk promptsStep 4: Process results
# Extract generated text
results = []
for output in outputs:
prompt = output.prompt
generated = output.outputs[0].text
results.append({
"prompt": prompt,
"generated": generated,
"tokens": len(output.outputs[0].token_ids)
})
# Save to file
import json
with open("results.jsonl", "w") as f:
for result in results:
f.write(json.dumps(result) + "\n")
print(f"Processed {len(results)} prompts")Fit large models in limited GPU memory.
Quantization Setup:
- [ ] Step 1: Choose quantization method
- [ ] Step 2: Find or create quantized model
- [ ] Step 3: Launch with quantization flag
- [ ] Step 4: Verify accuracyStep 1: Choose quantization method
Step 2: Find or create quantized model
Use pre-quantized models from HuggingFace:
# Search for AWQ models
# Example: TheBloke/Llama-2-70B-AWQStep 3: Launch with quantization flag
# Using pre-quantized model
vllm serve TheBloke/Llama-2-70B-AWQ \
--quantization awq \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.95
# Results: 70B model in ~40GB VRAMStep 4: Verify accuracy
Test outputs match expected quality:
# Compare quantized vs non-quantized responses
# Verify task-specific performance unchangedUse vLLM when:
Use alternatives instead:
Issue: Out of memory during model loading
Reduce memory usage:
vllm serve MODEL \
--gpu-memory-utilization 0.7 \
--max-model-len 4096Or use quantization:
vllm serve MODEL --quantization awqIssue: Slow first token (TTFT > 1 second)
Enable prefix caching for repeated prompts:
vllm serve MODEL --enable-prefix-cachingFor long prompts, enable chunked prefill:
vllm serve MODEL --enable-chunked-prefillIssue: Model not found error
Use --trust-remote-code for custom models:
vllm serve MODEL --trust-remote-codeIssue: Low throughput (<50 req/sec)
Increase concurrent sequences:
vllm serve MODEL --max-num-seqs 512Check GPU utilization with nvidia-smi - should be >80%.
Issue: Inference slower than expected
Verify tensor parallelism uses power of 2 GPUs:
vllm serve MODEL --tensor-parallel-size 4 # Not 3Enable speculative decoding for faster generation (pass config as JSON;
--speculative-model was removed in favor of --speculative-config):
vllm serve MODEL \
--speculative-config '{"model": "DRAFT_MODEL", "num_speculative_tokens": 5, "method": "draft_model"}'Server deployment patterns: See references/server-deployment.md for Docker, Kubernetes, and load balancing configurations.
Performance optimization: See references/optimization.md for PagedAttention tuning, continuous batching details, and benchmark results.
Quantization guide: See references/quantization.md for AWQ/GPTQ/FP8 setup, model preparation, and accuracy comparisons.
Troubleshooting: See references/troubleshooting.md for detailed error messages, debugging steps, and performance diagnostics.
Supported platforms: NVIDIA (primary), AMD ROCm, Intel GPUs, TPUs
© Luciole-Studio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm of Luciole-Studio/Misaka-Agent.
Open the folder on GitHubat commit 3bcf7a3
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Luciole-Studio/Misaka-Agent, which our catalogue first saw on October 7, 2026.
Serving LLMs Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Serving LLMs Vllm this skillLuciole-Studio/Misaka-Agent | 158 | 2 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Aider DelegateamElnagdy/delegate-skills | 2.3k | 2 repos | ~3k | Automated safety check: Pass | MIT | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Vllm Bench Random Syntheticvllm-project/vllm-skills | 103 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Vllm Bench Servevllm-project/vllm-skills | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 |
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
vllm-project/vllm-skills
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.
vllm-project/vllm-skills
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
Luciole-Studio/Misaka-Agent
Plan and run multi-agent video production pipelines. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
AST-aware structural code search and rewrite via ast-grep. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Drug discovery: ChEMBL search, drug-likeness, interactions. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Workout planning, macros, and body metrics via wger/USDA. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Render MP4/WebM videos from HTML compositions. An agent skill from Luciole-Studio/Misaka-Agent.
Luciole-Studio/Misaka-Agent
Follow the money via public records and sanctions data. An agent skill from Luciole-Studio/Misaka-Agent.
Categories
vLLM: high-throughput LLM serving, OpenAI API, quantization. Serving LLMs Vllm is an agent skill from Luciole-Studio/Misaka-Agent. vLLM: high-throughput LLM serving, OpenAI API, quantization.
Serving LLMs Vllm fits situations like: tasks that involve LLM inference and serving.
Run `npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a claude-code`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm in Luciole-Studio/Misaka-Agent) into .claude/skills/serving-llms-vllm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a codex`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/inference/serving-llms-vllm in Luciole-Studio/Misaka-Agent) into .agents/skills/serving-llms-vllm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Luciole-Studio/Misaka-Agent --skill serving-llms-vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/serving-llms-vllm, .gemini/skills/serving-llms-vllm, .github/skills/serving-llms-vllm and .opencode/skills/serving-llms-vllm in your project.
Going by SKILL.md and its folder, Serving LLMs Vllm needs the command-line tools its instructions call (pip, python, curl and docker). Our summary lists: Python 3; Docker.
SKILL.md names 3 domains. As links in the text: docs.vllm.ai, github.com and discuss.vllm.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Serving LLMs Vllm is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Serving LLMs Vllm: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Bench Random Synthetic (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Luciole-Studio (a GitHub organization) maintains it in Luciole-Studio/Misaka-Agent, which has 158 GitHub stars. The repository holds 77 skills in this directory. The repository was last updated on October 8, 2026.
Source: Luciole-Studio/Misaka-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.