AI model or service
vLLM agent skills for Claude Code, Codex and other agents.
- skills
- 163
- official
- 21
- Type
- AI model or service
- Website
- docs.vllm.ai
- Official GitHub
- vllm-project
- Reviews
- See vLLM on Enlisted
vLLM skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
Official (21 skills)
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. | huggingface/ | 11k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 2 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 3 | Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups. | huggingface/ | 11k | 1 repo | ~6.9k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 4 | Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. | NVIDIA/ | 15k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor). | intel/ | 1.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 6 | Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning. | aws-samples/ | 113 | — | ~5k | Automated safety check: Pass | MIT-0 | yesterday |
| 7 | Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling. | oracle/ | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 8 | Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK. | oracle/ | 125 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 9 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 10 | Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations. | oracle/ | 125 | — | ~1.8k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 11 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.5k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | today |
| 12 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.5k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters. | google/ | 21k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 14 | Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin. | NVIDIA/ | 3.5k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | today |
| 15 | Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson. | NVIDIA/ | 3.5k | 1 repo | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 16 | Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output. | NVIDIA/ | 3.5k | 1 repo | ~3.1k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN… | grafana/ | 278 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 18 | Set up AI Runway on AKS — from bare cluster to running model. | microsoft/ | 255 | 1 repo | ~1k | Automated safety check: Pass | MIT | today |
| 19 | Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck. | NVIDIA/ | 3.5k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether… | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | today |
| 21 | How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | today |
Community
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 22 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 23 | Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself. | amElnagdy/ | 2.3k | 2 repos | ~3k | Automated safety check: Pass | MIT | 17 days ago |
| 24 | Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints. | Orchestra-Research/ | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 25 | Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements. | vllm-project/ | 2.9k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm. | guqiong96/ | 464 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 27 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | today |
| 28 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 464 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 29 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 30 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 900 | — | ~2.5k | Automated safety check: Pass | No licence | 2 days ago |
| 31 | Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. | vllm-project/ | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 32 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 33 | 33.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 34 | 34.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 35 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 900 | — | ~2.8k | Automated safety check: Pass | No licence | 2 days ago |
| 36 | Low-token Codex session/thread title organizer. An agent skill from David-Lzy/codex_session_renamer. | David-Lzy/ | 108 | — | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 37 | Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems. | ModelCloud/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Unknown | today |
| 38 | 38.Add Model Adapt and port new LLM model architectures to this xinfer project. | guoqingbao/ | 333 | — | ~4.2k | Automated safety check: Notes | MIT | 28 days ago |
| 39 | Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu. | apache/ | 568 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 40 | 40.Resolve Always resolve Hugging Face models via model-shelf before any download. | alexziskind1/ | 130 | — | ~792 | Automated safety check: Pass | MIT | 1 mo ago |
| 41 | Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models. | Orchestra-Research/ | 13k | 10 repos | ~4k | Automated safety check: Pass | MIT | 3 mo ago |
| 42 | Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels. | MetaX-MACA/ | 179 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 43 | 43.Build Zendnn Build the standalone ZenDNN native library (zendnnl) from source with the alternate compute backends OFF (no oneDNN, libxsmm, parlooper, fbgemm), keeping AOCL DLP, which is the GEMM backend zendnnl… | amd/ | 158 | — | ~2k | Automated safety check: Pass | Unknown | yesterday |
| 44 | Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence. | ThinkFlowLab/ | 138 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 45 | 45.Check Model Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer. | guoqingbao/ | 333 | — | ~3.8k | Automated safety check: Pass | MIT | 28 days ago |
| 46 | Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats. | Orchestra-Research/ | 13k | 9 repos | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 47 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 48 | Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. | guqiong96/ | 464 | 1 repo | ~831 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
Questions, answered from the data.
What is the best vLLM skill?
SageMaker Serving Image Selection (official) from huggingface/skills ranks first of the 163 vLLM skills listed here, with the highest score: its repository has 11k GitHub stars, 1 other GitHub owner carry a copy, its SKILL.md loads about 4.6k tokens and it passes the automated safety check with no findings. Next come Hugging Face Local Model Evals and SageMaker Production Defaults.
Is there an official vLLM skill?
21 of the 163 vLLM skills are official, published by the vendor's own GitHub organization: SageMaker Serving Image Selection, Hugging Face Local Model Evals, SageMaker Production Defaults, Debug Inference, Add Export Format and 16 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.