AI model or service

vLLM agent skills for Claude Code, Codex and other agents.

High-throughput, memory-efficient inference and serving engine for large language models.
skills
163
official
21
Type
AI model or service
Website
docs.vllm.ai
Official GitHub
vllm-project
Reviews
See vLLM on Enlisted

vLLM skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Official (21 skills)

Official vLLM skills
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

huggingface/skills11k1 repo~4.6kAutomated safety check: PassApache-2.06 days ago
2

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.06 days ago
3

Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

huggingface/skills11k1 repo~6.9kAutomated safety check: PassApache-2.06 days ago
4
4.Debug InferenceOfficial

Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

NVIDIA/OpenShell15k—~1.9kAutomated safety check: PassApache-2.0today
5

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

intel/auto-round1.6k—~1.9kAutomated safety check: PassApache-2.02 days ago
6

Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning.

aws-samples/appmod-blueprints113—~5kAutomated safety check: PassMIT-0yesterday
7
7.Aqua DeploymentOfficial

Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

oracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.01 mo ago
8

Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

oracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.01 mo ago
9
9.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
10

Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations.

oracle/accelerated-data-science125—~1.8kAutomated safety check: PassUPL-1.01 mo ago
11

Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

NVIDIA/skills3.5k1 repo~2.3kAutomated safety check: NotesApache-2.0today
12
12.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.5k1 repo~1.8kAutomated safety check: PassApache-2.0today
13

Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0yesterday
14
14.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.5k2 repos~3kAutomated safety check: NotesApache-2.0today
15

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.5k1 repo~2.9kAutomated safety check: PassApache-2.0today
16

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.5k1 repo~3.1kAutomated safety check: PassApache-2.0today
17
17.ML AIOfficial

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

grafana/skills278—~1.3kAutomated safety check: PassApache-2.0yesterday
18

Set up AI Runway on AKS — from bare cluster to running model.

microsoft/GitHub-Copilot-for-Azure2551 repo~1kAutomated safety check: PassMITtoday
19

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

NVIDIA/skills3.5k1 repo~1.2kAutomated safety check: PassApache-2.0today
20
20.Deepstream SopOfficial

A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

NVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0today
21

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

NVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0today

Community

Community vLLM skills
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
22

Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

sgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0today
23

Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

amElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT17 days ago
24

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

Orchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT3 mo ago
25

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

vllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0today
26

Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

guqiong96/Lvllm4642 repos~349Automated safety check: PassApache-2.015 days ago
27

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0today
28

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4642 repos~1.5kAutomated safety check: PassApache-2.015 days ago
29

Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

wshobson/agents40k1 repo~2kAutomated safety check: PassMIT2 days ago
30

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~2.5kAutomated safety check: PassNo licence2 days ago
31

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

vllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0today
32

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMITyesterday
33

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.09 days ago
34

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
35

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNo licence2 days ago
36

Low-token Codex session/thread title organizer. An agent skill from David-Lzy/codex_session_renamer.

David-Lzy/codex_session_renamer108—~1.2kAutomated safety check: PassMIT3 mo ago
37

Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

ModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassUnknowntoday
38

Adapt and port new LLM model architectures to this xinfer project.

guoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT28 days ago
39

Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu.

apache/dubbo-go-pixiu568—~2.5kAutomated safety check: PassApache-2.04 days ago
40

Always resolve Hugging Face models via model-shelf before any download.

alexziskind1/model-shelf130—~792Automated safety check: PassMIT1 mo ago
41

Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.

Orchestra-Research/AI-Research-SKILLs13k10 repos~4kAutomated safety check: PassMIT3 mo ago
42

Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

MetaX-MACA/vLLM-metax179—~2.6kAutomated safety check: PassApache-2.08 days ago
43

Build the standalone ZenDNN native library (zendnnl) from source with the alternate compute backends OFF (no oneDNN, libxsmm, parlooper, fbgemm), keeping AOCL DLP, which is the GEMM backend zendnnl…

amd/ZenDNN158—~2kAutomated safety check: PassUnknownyesterday
44

Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence.

ThinkFlowLab/vllm-rlt138—~1.1kAutomated safety check: PassApache-2.07 days ago
45

Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

guoqingbao/xinfer333—~3.8kAutomated safety check: PassMIT28 days ago
46

Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats.

Orchestra-Research/AI-Research-SKILLs13k9 repos~1.2kAutomated safety check: PassMIT3 mo ago
47

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

Orchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT3 mo ago
48

Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

guqiong96/Lvllm4641 repo~831Automated safety check: PassApache-2.015 days ago

Questions, answered from the data.

What is the best vLLM skill?

SageMaker Serving Image Selection (official) from huggingface/skills ranks first of the 163 vLLM skills listed here, with the highest score: its repository has 11k GitHub stars, 1 other GitHub owner carry a copy, its SKILL.md loads about 4.6k tokens and it passes the automated safety check with no findings. Next come Hugging Face Local Model Evals and SageMaker Production Defaults.

Is there an official vLLM skill?

21 of the 163 vLLM skills are official, published by the vendor's own GitHub organization: SageMaker Serving Image Selection, Hugging Face Local Model Evals, SageMaker Production Defaults, Debug Inference, Add Export Format and 16 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.