Search
vLLM
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated… | open-thoughts/ | 301 | — | ~977 | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 146 | Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics. | benchflow-ai/ | 1.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 147 | How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 148 | End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. | amd/ | 182 | — | ~6.2k | Automated safety check: Pass | MIT | 13 days ago |
| 149 | This skill should be used when the user asks about "local models", "custom models", "fine-tuning", "self-hosting models", "model selection", "which model should I use", "data privacy and models"… | Habitat-Thinking/ | 114 | — | ~1k | Automated safety check: Pass | Unknown | 20 days ago |
| 150 | 150.Vllm Ascend vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 151 | Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode… | ascend-ai-coding/ | 174 | — | ~731 | Automated safety check: Pass | No licence | yesterday |
| 152 | 152.Litellm Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure… | magnus919/ | 115 | — | ~4.2k | Automated safety check: Notes | MIT | yesterday |
| 153 | 153.Vllm Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API… | magnus919/ | 115 | — | ~4.1k | Automated safety check: Notes | MIT | yesterday |
| 154 | 154.Vllm Bench Serve Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. | ascend-ai-coding/ | 174 | — | ~5.6k | Automated safety check: Pass | No licence | yesterday |
| 155 | 155.ML Research Lab Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability. | AnastasiyaW/ | 154 | — | ~794 | Automated safety check: Pass | MIT | yesterday |
| 156 | 156.Vllm Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support. | AlexAI-MCP/ | 135 | — | ~2.3k | Automated safety check: Pass | MIT | 6 mo ago |
| 157 | Deploy a GPU workload on CoreWeave with kubectl. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 158 | 158.LLM Gateway Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 159 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 160 | Build and operate Retrieval-Augmented Generation (RAG) infrastructure with vector stores, embedding pipelines, and hybrid search. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 161 | 161.Serving Systems LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 105 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 162 | 162.Ais Bench AISBench Benchmark - AI model evaluation tool for Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 163 | 163.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 164 | Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend… | VectorSpaceLab/ | 331 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 165 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 166 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 167 | Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 168 | Reference guide for using external AI models via claudish CLI. | MadAppGang/ | 285 | — | ~1.3k | Automated safety check: Pass | MIT | 7 mo ago |
| 169 | 169.Open Weights A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping. | ericrisco/ | 180 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 170 | 170.Unsloth A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the… | ericrisco/ | 180 | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 171 | 171.Vllm A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or… | ericrisco/ | 180 | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 172 | Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. | open-thoughts/ | 301 | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | 12 days ago |
| 173 | 173.ML Engineering Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity… | magnus919/ | 115 | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |