AI model or service
vLLM agent skills, page 4
vLLM skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | BagelHole/ | 1.1k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 146 | 146.Vllm Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support. | AlexAI-MCP/ | 135 | — | ~2.3k | Automated safety check: Pass | MIT | 6 mo ago |
| 147 | 147.Serving Systems LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 103 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 148 | 148.Ais Bench AISBench Benchmark - AI model evaluation tool for Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | today |
| 149 | 149.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |
| 150 | Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend… | VectorSpaceLab/ | 328 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 151 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | today |
| 152 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |
| 153 | Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | today |
| 154 | Reference guide for using external AI models via claudish CLI. | MadAppGang/ | 283 | — | ~1.3k | Automated safety check: Pass | MIT | 6 mo ago |
| 155 | Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval. | majiayu000/ | 666 | 1 repo | ~1.7k | Automated safety check: Pass | MIT | today |
| 156 | Build and operate Retrieval-Augmented Generation (RAG) infrastructure with vector stores, embedding pipelines, and hybrid search. | majiayu000/ | 666 | 1 repo | ~2.1k | Automated safety check: Pass | MIT | today |
| 157 | Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. | open-thoughts/ | 301 | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | 9 days ago |
| 158 | 158.Open Weights A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping. | ericrisco/ | 156 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 159 | 159.Unsloth A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the… | ericrisco/ | 156 | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 160 | 160.ML Engineering Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity… | magnus919/ | 111 | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 161 | 161.LLM App Builder Full pipeline where an agent team collaborates to develop an LLM app. | revfactory/ | 1.3k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 162 | 162.Ray LLM LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor). | pproenca/ | 214 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 163 | LLM serving expert: vLLM, TensorRT-LLM, Triton Inference Server, quantization (INT8/FP8/GPTQ/AWQ), continuous batching, PagedAttention, KV cache management. | theneoai/ | 183 | — | ~3.4k | Automated safety check: Pass | MIT | 4 mo ago |