AI model or service

vLLM agent skills, page 4

Skills #145–163 of 163, ranked by score.

vLLM skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

vLLM skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
145

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

BagelHole/DevOps-Security-Agent-Skills1.1k—~2kAutomated safety check: PassMIT4 mo ago
146
146.Vllm

Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

AlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT6 mo ago
147

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.

uw-syfi/vibesys103—~2.9kAutomated safety check: PassMITyesterday
148

AISBench Benchmark - AI model evaluation tool for Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licencetoday
149

自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licencetoday
150

Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend…

VectorSpaceLab/AREX-Skill328—~823Automated safety check: PassApache-2.01 mo ago
151

昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…

ascend-ai-coding/awesome-ascend-skills174—~1.2kAutomated safety check: PassNo licencetoday
152

模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licencetoday
153

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licencetoday
154

Reference guide for using external AI models via claudish CLI.

MadAppGang/claude-code283—~1.3kAutomated safety check: PassMIT6 mo ago
155

Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

majiayu000/claude-skill-registry6661 repo~1.7kAutomated safety check: PassMITtoday
156

Build and operate Retrieval-Augmented Generation (RAG) infrastructure with vector stores, embedding pipelines, and hybrid search.

majiayu000/claude-skill-registry6661 repo~2.1kAutomated safety check: PassMITtoday
157

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

open-thoughts/OpenThoughts-Agent301—~1.8kAutomated safety check: WarnApache-2.09 days ago
158

A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

ericrisco/rsc-harness156—~4.1kAutomated safety check: PassMITyesterday
159

A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

ericrisco/rsc-harness156—~3.6kAutomated safety check: PassMITyesterday
160

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

magnus919/agent-skills111—~1.5kAutomated safety check: PassMITyesterday
161

Full pipeline where an agent team collaborates to develop an LLM app.

revfactory/harness-1001.3k—~1.9kAutomated safety check: PassApache-2.06 mo ago
162

LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor).

pproenca/dot-skills214—~1.7kAutomated safety check: PassMIT1 mo ago
163

LLM serving expert: vLLM, TensorRT-LLM, Triton Inference Server, quantization (INT8/FP8/GPTQ/AWQ), continuous batching, PagedAttention, KV cache management.

theneoai/awesome-skills183—~3.4kAutomated safety check: PassMIT4 mo ago