Search

vLLM

173 skills found, page 4.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
145

EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated…

open-thoughts/OpenThoughts-Agent301—~977Automated safety check: PassApache-2.012 days ago
146

Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics.

benchflow-ai/skillsbench1.8k—~2.3kAutomated safety check: PassApache-2.02 mo ago
147

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

NVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.0yesterday
148

End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.

amd/Quark182—~6.2kAutomated safety check: PassMIT13 days ago
149

This skill should be used when the user asks about "local models", "custom models", "fine-tuning", "self-hosting models", "model selection", "which model should I use", "data privacy and models"…

Habitat-Thinking/ai-literacy-superpowers114—~1kAutomated safety check: PassUnknown20 days ago
150

vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday
151

Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…

ascend-ai-coding/awesome-ascend-skills174—~731Automated safety check: PassNo licenceyesterday
152

Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure…

magnus919/agent-skills115—~4.2kAutomated safety check: NotesMITyesterday
153
153.Vllm

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

magnus919/agent-skills115—~4.1kAutomated safety check: NotesMITyesterday
154

Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve.

ascend-ai-coding/awesome-ascend-skills174—~5.6kAutomated safety check: PassNo licenceyesterday
155

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

AnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMITyesterday
156
156.Vllm

Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

AlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT6 mo ago
157

Deploy a GPU workload on CoreWeave with kubectl. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMITyesterday
158

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago
159

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago
160

Build and operate Retrieval-Augmented Generation (RAG) infrastructure with vector stores, embedding pipelines, and hybrid search.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2.1kAutomated safety check: PassMIT4 mo ago
161

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.

uw-syfi/vibesys105—~2.9kAutomated safety check: PassMITyesterday
162

AISBench Benchmark - AI model evaluation tool for Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday
163

自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licenceyesterday
164

Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend…

VectorSpaceLab/AREX-Skill331—~823Automated safety check: PassApache-2.01 mo ago
165

昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…

ascend-ai-coding/awesome-ascend-skills174—~1.2kAutomated safety check: PassNo licenceyesterday
166

模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licenceyesterday
167

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday
168

Reference guide for using external AI models via claudish CLI.

MadAppGang/claude-code285—~1.3kAutomated safety check: PassMIT7 mo ago
169

A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

ericrisco/rsc-harness180—~4.1kAutomated safety check: PassMITyesterday
170

A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

ericrisco/rsc-harness180—~3.6kAutomated safety check: PassMITyesterday
171
171.Vllm

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

ericrisco/rsc-harness180—~3.6kAutomated safety check: PassMITyesterday
172

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

open-thoughts/OpenThoughts-Agent301—~1.8kAutomated safety check: WarnApache-2.012 days ago
173

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

magnus919/agent-skills115—~1.5kAutomated safety check: PassMITyesterday