Search

AI & LLM Engineering · vLLM

166 skills found, page 4.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
145

Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure…

magnus919/agent-skills115—~4.2kAutomated safety check: NotesMITyesterday
146
146.Vllm

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

magnus919/agent-skills115—~4.1kAutomated safety check: NotesMITyesterday
147

Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve.

ascend-ai-coding/awesome-ascend-skills174—~5.6kAutomated safety check: PassNo licenceyesterday
148

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

AnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMIT2 days ago
149
149.Vllm

Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

AlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT6 mo ago
150

Deploy a GPU workload on CoreWeave with kubectl. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMITyesterday
151

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago
152

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago
153

Build and operate Retrieval-Augmented Generation (RAG) infrastructure with vector stores, embedding pipelines, and hybrid search.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2.1kAutomated safety check: PassMIT4 mo ago
154

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.

uw-syfi/vibesys105—~2.9kAutomated safety check: PassMITyesterday
155

AISBench Benchmark - AI model evaluation tool for Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday
156

自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licenceyesterday
157

Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend…

VectorSpaceLab/AREX-Skill331—~823Automated safety check: PassApache-2.01 mo ago
158

昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…

ascend-ai-coding/awesome-ascend-skills174—~1.2kAutomated safety check: PassNo licenceyesterday
159

模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licenceyesterday
160

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday
161

Reference guide for using external AI models via claudish CLI.

MadAppGang/claude-code285—~1.3kAutomated safety check: PassMIT7 mo ago
162

A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

ericrisco/rsc-harness180—~4.1kAutomated safety check: PassMIT2 days ago
163

A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

ericrisco/rsc-harness180—~3.6kAutomated safety check: PassMIT2 days ago
164
164.Vllm

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

ericrisco/rsc-harness180—~3.6kAutomated safety check: PassMIT2 days ago
165

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

open-thoughts/OpenThoughts-Agent301—~1.8kAutomated safety check: WarnApache-2.013 days ago
166

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

magnus919/agent-skills115—~1.5kAutomated safety check: PassMITyesterday