Search
AI & LLM Engineering · vLLM
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | 145.Litellm Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure… | magnus919/ | 115 | — | ~4.2k | Automated safety check: Notes | MIT | yesterday |
| 146 | 146.Vllm Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API… | magnus919/ | 115 | — | ~4.1k | Automated safety check: Notes | MIT | yesterday |
| 147 | 147.Vllm Bench Serve Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. | ascend-ai-coding/ | 174 | — | ~5.6k | Automated safety check: Pass | No licence | yesterday |
| 148 | 148.ML Research Lab Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability. | AnastasiyaW/ | 154 | — | ~794 | Automated safety check: Pass | MIT | 2 days ago |
| 149 | 149.Vllm Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support. | AlexAI-MCP/ | 135 | — | ~2.3k | Automated safety check: Pass | MIT | 6 mo ago |
| 150 | Deploy a GPU workload on CoreWeave with kubectl. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 151 | 151.LLM Gateway Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 152 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 153 | Build and operate Retrieval-Augmented Generation (RAG) infrastructure with vector stores, embedding pipelines, and hybrid search. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 154 | 154.Serving Systems LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 105 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 155 | 155.Ais Bench AISBench Benchmark - AI model evaluation tool for Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 156 | 156.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 157 | Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend… | VectorSpaceLab/ | 331 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 158 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 159 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 160 | Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 161 | Reference guide for using external AI models via claudish CLI. | MadAppGang/ | 285 | — | ~1.3k | Automated safety check: Pass | MIT | 7 mo ago |
| 162 | 162.Open Weights A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping. | ericrisco/ | 180 | — | ~4.1k | Automated safety check: Pass | MIT | 2 days ago |
| 163 | 163.Unsloth A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the… | ericrisco/ | 180 | — | ~3.6k | Automated safety check: Pass | MIT | 2 days ago |
| 164 | 164.Vllm A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or… | ericrisco/ | 180 | — | ~3.6k | Automated safety check: Pass | MIT | 2 days ago |
| 165 | Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. | open-thoughts/ | 301 | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | 13 days ago |
| 166 | 166.ML Engineering Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity… | magnus919/ | 115 | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |