Search
LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 337 | 337.Llama Cpp Run quantized LLMs locally with llama.cpp — CPU+GPU inference, GGUF format, OpenAI-compatible server, and Python bindings. | AlexAI-MCP/ | 135 | — | ~2.3k | Automated safety check: Pass | MIT | 6 mo ago |
| 338 | 338.Vllm Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support. | AlexAI-MCP/ | 135 | — | ~2.3k | Automated safety check: Pass | MIT | 6 mo ago |
| 339 | Build model quantization tool operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~571 | Automated safety check: Pass | MIT | yesterday |
| 340 | Deploy KServe InferenceService on CoreWeave with autoscaling and GPU scheduling. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 341 | Set up local development workflow for CoreWeave GPU deployments. | jeremylongshore/ | 2.8k | — | ~927 | Automated safety check: Notes | MIT | yesterday |
| 342 | Optimize CoreWeave GPU inference latency and throughput. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 343 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 344 | Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. | BagelHole/ | 1.2k | — | ~2.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 345 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 346 | 346.LLM Aiops Guide Papers on LLMs for IT operations and AIOps research. An agent skill from wentorai/research-plugins. | wentorai/ | 298 | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 347 | 347.Local LLM Router Route AI coding queries to local LLMs in air-gapped networks. | hoodini/ | 282 | — | ~20k | Automated safety check: Pass | No licence | 3 mo ago |
| 348 | 348.Serving Systems LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 105 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 349 | Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring. | curiositech/ | 244 | — | ~3.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 350 | 350.Ais Bench AISBench Benchmark - AI model evaluation tool for Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 351 | 351.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 352 | Cloudflare Workers AI for serverless GPU inference. An agent skill from secondsky/claude-skills. | secondsky/ | 227 | — | ~2.4k | Automated safety check: Pass | MIT | 13 days ago |
| 353 | Use this sub-skill for Torch-TensorRT model compilation, dynamic input planning, torch.export workflows, save/load formats, raw TensorRT engines, and compile-time troubleshooting. | VectorSpaceLab/ | 331 | — | ~1.3k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 354 | Use torchtune generation, Eleuther evaluation, and quantization workflows safely after checkpoints exist. | VectorSpaceLab/ | 331 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 355 | Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark… | VectorSpaceLab/ | 331 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 356 | Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend… | VectorSpaceLab/ | 331 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 357 | 357.Torch Tensorrt A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging… | VectorSpaceLab/ | 331 | — | ~1.5k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 358 | A skill your agent uses when planning or auditing TorchVision reference training/evaluation workflows for classification, quantization, detection, segmentation, video classification, optical flow… | VectorSpaceLab/ | 331 | — | ~684 | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 359 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 360 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 361 | Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 362 | 362.Model Discovery Fetch current model names from AI providers (Anthropic, OpenAI, Gemini, Ollama), classify them into tiers (fast/default/heavy), and detect new models. | aiskillstore/ | 433 | — | ~1.9k | Automated safety check: Pass | No licence | yesterday |
| 363 | Select appropriate Ollama models for processing sensitive but legal content. | divinevideo/ | 266 | — | ~1.3k | Automated safety check: Pass | MPL-2.0 | yesterday |
| 364 | 364.Open Weights A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping. | ericrisco/ | 180 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 365 | 365.Unsloth A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the… | ericrisco/ | 180 | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 366 | 366.Vllm A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or… | ericrisco/ | 180 | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 367 | 367.Tmux Agents Manage background coding agents in tmux sessions. An agent skill from sundial-org/awesome-openclaw-skills. | sundial-org/ | 663 | — | ~933 | Automated safety check: Pass | No licence | 7 mo ago |
| 368 | Delegates tasks to a locally served Muse Glimmer via ollama. | athola/ | 341 | — | ~988 | Automated safety check: Notes | MIT | yesterday |
| 369 | 369.Llama Cpp Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems. | magnus919/ | 115 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 370 | 370.ML Engineering Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity… | magnus919/ | 115 | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 371 | Offline experimental post-call feedback scorer using supplied ratings and text heuristics. | CALLE-AI/ | 107 | — | ~727 | Automated safety check: Pass | MIT | yesterday |
| 372 | Build machine vision inspection systems with MATLAB Visual Inspection Toolbox. | matlab/ | 1.1k | — | ~3.1k | Automated safety check: Pass | Unknown | 2 days ago |
| 373 | Fully local multi-agent swarm intelligence simulation engine using Neo4j + Ollama for public opinion, market sentiment, and social dynamics prediction. | LeoYeAI/ | 2.2k | — | ~4.1k | Automated safety check: Notes | MIT | 2 mo ago |