Topic · AI & LLM Engineering
Best LLM inference and serving skills, page 8
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 337 | Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark… | VectorSpaceLab/ | 328 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 338 | Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend… | VectorSpaceLab/ | 328 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 339 | 339.Torch Tensorrt A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging… | VectorSpaceLab/ | 328 | — | ~1.5k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 340 | A skill your agent uses when planning or auditing TorchVision reference training/evaluation workflows for classification, quantization, detection, segmentation, video classification, optical flow… | VectorSpaceLab/ | 328 | — | ~684 | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 341 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | today |
| 342 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |
| 343 | Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | today |
| 344 | 344.Tmux Agents Manage background coding agents in tmux sessions. An agent skill from sundial-org/awesome-openclaw-skills. | sundial-org/ | 663 | — | ~933 | Automated safety check: Pass | No licence | 7 mo ago |
| 345 | Select appropriate Ollama models for processing sensitive but legal content. | divinevideo/ | 265 | — | ~1.3k | Automated safety check: Pass | MPL-2.0 | today |
| 346 | 346.Groq Expert guidance for Groq, the LLM inference platform that provides the fastest token generation speeds available, powered by custom LPU (Language Processing Unit) hardware. | majiayu000/ | 666 | 1 repo | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 347 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | majiayu000/ | 666 | 1 repo | ~2.1k | Automated safety check: Pass | MIT | today |
| 348 | 348.Ray Framework for scaling Python applications from a laptop to a cluster. | majiayu000/ | 666 | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 349 | NVIDIA TensorRT model optimization and deployment. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 1 repo | ~2.4k | Automated safety check: Notes | MIT | today |
| 350 | Delegates tasks to a locally served Muse Glimmer via ollama. | athola/ | 342 | — | ~988 | Automated safety check: Notes | MIT | yesterday |
| 351 | 351.Open Weights A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping. | ericrisco/ | 156 | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 352 | 352.Unsloth A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the… | ericrisco/ | 156 | — | ~3.6k | Automated safety check: Pass | MIT | today |
| 353 | Offline experimental post-call feedback scorer using supplied ratings and text heuristics. | CALLE-AI/ | 106 | — | ~727 | Automated safety check: Pass | MIT | 3 days ago |
| 354 | Build machine vision inspection systems with MATLAB Visual Inspection Toolbox. | matlab/ | 1.1k | — | ~3.1k | Automated safety check: Pass | Unknown | 7 days ago |
| 355 | 355.Llama Cpp Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems. | magnus919/ | 111 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 356 | 356.ML Engineering Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity… | magnus919/ | 111 | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 357 | Fully local multi-agent swarm intelligence simulation engine using Neo4j + Ollama for public opinion, market sentiment, and social dynamics prediction. | LeoYeAI/ | 2.2k | — | ~4.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 358 | 358.LLM App Builder Full pipeline where an agent team collaborates to develop an LLM app. | revfactory/ | 1.3k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 359 | 359.Ray Production Ray (open-source, pinned to 2.57) for classic-ML workloads from training to serving — Ray Train, Tune, Data, Serve, Core, and cluster deployment on KubeRay. | pproenca/ | 214 | — | ~2.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 360 | 360.Ray LLM LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor). | pproenca/ | 214 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 361 | 361.AI ML Engineer Expert AI/ML Engineer with deep MLOps expertise. An agent skill from theneoai/awesome-skills. | theneoai/ | 183 | — | ~2.9k | Automated safety check: Pass | MIT | 4 mo ago |
| 362 | LLM serving expert: vLLM, TensorRT-LLM, Triton Inference Server, quantization (INT8/FP8/GPTQ/AWQ), continuous batching, PagedAttention, KV cache management. | theneoai/ | 183 | — | ~3.4k | Automated safety check: Pass | MIT | 4 mo ago |
| 363 | 363.Mlflow Expert MLflow expert: experiment tracking, model registry, autologging, MLflow Projects, MLflow Models, model serving, A/B testing, feature store integration. | theneoai/ | 183 | — | ~3.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 364 | Expert robot perception engineer specializing in 3D point cloud processing, multi-modal sensor fusion (camera+LiDAR+IMU), real-time SLAM, and edge-optimized deep learning inference via TensorRT/ONNX… | theneoai/ | 183 | — | ~1.5k | Automated safety check: Pass | MIT | 4 mo ago |
Explore related skills
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23