Topic · AI & LLM Engineering

Best LLM inference and serving skills, page 8

Skills #337–364 of 364, ranked by score.

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
337

Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark…

VectorSpaceLab/AREX-Skill328—~1kAutomated safety check: PassBSD-3-Clause1 mo ago
338

Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend…

VectorSpaceLab/AREX-Skill328—~823Automated safety check: PassApache-2.01 mo ago
339

A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging…

VectorSpaceLab/AREX-Skill328—~1.5kAutomated safety check: PassBSD-3-Clause1 mo ago
340

A skill your agent uses when planning or auditing TorchVision reference training/evaluation workflows for classification, quantization, detection, segmentation, video classification, optical flow…

VectorSpaceLab/AREX-Skill328—~684Automated safety check: PassBSD-3-Clause1 mo ago
341

昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…

ascend-ai-coding/awesome-ascend-skills174—~1.2kAutomated safety check: PassNo licencetoday
342

模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licencetoday
343

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licencetoday
344

Manage background coding agents in tmux sessions. An agent skill from sundial-org/awesome-openclaw-skills.

sundial-org/awesome-openclaw-skills663—~933Automated safety check: PassNo licence7 mo ago
345

Select appropriate Ollama models for processing sensitive but legal content.

divinevideo/divine-mobile265—~1.3kAutomated safety check: PassMPL-2.0today
346
346.Groq

Expert guidance for Groq, the LLM inference platform that provides the fastest token generation speeds available, powered by custom LPU (Language Processing Unit) hardware.

majiayu000/claude-skill-registry6661 repo~2.3kAutomated safety check: PassApache-2.0today
347

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

majiayu000/claude-skill-registry6661 repo~2.1kAutomated safety check: PassMITtoday
348
348.Ray

Framework for scaling Python applications from a laptop to a cluster.

majiayu000/claude-skill-registry6661 repo~1.8kAutomated safety check: PassApache-2.0today
349

NVIDIA TensorRT model optimization and deployment. An agent skill from majiayu000/claude-skill-registry.

majiayu000/claude-skill-registry6661 repo~2.4kAutomated safety check: NotesMITtoday
350

Delegates tasks to a locally served Muse Glimmer via ollama.

athola/claude-night-market342—~988Automated safety check: NotesMITyesterday
351

A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

ericrisco/rsc-harness156—~4.1kAutomated safety check: PassMITtoday
352

A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

ericrisco/rsc-harness156—~3.6kAutomated safety check: PassMITtoday
353

Offline experimental post-call feedback scorer using supplied ratings and text heuristics.

CALLE-AI/awesome-phone-call-agents106—~727Automated safety check: PassMIT3 days ago
354

Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

matlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassUnknown7 days ago
355

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

magnus919/agent-skills111—~2.3kAutomated safety check: PassMITyesterday
356

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

magnus919/agent-skills111—~1.5kAutomated safety check: PassMITyesterday
357

Fully local multi-agent swarm intelligence simulation engine using Neo4j + Ollama for public opinion, market sentiment, and social dynamics prediction.

LeoYeAI/openclaw-master-skills2.2k—~4.1kAutomated safety check: NotesMIT2 mo ago
358

Full pipeline where an agent team collaborates to develop an LLM app.

revfactory/harness-1001.3k—~1.9kAutomated safety check: PassApache-2.06 mo ago
359
359.Ray

Production Ray (open-source, pinned to 2.57) for classic-ML workloads from training to serving — Ray Train, Tune, Data, Serve, Core, and cluster deployment on KubeRay.

pproenca/dot-skills214—~2.2kAutomated safety check: PassMIT1 mo ago
360

LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor).

pproenca/dot-skills214—~1.7kAutomated safety check: PassMIT1 mo ago
361

Expert AI/ML Engineer with deep MLOps expertise. An agent skill from theneoai/awesome-skills.

theneoai/awesome-skills183—~2.9kAutomated safety check: PassMIT4 mo ago
362

LLM serving expert: vLLM, TensorRT-LLM, Triton Inference Server, quantization (INT8/FP8/GPTQ/AWQ), continuous batching, PagedAttention, KV cache management.

theneoai/awesome-skills183—~3.4kAutomated safety check: PassMIT4 mo ago
363

MLflow expert: experiment tracking, model registry, autologging, MLflow Projects, MLflow Models, model serving, A/B testing, feature store integration.

theneoai/awesome-skills183—~3.5kAutomated safety check: PassMIT4 mo ago
364

Expert robot perception engineer specializing in 3D point cloud processing, multi-modal sensor fusion (camera+LiDAR+IMU), real-time SLAM, and edge-optimized deep learning inference via TensorRT/ONNX…

theneoai/awesome-skills183—~1.5kAutomated safety check: PassMIT4 mo ago