Search

LLM inference and serving

373 skills found, page 8.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
337

Run quantized LLMs locally with llama.cpp — CPU+GPU inference, GGUF format, OpenAI-compatible server, and Python bindings.

AlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT6 mo ago
338
338.Vllm

Deploy and serve LLMs with vLLM — OpenAI-compatible inference server with PagedAttention, continuous batching, and quantization support.

AlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT6 mo ago
339

Build model quantization tool operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~571Automated safety check: PassMITyesterday
340

Deploy KServe InferenceService on CoreWeave with autoscaling and GPU scheduling.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.2kAutomated safety check: PassMITyesterday
341

Set up local development workflow for CoreWeave GPU deployments.

jeremylongshore/tons-of-skills-marketplace2.8k—~927Automated safety check: NotesMITyesterday
342

Optimize CoreWeave GPU inference latency and throughput. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMITyesterday
343

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago
344

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2.2kAutomated safety check: PassMIT4 mo ago
345

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2.1kAutomated safety check: PassMIT4 mo ago
346

Papers on LLMs for IT operations and AIOps research. An agent skill from wentorai/research-plugins.

wentorai/research-plugins2981 repo~3.3kAutomated safety check: PassMIT3 mo ago
347

Route AI coding queries to local LLMs in air-gapped networks.

hoodini/ai-agents-skills282—~20kAutomated safety check: PassNo licence3 mo ago
348

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.

uw-syfi/vibesys105—~2.9kAutomated safety check: PassMITyesterday
349

Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring.

curiositech/some_claude_skills244—~3.4kAutomated safety check: PassMIT1 mo ago
350

AISBench Benchmark - AI model evaluation tool for Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday
351

自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licenceyesterday
352

Cloudflare Workers AI for serverless GPU inference. An agent skill from secondsky/claude-skills.

secondsky/claude-skills227—~2.4kAutomated safety check: PassMIT13 days ago
353

Use this sub-skill for Torch-TensorRT model compilation, dynamic input planning, torch.export workflows, save/load formats, raw TensorRT engines, and compile-time troubleshooting.

VectorSpaceLab/AREX-Skill331—~1.3kAutomated safety check: PassBSD-3-Clause1 mo ago
354

Use torchtune generation, Eleuther evaluation, and quantization workflows safely after checkpoints exist.

VectorSpaceLab/AREX-Skill331—~1kAutomated safety check: PassBSD-3-Clause1 mo ago
355

Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark…

VectorSpaceLab/AREX-Skill331—~1kAutomated safety check: PassBSD-3-Clause1 mo ago
356

Use this vLLM sub-skill for structured outputs, JSON/regex/grammar constraints, tool calling, reasoning parsers, chat-template/tool-parser routing, streaming tool-call deltas, and parser/backend…

VectorSpaceLab/AREX-Skill331—~823Automated safety check: PassApache-2.01 mo ago
357

A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging…

VectorSpaceLab/AREX-Skill331—~1.5kAutomated safety check: PassBSD-3-Clause1 mo ago
358

A skill your agent uses when planning or auditing TorchVision reference training/evaluation workflows for classification, quantization, detection, segmentation, video classification, optical flow…

VectorSpaceLab/AREX-Skill331—~684Automated safety check: PassBSD-3-Clause1 mo ago
359

昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…

ascend-ai-coding/awesome-ascend-skills174—~1.2kAutomated safety check: PassNo licenceyesterday
360

模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor…

ascend-ai-coding/awesome-ascend-skills174—~2.4kAutomated safety check: PassNo licenceyesterday
361

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

ascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNo licenceyesterday
362

Fetch current model names from AI providers (Anthropic, OpenAI, Gemini, Ollama), classify them into tiers (fast/default/heavy), and detect new models.

aiskillstore/marketplace433—~1.9kAutomated safety check: PassNo licenceyesterday
363

Select appropriate Ollama models for processing sensitive but legal content.

divinevideo/divine-mobile266—~1.3kAutomated safety check: PassMPL-2.0yesterday
364

A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

ericrisco/rsc-harness180—~4.1kAutomated safety check: PassMITyesterday
365

A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

ericrisco/rsc-harness180—~3.6kAutomated safety check: PassMITyesterday
366
366.Vllm

A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…

ericrisco/rsc-harness180—~3.6kAutomated safety check: PassMITyesterday
367

Manage background coding agents in tmux sessions. An agent skill from sundial-org/awesome-openclaw-skills.

sundial-org/awesome-openclaw-skills663—~933Automated safety check: PassNo licence7 mo ago
368

Delegates tasks to a locally served Muse Glimmer via ollama.

athola/claude-night-market341—~988Automated safety check: NotesMITyesterday
369

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

magnus919/agent-skills115—~2.3kAutomated safety check: PassMITyesterday
370

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

magnus919/agent-skills115—~1.5kAutomated safety check: PassMITyesterday
371

Offline experimental post-call feedback scorer using supplied ratings and text heuristics.

CALLE-AI/awesome-phone-call-agents107—~727Automated safety check: PassMITyesterday
372

Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

matlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassUnknown2 days ago
373

Fully local multi-agent swarm intelligence simulation engine using Neo4j + Ollama for public opinion, market sentiment, and social dynamics prediction.

LeoYeAI/openclaw-master-skills2.2k—~4.1kAutomated safety check: NotesMIT2 mo ago