Topic · AI & LLM Engineering

Best LLM inference and serving skills, page 4

Skills #145–192 of 372, ranked by score.

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
145

Optimize Ollama configuration for the current machine's hardware.

luongnv89/skills131—~4.1kAutomated safety check: NotesMITyesterday
146

Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

amd/skills398—~1.7kAutomated safety check: NotesMITtoday
147

Úsalo cuando el usuario pida redactar una factura, una cuenta de cobro o una nota de cobro.

gustavoeenriquez/MakerAi212—~127Automated safety check: PassMIT3 days ago
148

Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades.

MetaX-MACA/vLLM-metax179—~1.1kAutomated safety check: PassApache-2.09 days ago
149

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS911—~7.5kAutomated safety check: PassNo licence3 days ago
150

Deploys and manages VSS through setup.sh and its Docker Compose overlays.

open-edge-platform/edge-ai-libraries169—~4.1kAutomated safety check: PassApache-2.0today
151

Installation and configuration skill for Agent Brain document search system.

SpillwaveSolutions/agent-brain119—~7kAutomated safety check: NotesMIT18 days ago
152

Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks.

davila7/claude-code-templates32k10 repos~2.3kAutomated safety check: PassMITtoday
153

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark181—~1.4kAutomated safety check: PassMIT10 days ago
154

A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.

heypinchy/pinchy182—~3.9kAutomated safety check: NotesAGPL-3.017 days ago
155

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

vllm-project/vllm-skills103—~1.4kAutomated safety check: PassApache-2.06 mo ago
156

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills398—~2.3kAutomated safety check: PassMITtoday
157

Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.

NVIDIA-NeMo/Nemotron2.1k—~1.8kAutomated safety check: PassApache-2.02 days ago
158

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS911—~1.5kAutomated safety check: PassNo licence3 days ago
159
159.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
160

Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.

majiayu000/claude-skill-registry6664 repos~2.4kAutomated safety check: PassMITtoday
161

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

open-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.010 days ago
162

llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes.

Tommy-yw/RunbookHermes5465 repos~2.2kAutomated safety check: PassMIT4 mo ago
163

Create or debug an AReno multi-turn agentic dataset, runagent implementation, tool schemas, tool execution, reward, loss masks, interactive TUI game with OpenAI-compatible LLM inference, or agentic…

inclusionAI/AReno323—~435Automated safety check: PassApache-2.014 days ago
164

A skill your agent uses when learning, configuring, or troubleshooting Prime Agent (PrimeIntellect-ai/prime-agent), including installation, providers, custom OpenAI-compatible models, local…

wcygan/dotfiles196—~1.8kAutomated safety check: PassNo licencetoday
165

Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama.

artokun/comfyui-mcp795—~897Automated safety check: NotesMIT3 days ago
166

A skill your agent uses when running quantization of a BF16/FP16 GGUF repo and Skippy layer-package creation as one local or Hugging Face Jobs workflow, publishing both artifacts to Hugging Face.

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
167

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

sickn33/agentic-awesome-skills47k1 repo~1.9kAutomated safety check: PassApache-2.0yesterday
168

Organizes research with the self-hosted Open Notebook alternative to NotebookLM.

K-Dense-AI/scientific-agent-skills48k1 repo~2.8kAutomated safety check: PassMIT3 days ago
169

3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment…

jaccen/Awesome-Gaussian-Skills161—~5.5kAutomated safety check: PassApache-2.0today
170

Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs.

MetaX-MACA/vLLM-metax179—~2.4kAutomated safety check: PassApache-2.09 days ago
171

Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.

waybarrios/opencode-power-pack533—~4.6kAutomated safety check: PassApache-2.02 days ago
172

L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

amd/Quark181—~3.4kAutomated safety check: PassMIT10 days ago
173

Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations.

oracle/accelerated-data-science125—~1.8kAutomated safety check: PassUPL-1.01 mo ago
174
174.App

Opinionated app components building on top of ./ui primitives

JakeATX/llamAmpere148—~146Automated safety check: PassMITyesterday
175

Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

Tommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: NotesMIT4 mo ago
176

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

ai-dynamo/dynamo8.2k—~4.9kAutomated safety check: PassApache-2.0today
177

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

vllm-project/vllm-omni7.1k—~3kAutomated safety check: PassApache-2.0today
178

Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.

darknecrocities/DomoDomo---All-in-one-Tool239—~17kAutomated safety check: PassNo licence2 days ago
179

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.06 mo ago
180

Builds with and operates Pi, the minimal terminal coding harness.

K-Dense-AI/scientific-agent-skills48k1 repo~2.1kAutomated safety check: PassMIT3 days ago
181

Select and verify the current region-specific serving container URI for a SageMaker model deployment.

waybarrios/opencode-power-pack533—~4.3kAutomated safety check: PassApache-2.02 days ago
182

Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts.

amd/Quark181—~4.8kAutomated safety check: PassMIT10 days ago
183

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

open-thoughts/OpenThoughts-Agent301—~1.2kAutomated safety check: PassApache-2.010 days ago
184

LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.

ancoleman/ai-design-components5261 repo~3.4kAutomated safety check: PassMIT10 mo ago
185

Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex.

databricks/databricks-agent-skills345—~3kAutomated safety check: PassUnknowntoday
186

Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.

mathruffian-dot/claude-code-lazy-packs254—~118Automated safety check: PassMIT1 mo ago
187

Index of all PAI skills.

nirholas/PAI113—~333Automated safety check: PassUnknown23 days ago
188

Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp.

dpearson2699/swift-ios-skills1.2k—~3.4kAutomated safety check: PassUnknown2 mo ago
189

A skill your agent uses when changing mesh-llm's llama.cpp patch queue, upstream pin, prepare/build scripts, or carried RPC, MoE, and mesh-hook llama.cpp patches.

Mesh-LLM/mesh-llm3.5k—~1.9kAutomated safety check: PassApache-2.0today
190

A skill your agent uses when changing mesh-llm's patched llama.cpp Skippy ABI, runtime hooks, model introspection, tensor filtering, activation-frame execution, GGUF writer surface, upstream pin, or…

Mesh-LLM/mesh-llm3.5k—~2.6kAutomated safety check: PassApache-2.0today
191

Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows.

maziyarpanahi/openmed5.5k—~2kAutomated safety check: PassApache-2.0today
192

Inspect the availability of model serving on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed serving manifest.

affaan-m/ECC275k1 repo~1.5kAutomated safety check: PassMIT3 days ago