Topic · AI & LLM Engineering

Best LLM inference and serving skills, page 5

Skills #193–240 of 372, ranked by score.

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
193

Integrates local AI capabilities into applications using Embeddable Lemonade.

amd/skills406—~6kAutomated safety check: PassMITtoday
194

Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators.

lazyFrogLOL/Harness_Engineering128—~2.2kAutomated safety check: PassNo licence4 mo ago
195

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

sickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMITtoday
196

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMITtoday
197

Extract structured domain knowledge from AI models in-session or from local open-source models via Ollama.

sickn33/agentic-awesome-skills47k2 repos~926Automated safety check: PassMITtoday
198

Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.5kAutomated safety check: PassApache-2.0today
199

Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.

mirage-project/mirage2.5k—~4.6kAutomated safety check: PassApache-2.0yesterday
200

Measure local AI task latency, token usage, errors and verified outcomes using Pudu AI hardware evidence and installed Ollama models.

davila7/claude-code-templates32k—~1.3kAutomated safety check: PassMITtoday
201

Reduce the size of glTF/GLB 3D assets to cut Git LFS bandwidth/storage while keeping them loadable by the editor's Bevy 0.18 glTF loader.

elodin-sys/elodin547—~1kAutomated safety check: PassApache-2.0today
202

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom218—~7.2kAutomated safety check: NotesUnknowntoday
203
203.Ollama

Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.

Prism-Shadow/penguin-harness2.5k—~839Automated safety check: NotesApache-2.0yesterday
204
204.Vllm

Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.

Prism-Shadow/penguin-harness2.5k—~1kAutomated safety check: PassApache-2.0yesterday
205

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT4 days ago
206

Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard.

Xilinx/mlir-air150—~3.4kAutomated safety check: PassMITtoday
207

LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.

ancoleman/ai-design-components526—~3.4kAutomated safety check: PassMIT10 mo ago
208

Diagnoses Qdrant search quality issues. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~928Automated safety check: PassMITtoday
209

Optimize AgentDB performance with quantization (4-32x memory reduction), HNSW indexing (150x faster search), caching, and batch operations.

aiskillstore/marketplace4306 repos~3kAutomated safety check: PassNo licencetoday
210

Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra.

pytorch/test-infra113—~2.4kAutomated safety check: PassUnknowntoday
211

Add and manage evaluation results in Hugging Face model cards.

sickn33/agentic-awesome-skills47k2 repos~418Automated safety check: PassMITtoday
212
212.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.5k1 repo~1.8kAutomated safety check: PassApache-2.0today
213

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

sickn33/agentic-awesome-skills47k1 repo~2.5kAutomated safety check: PassMITtoday
214

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMITtoday
215

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

sickn33/agentic-awesome-skills47k1 repo~2.3kAutomated safety check: PassMITtoday
216

Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0today
217

vLLM: high-throughput LLM serving, OpenAI API, quantization.

Luciole-Studio/Misaka-Agent1582 repos~2.3kAutomated safety check: PassMITyesterday
218

Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph…

agentsope/SkillAlchemy466—~5.8kAutomated safety check: PassMITtoday
219

Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~336Automated safety check: PassMITtoday
220

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

wshobson/agents40k—~2kAutomated safety check: PassMIT4 days ago
221

Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer.

pytorch/test-infra113—~1.3kAutomated safety check: PassUnknowntoday
222

Command-line interface for Ollama - Local LLM inference and model management via Ollama REST API.

HKUDS/CLI-Anything52k—~1.2kAutomated safety check: PassApache-2.017 days ago
223

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.5k1 repo~2.9kAutomated safety check: PassApache-2.0today
224

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.5k1 repo~3.1kAutomated safety check: PassApache-2.0today
225

Google platform decision and setup guidance, loaded on demand from Google's skill catalog.

google/skills21k—~1.7kAutomated safety check: PassApache-2.0today
226
226.Gke InferenceOfficial

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

google/skills21k—~2kAutomated safety check: PassApache-2.0today
227

Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning.

amd/Quark181—~4.3kAutomated safety check: PassMIT11 days ago
228

Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow…

open-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.010 days ago
229

TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.

secondsky/claude-skills227—~3.6kAutomated safety check: NotesMIT11 days ago
230
230.ML AIOfficial

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

grafana/skills281—~1.3kAutomated safety check: PassApache-2.0today
231

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

ruvnet/ruflo74k—~336Automated safety check: NotesMITtoday
232

Validate ChatQnA Core REST APIs from docs/user-guide/api-reference.md using repeatable curl-based smoke tests, runtime-specific endpoint checks (OpenVINO or Ollama), and concise pass/fail evidence.

open-edge-platform/edge-ai-libraries169—~1.6kAutomated safety check: PassApache-2.0today
233

Build Chat Question and Answer Core Docker images from source using direct Docker or Docker Compose build commands (backend CPU, backend GPU, backend Ollama, and UI).

open-edge-platform/edge-ai-libraries169—~1.2kAutomated safety check: PassApache-2.0today
234

Deploy Chat Question-and-Answer Core with Docker Compose (OpenVINO CPU, OpenVINO GPU, or Ollama CPU), including env setup, profile selection, startup verification, health checks, and teardown.

open-edge-platform/edge-ai-libraries169—~2.9kAutomated safety check: PassApache-2.0today
235

Troubleshoot Chat Question-and-Answer Core end-to-end across Docker Compose and Helm deployments, including startup failures, health/API errors, runtime mismatches (OpenVINO vs Ollama), model/config…

open-edge-platform/edge-ai-libraries169—~2.6kAutomated safety check: PassApache-2.0today
236

High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1581 repo~1.3kAutomated safety check: PassMITyesterday
237

Search PubMed, preserve complete abstracts and sections, and produce research briefings with abstract openings or optional OpenAI/Ollama summaries.

ClawBio/ClawBio1.2k—~3.3kAutomated safety check: PassMITtoday
238
238.Airunway Aks SetupOfficial

Set up AI Runway on AKS — from bare cluster to running model.

microsoft/GitHub-Copilot-for-Azure2551 repo~1.1kAutomated safety check: PassMITtoday
239

Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation.

microsoft/GitHub-Copilot-for-Azure255—~3.2kAutomated safety check: PassMITtoday
240

EDM and club music as a style layer (电子舞曲风格层) - a form built on layer addition and subtraction rather than verses and choruses.

jtydhr88/music-composition-skills151—~2.1kAutomated safety check: PassMIT16 days ago