Topic · AI & LLM Engineering
Best LLM inference and serving skills, page 5
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 193 | Integrates local AI capabilities into applications using Embeddable Lemonade. | amd/ | 406 | — | ~6k | Automated safety check: Pass | MIT | today |
| 194 | Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators. | lazyFrogLOL/ | 128 | — | ~2.2k | Automated safety check: Pass | No licence | 4 mo ago |
| 195 | 195.Local LLM Expert Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. | sickn33/ | 47k | 2 repos | ~1.6k | Automated safety check: Pass | MIT | today |
| 196 | 196.Vllm Server Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 197 | Extract structured domain knowledge from AI models in-session or from local open-source models via Ollama. | sickn33/ | 47k | 2 repos | ~926 | Automated safety check: Pass | MIT | today |
| 198 | Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure. | NVIDIA-AI-Blueprints/ | 1.9k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 199 | Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation. | mirage-project/ | 2.5k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 200 | Measure local AI task latency, token usage, errors and verified outcomes using Pudu AI hardware evidence and installed Ollama models. | davila7/ | 32k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 201 | Reduce the size of glTF/GLB 3D assets to cut Git LFS bandwidth/storage while keeping them loadable by the editor's Bevy 0.18 glTF loader. | elodin-sys/ | 547 | — | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 202 | 202.Hyperloom Setup Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom. | AMD-AGI/ | 218 | — | ~7.2k | Automated safety check: Notes | Unknown | today |
| 203 | 203.Ollama Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents. | Prism-Shadow/ | 2.5k | — | ~839 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 204 | 204.Vllm Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads. | Prism-Shadow/ | 2.5k | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 205 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 4 days ago |
| 206 | Phase 1 of LLM deployment — for every leaf kernel × shape the model needs, verify numerical correctness on real NPU2 against the registry's GPU/vLLM-aligned standard. | Xilinx/ | 150 | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 207 | 207.Model Serving LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components. | ancoleman/ | 526 | — | ~3.4k | Automated safety check: Pass | MIT | 10 mo ago |
| 208 | Diagnoses Qdrant search quality issues. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~928 | Automated safety check: Pass | MIT | today |
| 209 | Optimize AgentDB performance with quantization (4-32x memory reduction), HNSW indexing (150x faster search), caching, and batch operations. | aiskillstore/ | 430 | 6 repos | ~3k | Automated safety check: Pass | No licence | today |
| 210 | Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra. | pytorch/ | 113 | — | ~2.4k | Automated safety check: Pass | Unknown | today |
| 211 | Add and manage evaluation results in Hugging Face model cards. | sickn33/ | 47k | 2 repos | ~418 | Automated safety check: Pass | MIT | today |
| 212 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.5k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 213 | Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. | sickn33/ | 47k | 1 repo | ~2.5k | Automated safety check: Pass | MIT | today |
| 214 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | today |
| 215 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | today |
| 216 | Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters. | google/ | 21k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | today |
| 217 | vLLM: high-throughput LLM serving, OpenAI API, quantization. | Luciole-Studio/ | 158 | 2 repos | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 218 | Neutral, framework-agnostic decision tree for project kickoff: "which agent / RAG / LLM framework should I reach for?" Synthesizes the ecosystem sections of 7 landmark-project SOPs (LangGraph… | agentsope/ | 466 | — | ~5.8k | Automated safety check: Pass | MIT | today |
| 219 | Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~336 | Automated safety check: Pass | MIT | today |
| 220 | 220.Quantized Export Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 4 days ago |
| 221 | Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer. | pytorch/ | 113 | — | ~1.3k | Automated safety check: Pass | Unknown | today |
| 222 | Command-line interface for Ollama - Local LLM inference and model management via Ollama REST API. | HKUDS/ | 52k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 17 days ago |
| 223 | Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson. | NVIDIA/ | 3.5k | 1 repo | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 224 | Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output. | NVIDIA/ | 3.5k | 1 repo | ~3.1k | Automated safety check: Pass | Apache-2.0 | today |
| 225 | Google platform decision and setup guidance, loaded on demand from Google's skill catalog. | google/ | 21k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 226 | Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. | google/ | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 227 | Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning. | amd/ | 181 | — | ~4.3k | Automated safety check: Pass | MIT | 11 days ago |
| 228 | 228.Datagen Launch Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow… | open-thoughts/ | 301 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 229 | 229.Tanstack AI TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama. | secondsky/ | 227 | — | ~3.6k | Automated safety check: Notes | MIT | 11 days ago |
| 230 | Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN… | grafana/ | 281 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 231 | 231.Cost Local Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a… | ruvnet/ | 74k | — | ~336 | Automated safety check: Notes | MIT | today |
| 232 | Validate ChatQnA Core REST APIs from docs/user-guide/api-reference.md using repeatable curl-based smoke tests, runtime-specific endpoint checks (OpenVINO or Ollama), and concise pass/fail evidence. | open-edge-platform/ | 169 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 233 | 233.Chatqna Build Build Chat Question and Answer Core Docker images from source using direct Docker or Docker Compose build commands (backend CPU, backend GPU, backend Ollama, and UI). | open-edge-platform/ | 169 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 234 | Deploy Chat Question-and-Answer Core with Docker Compose (OpenVINO CPU, OpenVINO GPU, or Ollama CPU), including env setup, profile selection, startup verification, health checks, and teardown. | open-edge-platform/ | 169 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 235 | Troubleshoot Chat Question-and-Answer Core end-to-end across Docker Compose and Helm deployments, including startup failures, health/API errors, runtime mismatches (OpenVINO vs Ollama), model/config… | open-edge-platform/ | 169 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 236 | 236.Tensorrt LLM High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 158 | 1 repo | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 237 | Search PubMed, preserve complete abstracts and sections, and produce research briefings with abstract openings or optional OpenAI/Ollama summaries. | ClawBio/ | 1.2k | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 238 | Set up AI Runway on AKS — from bare cluster to running model. | microsoft/ | 255 | 1 repo | ~1.1k | Automated safety check: Pass | MIT | today |
| 239 | Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation. | microsoft/ | 255 | — | ~3.2k | Automated safety check: Pass | MIT | today |
| 240 | 240.Mc Style Edm EDM and club music as a style layer (电子舞曲风格层) - a form built on layer addition and subtraction rather than verses and choruses. | jtydhr88/ | 151 | — | ~2.1k | Automated safety check: Pass | MIT | 16 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents546
- Deep learning408
- Embeddings373
- Prompt engineering348
- Retrieval-augmented generation348
- LLM evaluation308
- Fine-tuning308
- Speech recognition and synthesis305
- Structured output and tool calling268
- Model routing and gateways257
- LLM cost and token optimization254
- LLM API integration250
- LLM observability242
- LLM guardrails217
- Computer vision195
- Model hubs and datasets179
- GPU and accelerator computing172
- Diffusion and image models165
- Natural language processing134
- Reinforcement learning66
- AI interpretability23