Topic · AI & LLM Engineering
Best LLM inference and serving skills, page 4
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | 145.Ollama Optimizer Optimize Ollama configuration for the current machine's hardware. | luongnv89/ | 131 | — | ~4.1k | Automated safety check: Notes | MIT | yesterday |
| 146 | Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. | amd/ | 398 | — | ~1.7k | Automated safety check: Notes | MIT | today |
| 147 | 147.Facturas Úsalo cuando el usuario pida redactar una factura, una cuenta de cobro o una nota de cobro. | gustavoeenriquez/ | 212 | — | ~127 | Automated safety check: Pass | MIT | 3 days ago |
| 148 | Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades. | MetaX-MACA/ | 179 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 149 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 911 | — | ~7.5k | Automated safety check: Pass | No licence | 3 days ago |
| 150 | 150.Vss Deploy Deploys and manages VSS through setup.sh and its Docker Compose overlays. | open-edge-platform/ | 169 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | today |
| 151 | Installation and configuration skill for Agent Brain document search system. | SpillwaveSolutions/ | 119 | — | ~7k | Automated safety check: Notes | MIT | 18 days ago |
| 152 | 152.ML Engineer Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. | davila7/ | 32k | 10 repos | ~2.3k | Automated safety check: Pass | MIT | today |
| 153 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 181 | — | ~1.4k | Automated safety check: Pass | MIT | 10 days ago |
| 154 | A skill your agent uses when a new Ollama Cloud model is announced or available (e.g. | heypinchy/ | 182 | — | ~3.9k | Automated safety check: Notes | AGPL-3.0 | 17 days ago |
| 155 | This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. | vllm-project/ | 103 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 156 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 398 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 157 | 157.Nemotron Ultra Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference. | NVIDIA-NeMo/ | 2.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 158 | A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. | BBuf/ | 911 | — | ~1.5k | Automated safety check: Pass | No licence | 3 days ago |
| 159 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 160 | 160.Open Notebook Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis. | majiayu000/ | 666 | 4 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 161 | DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity… | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 162 | 162.Llama Cpp llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes. | Tommy-yw/ | 546 | 5 repos | ~2.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 163 | Create or debug an AReno multi-turn agentic dataset, runagent implementation, tool schemas, tool execution, reward, loss masks, interactive TUI game with OpenAI-compatible LLM inference, or agentic… | inclusionAI/ | 323 | — | ~435 | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 164 | 164.Prime Agent A skill your agent uses when learning, configuring, or troubleshooting Prime Agent (PrimeIntellect-ai/prime-agent), including installation, providers, custom OpenAI-compatible models, local… | wcygan/ | 196 | — | ~1.8k | Automated safety check: Pass | No licence | today |
| 165 | 165.Local LLM Free Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. | artokun/ | 795 | — | ~897 | Automated safety check: Notes | MIT | 3 days ago |
| 166 | A skill your agent uses when running quantization of a BF16/FP16 GGUF repo and Skippy layer-package creation as one local or Hugging Face Jobs workflow, publishing both artifacts to Hugging Face. | Mesh-LLM/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 167 | Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware. | sickn33/ | 47k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 168 | 168.Open Notebook Organizes research with the self-hosted Open Notebook alternative to NotebookLM. | K-Dense-AI/ | 48k | 1 repo | ~2.8k | Automated safety check: Pass | MIT | 3 days ago |
| 169 | 3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment… | jaccen/ | 161 | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 170 | Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs. | MetaX-MACA/ | 179 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 171 | Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags. | waybarrios/ | 533 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 172 | L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script… | amd/ | 181 | — | ~3.4k | Automated safety check: Pass | MIT | 10 days ago |
| 173 | Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations. | oracle/ | 125 | — | ~1.8k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 174 | 174.App Opinionated app components building on top of ./ui primitives | JakeATX/ | 148 | — | ~146 | Automated safety check: Pass | MIT | yesterday |
| 175 | 175.Page Agent Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with… | Tommy-yw/ | 546 | 3 repos | ~2.3k | Automated safety check: Notes | MIT | 4 mo ago |
| 176 | Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and… | ai-dynamo/ | 8.2k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 177 | Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic. | vllm-project/ | 7.1k | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
| 178 | Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces. | darknecrocities/ | 239 | — | ~17k | Automated safety check: Pass | No licence | 2 days ago |
| 179 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 180 | 180.Pi Agent Builds with and operates Pi, the minimal terminal coding harness. | K-Dense-AI/ | 48k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 3 days ago |
| 181 | Select and verify the current region-specific serving container URI for a SageMaker model deployment. | waybarrios/ | 533 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 182 | 182.Quark Onnx Debug Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts. | amd/ | 181 | — | ~4.8k | Automated safety check: Pass | MIT | 10 days ago |
| 183 | 183.Crud Archive Run Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL… | open-thoughts/ | 301 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 184 | 184.Model Serving LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components. | ancoleman/ | 526 | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 10 mo ago |
| 185 | Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex. | databricks/ | 345 | — | ~3k | Automated safety check: Pass | Unknown | today |
| 186 | 186.Cc Ollama Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs. | mathruffian-dot/ | 254 | — | ~118 | Automated safety check: Pass | MIT | 1 mo ago |
| 187 | 187.Pai Skills Root Index of all PAI skills. | nirholas/ | 113 | — | ~333 | Automated safety check: Pass | Unknown | 23 days ago |
| 188 | Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp. | dpearson2699/ | 1.2k | — | ~3.4k | Automated safety check: Pass | Unknown | 2 mo ago |
| 189 | A skill your agent uses when changing mesh-llm's llama.cpp patch queue, upstream pin, prepare/build scripts, or carried RPC, MoE, and mesh-hook llama.cpp patches. | Mesh-LLM/ | 3.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 190 | A skill your agent uses when changing mesh-llm's patched llama.cpp Skippy ABI, runtime hooks, model introspection, tensor filtering, activation-frame execution, GGUF writer surface, upstream pin, or… | Mesh-LLM/ | 3.5k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 191 | Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 192 | 192.Ito Inference Inspect the availability of model serving on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed serving manifest. | affaan-m/ | 275k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | 3 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23