Topic · AI & LLM Engineering
Best LLM inference and serving skills, page 3
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling. | oracle/ | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 98 | Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch). | amd/ | 158 | — | ~5.2k | Automated safety check: Pass | Unknown | yesterday |
| 99 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 398 | — | ~4k | Automated safety check: Notes | MIT | today |
| 100 | 100.Nemotron Super3 Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes. | NVIDIA-NeMo/ | 2.1k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 101 | A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 102 | Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling. | intel/ | 1.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 103 | 103.Awq Quantization Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. | Orchestra-Research/ | 13k | 3 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 104 | 104.Gptq Post-training 4-bit quantization for LLMs with minimal accuracy loss. | Orchestra-Research/ | 13k | 3 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 105 | 105.Hqq Quantization Half-Quadratic Quantization for LLMs without calibration data. | Orchestra-Research/ | 13k | 3 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 106 | Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. | Orchestra-Research/ | 13k | 3 repos | ~2.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 107 | 107.Moe Training Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. | Orchestra-Research/ | 13k | 3 repos | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 108 | High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 3 repos | ~2.1k | Automated safety check: Notes | MIT | 3 mo ago |
| 109 | Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model. | Orchestra-Research/ | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 110 | Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads. | Orchestra-Research/ | 13k | 3 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 111 | Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits. | Orchestra-Research/ | 13k | 3 repos | ~3.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 112 | Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor… | agentsope/ | 459 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 113 | 113.Example Tool A template skill demonstrating the PAI SKILL.md format. An agent skill from nirholas/PAI. | nirholas/ | 113 | — | ~919 | Automated safety check: Pass | Unknown | 23 days ago |
| 114 | 114.Model Builder QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder. | qualcomm/ | 246 | — | ~4.1k | Automated safety check: Pass | BSD-3-Clause | today |
| 115 | Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills. | databricks/ | 345 | 1 repo | ~3.3k | Automated safety check: Pass | Unknown | today |
| 116 | Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes. | vllm-project/ | 7.1k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 117 | 117.Recif Eval Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. | sohu-mptc/ | 107 | — | ~915 | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 118 | 118.Vllm Bench Serve Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. | vllm-project/ | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 119 | Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed. | scragnog/ | 171 | — | ~4.9k | Automated safety check: Pass | MIT | today |
| 120 | Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. | amd/ | 181 | — | ~3k | Automated safety check: Pass | MIT | 10 days ago |
| 121 | Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK. | oracle/ | 125 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 122 | 122.Hipfire Tester Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs. | warpfront/ | 653 | — | ~1.5k | Automated safety check: Pass | Unknown | today |
| 123 | 123.Dev Bump Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron. | Netis/ | 102 | — | ~983 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 124 | Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files… | intel/ | 1.6k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 125 | Summarize a document at a requested length. An agent skill from nirholas/PAI. | nirholas/ | 113 | — | ~992 | Automated safety check: Pass | Unknown | 23 days ago |
| 126 | Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs. | MetaX-MACA/ | 179 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 127 | This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms. | GoogleCloudPlatform/ | 106 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 128 | Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting. | Orchestra-Research/ | 13k | 3 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 129 | 129.Fastllm Backends Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding… | azrtydxb/ | 108 | — | ~980 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 130 | Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. | vllm-project/ | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 131 | 131.Code Review Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. | JakeATX/ | 148 | — | ~5.2k | Automated safety check: Pass | MIT | yesterday |
| 132 | Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda. | amd/ | 398 | — | ~5.7k | Automated safety check: Notes | MIT | today |
| 133 | 133.Vllm Deploy K8s Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. | vllm-project/ | 103 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 134 | Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering. | Krusty84/ | 106 | — | ~597 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 135 | Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron. | Netis/ | 102 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 136 | 136.Scholar RAG Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review. | joshzyj/ | 168 | — | ~7.4k | Automated safety check: Notes | Unknown | 19 days ago |
| 137 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 3 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 138 | Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision. | MetaX-MACA/ | 179 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 139 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 911 | — | ~3.9k | Automated safety check: Pass | No licence | 2 days ago |
| 140 | 140.Vllm Omni Test Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4). | vllm-project/ | 7.1k | — | ~17k | Automated safety check: Pass | Apache-2.0 | today |
| 141 | Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills. | huggingface/ | 11k | 1 repo | ~2.1k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 142 | Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers. | amd/ | 181 | — | ~2.9k | Automated safety check: Pass | MIT | 10 days ago |
| 143 | Adds an MCP server so the NanoClaw container agent can send prompts to local Ollama models, with optional tools to manage the model library. | nanocoai/ | 31k | 1 repo | ~3k | Automated safety check: Notes | MIT | yesterday |
| 144 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
Explore related skills
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23