Topic · AI & LLM Engineering

Best LLM inference and serving skills, page 3

Skills #97–144 of 372, ranked by score.

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97
97.Aqua DeploymentOfficial

Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

oracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.01 mo ago
98

Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch).

amd/ZenDNN158—~5.2kAutomated safety check: PassUnknownyesterday
99

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

amd/skills398—~4kAutomated safety check: NotesMITtoday
100

Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.0yesterday
101

A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.4kAutomated safety check: PassApache-2.0today
102
102.Add Vlm ModelOfficial

Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

intel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0today
103

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.1kAutomated safety check: PassMIT3 mo ago
104
104.Gptq

Post-training 4-bit quantization for LLMs with minimal accuracy loss.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT3 mo ago
105

Half-Quadratic Quantization for LLMs without calibration data.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT3 mo ago
106

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.2kAutomated safety check: PassMIT3 mo ago
107

Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace.

Orchestra-Research/AI-Research-SKILLs13k3 repos~3.7kAutomated safety check: PassMIT3 mo ago
108

High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.1kAutomated safety check: NotesMIT3 mo ago
109

Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT3 mo ago
110

Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT3 mo ago
111

Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits.

Orchestra-Research/AI-Research-SKILLs13k3 repos~3.5kAutomated safety check: PassMIT3 mo ago
112

Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

agentsope/SkillAlchemy459—~3kAutomated safety check: PassMIT1 mo ago
113

A template skill demonstrating the PAI SKILL.md format. An agent skill from nirholas/PAI.

nirholas/PAI113—~919Automated safety check: PassUnknown23 days ago
114

QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

qualcomm/qai-appbuilder246—~4.1kAutomated safety check: PassBSD-3-Clausetoday
115

Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills.

databricks/databricks-agent-skills3451 repo~3.3kAutomated safety check: PassUnknowntoday
116

Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes.

vllm-project/vllm-omni7.1k—~2.7kAutomated safety check: PassApache-2.0today
117

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

sohu-mptc/FlashRec107—~915Automated safety check: PassApache-2.06 days ago
118

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

vllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.06 mo ago
119

Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

scragnog/HOT-Step-CPP171—~4.9kAutomated safety check: PassMITtoday
120

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

amd/Quark181—~3kAutomated safety check: PassMIT10 days ago
121

Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

oracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.01 mo ago
122

Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs.

warpfront/hipfire653—~1.5kAutomated safety check: PassUnknowntoday
123

Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron.

Netis/heron102—~983Automated safety check: PassApache-2.03 days ago
124
124.Review PROfficial

Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files…

intel/auto-round1.6k—~1.7kAutomated safety check: PassApache-2.0today
125

Summarize a document at a requested length. An agent skill from nirholas/PAI.

nirholas/PAI113—~992Automated safety check: PassUnknown23 days ago
126

Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

MetaX-MACA/vLLM-metax179—~2.3kAutomated safety check: PassApache-2.08 days ago
127

This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms.

GoogleCloudPlatform/accelerated-platforms106—~1.2kAutomated safety check: PassApache-2.02 days ago
128

Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting.

Orchestra-Research/AI-Research-SKILLs13k3 repos~1.8kAutomated safety check: PassMIT3 mo ago
129

Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

azrtydxb/Fastllm-proxy108—~980Automated safety check: PassApache-2.02 days ago
130

Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.

vllm-project/vllm-omni7.1k—~5.5kAutomated safety check: PassApache-2.0today
131

Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

JakeATX/llamAmpere148—~5.2kAutomated safety check: PassMITyesterday
132

Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

amd/skills398—~5.7kAutomated safety check: NotesMITtoday
133

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

vllm-project/vllm-skills103—~2kAutomated safety check: PassApache-2.06 mo ago
134

Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

Krusty84/triton-ascend-agent-dev-kit106—~597Automated safety check: PassApache-2.01 mo ago
135

Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron.

Netis/heron102—~1.4kAutomated safety check: PassApache-2.03 days ago
136

Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

joshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesUnknown19 days ago
137

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.3kAutomated safety check: PassMIT3 mo ago
138

Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

MetaX-MACA/vLLM-metax179—~2.8kAutomated safety check: PassApache-2.08 days ago
139

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

BBuf/AI-Infra-Auto-Driven-SKILLS911—~3.9kAutomated safety check: PassNo licence2 days ago
140

Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4).

vllm-project/vllm-omni7.1k—~17kAutomated safety check: PassApache-2.0today
141

Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.

huggingface/skills11k1 repo~2.1kAutomated safety check: PassApache-2.06 days ago
142

Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

amd/Quark181—~2.9kAutomated safety check: PassMIT10 days ago
143

Adds an MCP server so the NanoClaw container agent can send prompts to local Ollama models, with optional tools to manage the model library.

nanocoai/nanoclaw31k1 repo~3kAutomated safety check: NotesMITyesterday
144

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

vllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.06 mo ago