Topic · AI & LLM Engineering

Best LLM inference and serving skills, page 2

Skills #49–96 of 364, ranked by score.

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Always resolve Hugging Face models via model-shelf before any download.

alexziskind1/model-shelf130—~792Automated safety check: PassMIT1 mo ago
50

Maintain a persistent AI-tended research wiki. An agent skill from frankchu91/mindbase-llm-wiki.

frankchu91/mindbase-llm-wiki128—~1.1kAutomated safety check: PassMIT23 days ago
51

Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.

Orchestra-Research/AI-Research-SKILLs13k10 repos~4kAutomated safety check: PassMIT3 mo ago
52

Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

MetaX-MACA/vLLM-metax179—~2.6kAutomated safety check: PassApache-2.08 days ago
53

CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models).

MadAppGang/claudish1k—~9kAutomated safety check: PassNo licencetoday
54

Build the standalone ZenDNN native library (zendnnl) from source with the alternate compute backends OFF (no oneDNN, libxsmm, parlooper, fbgemm), keeping AOCL DLP, which is the GEMM backend zendnnl…

amd/ZenDNN158—~2kAutomated safety check: PassUnknownyesterday
55

Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence.

ThinkFlowLab/vllm-rlt138—~1.1kAutomated safety check: PassApache-2.07 days ago
56

Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

intel/auto-round1.6k—~2.3kAutomated safety check: PassApache-2.0yesterday
57

Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

guoqingbao/xinfer333—~3.8kAutomated safety check: PassMIT28 days ago
58

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

Orchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT3 mo ago
59

Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

zhongkaifu/TensorSharp553—~1kAutomated safety check: PassBSD-3-Clausetoday
60

用于在 cc-router 仓库新增一个 LLM provider(即在 src-tauri/providers/ 下添加 YAML 描述符并完成配套的同步改动)。当用户说「加 provider」「接入 XX 厂商」「新增订阅源」「provider YAML」「让 cc-router 支持 OpenRouter/Together/Groq/Ollama 之类」时必须触发本…

finch-xu/cc-router270—~1.6kAutomated safety check: PassMITyesterday
61

Regenerate the README UI screenshots for nyx-local-ai from the browser harness — build the webview, serve the harness scenes, capture and crop the images into docs/.

sthamann/nyx-local-ai133—~471Automated safety check: PassMIT3 mo ago
62

Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~4.5kAutomated safety check: PassNo licence2 days ago
63

Deploy BankMCP™ to a small server so it works in claude.ai and on the phone: Railway or Fly.io, volume, domain, setup page, connector.

noskillish/bankmcp276—~744Automated safety check: PassMIT9 days ago
64

TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.

secondsky/claude-skills2271 repo~3.6kAutomated safety check: NotesMIT9 days ago
65

Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

google/skills21k—~5.1kAutomated safety check: PassApache-2.0today
66

The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models.

oxbshw/watch-skill452—~509Automated safety check: NotesMIT22 days ago
67

Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

Aseiel/VideoHighlighter157—~839Automated safety check: PassAGPL-3.0today
68

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT3 mo ago
69

Cross-drive storage audit and cleanup. An agent skill from SethGammon/Citadel.

SethGammon/Citadel922—~2.2kAutomated safety check: PassMIT6 days ago
70

Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

guqiong96/Lvllm4641 repo~831Automated safety check: PassApache-2.015 days ago
71

AIPC, AI Porting Conversion. An agent skill from qualcomm/qai-appbuilder.

qualcomm/qai-appbuilder246—~5.7kAutomated safety check: NotesUnknowntoday
72

Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment.

NVIDIA-NeMo/Nemotron2.1k—~1.9kAutomated safety check: PassApache-2.0yesterday
73

Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs.

MetaX-MACA/vLLM-metax179—~1.8kAutomated safety check: PassApache-2.08 days ago
74

Interact with a Meshtastic LoRa mesh network through MESH-API — list nodes, read messages, send texts, and check connection status.

mr-tbot/mesh-api179—~1.8kAutomated safety check: PassGPL-3.02 mo ago
75

Add a new quantization data type to AutoRound (e.g., INT, FP8, MXFP, NVFP, GGUF variants).

intel/auto-round1.6k—~1.5kAutomated safety check: PassApache-2.0yesterday
76

Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning.

aws-samples/appmod-blueprints113—~5kAutomated safety check: PassMIT-0today
77

Test LLM models served by xinfer for correctness, output quality, and performance.

guoqingbao/xinfer333—~2.6kAutomated safety check: PassMIT28 days ago
78

Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

marin-community/marin3.9k—~1.1kAutomated safety check: PassApache-2.0today
79

Run a structured debate between agent CLIs (claude, codex, agy, grok) and the user's configured API or local models (OpenAI, Anthropic, Google, xAI, OpenRouter, Ollama and more) through the Quorum…

Detrol/quorum-cli119—~807Automated safety check: NotesUnknown10 days ago
80

Naming conventions for SGLang speculative decoding identifiers.

sgl-project/sglang37k2 repos~1.6kAutomated safety check: PassApache-2.0today
81

Validate changed quantization math, packed formats, and inference kernels against a Torch oracle; adjudicate near-tie mismatches before rejecting optimized arithmetic.

ModelCloud/GPTQModel1.3k—~1.4kAutomated safety check: PassUnknowntoday
82

A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages.

zhongkaifu/TensorSharp553—~2.3kAutomated safety check: WarnBSD-3-Clausetoday
83

Invoke ML models, run vector search, and connect to MCP servers from Databricks Apps.

databricks-solutions/databricks-apps-cookbook183—~1.7kAutomated safety check: PassUnknown3 days ago
84

Ship a new nyx-local-ai release end to end — bump versions consistently, run the quality gates (typecheck, smoke tests, package), install locally, tag and push so CI publishes the installer artifacts.

sthamann/nyx-local-ai133—~609Automated safety check: PassMIT3 mo ago
85

GGUF format and llama.cpp quantization for efficient CPU/GPU inference.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.6kAutomated safety check: PassMIT3 mo ago
86

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.5kAutomated safety check: PassMIT3 mo ago
87

Set up BankMCP™ on this machine step by step: check Node, register the local MCP server, walk through the Enable Banking application, finish the setup page and connect the first bank.

noskillish/bankmcp276—~805Automated safety check: PassMIT9 days ago
88

Guided workflow for adding a new model architecture to llama.cpp.

JakeATX/llamAmpere148—~4.1kAutomated safety check: PassMITtoday
89

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

vllm-project/vllm-skills103—~1.5kAutomated safety check: PassApache-2.06 mo ago
90

Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion.

SouthpawIN/turbofit106—~1.9kAutomated safety check: PassMIT3 days ago
91

A skill your agent uses when the user wants to consult an AI persona "clone" for advice, strategy, analysis, or domain expertise—especially in startup, VC, tech, growth, HR, or business contexts.

team-attention/openclone130—~712Automated safety check: PassUnknown1 mo ago
92
92.Aqua DeploymentOfficial

Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

oracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.01 mo ago
93

Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch).

amd/ZenDNN158—~5.2kAutomated safety check: PassUnknownyesterday
94

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

amd/skills395—~4kAutomated safety check: NotesMITtoday
95

Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.0yesterday
96

A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.4kAutomated safety check: PassApache-2.0today