Topic · AI & LLM Engineering
Best LLM inference and serving skills, page 2
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Resolve Always resolve Hugging Face models via model-shelf before any download. | alexziskind1/ | 130 | — | ~792 | Automated safety check: Pass | MIT | 1 mo ago |
| 50 | 50.Mindbase Maintain a persistent AI-tended research wiki. An agent skill from frankchu91/mindbase-llm-wiki. | frankchu91/ | 128 | — | ~1.1k | Automated safety check: Pass | MIT | 23 days ago |
| 51 | Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models. | Orchestra-Research/ | 13k | 10 repos | ~4k | Automated safety check: Pass | MIT | 3 mo ago |
| 52 | Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels. | MetaX-MACA/ | 179 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 53 | CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models). | MadAppGang/ | 1k | — | ~9k | Automated safety check: Pass | No licence | today |
| 54 | 54.Build Zendnn Build the standalone ZenDNN native library (zendnnl) from source with the alternate compute backends OFF (no oneDNN, libxsmm, parlooper, fbgemm), keeping AOCL DLP, which is the GEMM backend zendnnl… | amd/ | 158 | — | ~2k | Automated safety check: Pass | Unknown | yesterday |
| 55 | Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence. | ThinkFlowLab/ | 138 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 56 | Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK). | intel/ | 1.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 57 | 57.Check Model Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer. | guoqingbao/ | 333 | — | ~3.8k | Automated safety check: Pass | MIT | 28 days ago |
| 58 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 59 | 59.Market Data Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares). | zhongkaifu/ | 553 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | today |
| 60 | 60.New Provider 用于在 cc-router 仓库新增一个 LLM provider(即在 src-tauri/providers/ 下添加 YAML 描述符并完成配套的同步改动)。当用户说「加 provider」「接入 XX 厂商」「新增订阅源」「provider YAML」「让 cc-router 支持 OpenRouter/Together/Groq/Ollama 之类」时必须触发本… | finch-xu/ | 270 | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 61 | Regenerate the README UI screenshots for nyx-local-ai from the browser harness — build the webview, serve the harness scenes, capture and crop the images into docs/. | sthamann/ | 133 | — | ~471 | Automated safety check: Pass | MIT | 3 mo ago |
| 62 | Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks. | BBuf/ | 900 | — | ~4.5k | Automated safety check: Pass | No licence | 2 days ago |
| 63 | 63.Deploy Deploy BankMCP™ to a small server so it works in claude.ai and on the phone: Railway or Fly.io, volume, domain, setup page, connector. | noskillish/ | 276 | — | ~744 | Automated safety check: Pass | MIT | 9 days ago |
| 64 | 64.Tanstack AI TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama. | secondsky/ | 227 | 1 repo | ~3.6k | Automated safety check: Notes | MIT | 9 days ago |
| 65 | Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change. | google/ | 21k | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | today |
| 66 | The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models. | oxbshw/ | 452 | — | ~509 | Automated safety check: Notes | MIT | 22 days ago |
| 67 | Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better. | Aseiel/ | 157 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | today |
| 68 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 69 | 69.Houseclean Cross-drive storage audit and cleanup. An agent skill from SethGammon/Citadel. | SethGammon/ | 922 | — | ~2.2k | Automated safety check: Pass | MIT | 6 days ago |
| 70 | Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. | guqiong96/ | 464 | 1 repo | ~831 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 71 | 71.Aipc Toolkit AIPC, AI Porting Conversion. An agent skill from qualcomm/qai-appbuilder. | qualcomm/ | 246 | — | ~5.7k | Automated safety check: Notes | Unknown | today |
| 72 | Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. | NVIDIA-NeMo/ | 2.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 73 | Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs. | MetaX-MACA/ | 179 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 74 | 74.Mesh API Interact with a Meshtastic LoRa mesh network through MESH-API — list nodes, read messages, send texts, and check connection status. | mr-tbot/ | 179 | — | ~1.8k | Automated safety check: Pass | GPL-3.0 | 2 mo ago |
| 75 | Add a new quantization data type to AutoRound (e.g., INT, FP8, MXFP, NVFP, GGUF variants). | intel/ | 1.6k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 76 | Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning. | aws-samples/ | 113 | — | ~5k | Automated safety check: Pass | MIT-0 | today |
| 77 | 77.Test Model Test LLM models served by xinfer for correctness, output quality, and performance. | guoqingbao/ | 333 | — | ~2.6k | Automated safety check: Pass | MIT | 28 days ago |
| 78 | Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding. | marin-community/ | 3.9k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 79 | 79.Quorum Run a structured debate between agent CLIs (claude, codex, agy, grok) and the user's configured API or local models (OpenAI, Anthropic, Google, xAI, OpenRouter, Ollama and more) through the Quorum… | Detrol/ | 119 | — | ~807 | Automated safety check: Notes | Unknown | 10 days ago |
| 80 | Naming conventions for SGLang speculative decoding identifiers. | sgl-project/ | 37k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 81 | Validate changed quantization math, packed formats, and inference kernels against a Torch oracle; adjudicate near-tie mismatches before rejecting optimized arithmetic. | ModelCloud/ | 1.3k | — | ~1.4k | Automated safety check: Pass | Unknown | today |
| 82 | 82.Research A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages. | zhongkaifu/ | 553 | — | ~2.3k | Automated safety check: Warn | BSD-3-Clause | today |
| 83 | Invoke ML models, run vector search, and connect to MCP servers from Databricks Apps. | databricks-solutions/ | 183 | — | ~1.7k | Automated safety check: Pass | Unknown | 3 days ago |
| 84 | 84.Release Nyx Ship a new nyx-local-ai release end to end — bump versions consistently, run the quality gates (typecheck, smoke tests, package), install locally, tag and push so CI publishes the installer artifacts. | sthamann/ | 133 | — | ~609 | Automated safety check: Pass | MIT | 3 mo ago |
| 85 | GGUF format and llama.cpp quantization for efficient CPU/GPU inference. | Orchestra-Research/ | 13k | 4 repos | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 86 | 86.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 87 | 87.Setup Set up BankMCP™ on this machine step by step: check Node, register the local MCP server, walk through the Enable Banking application, finish the setup page and connect the first bank. | noskillish/ | 276 | — | ~805 | Automated safety check: Pass | MIT | 9 days ago |
| 88 | Guided workflow for adding a new model architecture to llama.cpp. | JakeATX/ | 148 | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 89 | Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. | vllm-project/ | 103 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 90 | 90.Turbofit Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion. | SouthpawIN/ | 106 | — | ~1.9k | Automated safety check: Pass | MIT | 3 days ago |
| 91 | A skill your agent uses when the user wants to consult an AI persona "clone" for advice, strategy, analysis, or domain expertise—especially in startup, VC, tech, growth, HR, or business contexts. | team-attention/ | 130 | — | ~712 | Automated safety check: Pass | Unknown | 1 mo ago |
| 92 | Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling. | oracle/ | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 93 | Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch). | amd/ | 158 | — | ~5.2k | Automated safety check: Pass | Unknown | yesterday |
| 94 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 395 | — | ~4k | Automated safety check: Notes | MIT | today |
| 95 | Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes. | NVIDIA-NeMo/ | 2.1k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 96 | A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23