Topic · AI & LLM Engineering

Best LLM inference and serving skills for Claude Code, Codex and other agents.

Skills that serve and speed up language models, including local and quantised inference.
skills
364
official
48

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

sgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0today
2

Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

huggingface/skills11k1 repo~4.6kAutomated safety check: PassApache-2.06 days ago
3

Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

amElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT16 days ago
4

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

Jeffallan/claude-skills12k1 repo~1.7kAutomated safety check: PassMIT4 days ago
5

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.06 days ago
6

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

vllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0today
7

Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

raullenchai/Rapid-MLX3.9k—~1.6kAutomated safety check: NotesUnknowntoday
8

Bump or upgrade the pinned versions of Helmor's bundled agent CLIs, SDKs, and supporting binaries — Claude Code + claude-agent-sdk (lockstep), Codex, Cursor SDK, OpenCode, Kimi, Pi, and gh / glab /…

dohooo/helmor1.3k—~2.1kAutomated safety check: PassApache-2.01 mo ago
9

Run the pre-release end-to-end sweep of every user-facing surface — the 33 samples under samples/ (booted from their packaged artifacts and driven in a real browser via chrome-devtools MCP), the…

Atmosphere/atmosphere3.8k—~4.2kAutomated safety check: PassApache-2.0yesterday
10

Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

R6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT2 mo ago
11

Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

guqiong96/Lvllm4642 repos~349Automated safety check: PassApache-2.015 days ago
12

Create a new skill in the current repository. An agent skill from Hyk260/PureChat.

Hyk260/PureChat5461 repo~823Automated safety check: PassMIT20 days ago
13

Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~2.3kAutomated safety check: PassNo licence2 days ago
14

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0today
15

Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

perminder-klair/subwave1.4k—~2.4kAutomated safety check: NotesMITtoday
16

Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

intel/auto-round1.6k—~2.8kAutomated safety check: PassApache-2.0yesterday
17

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4642 repos~1.5kAutomated safety check: PassApache-2.015 days ago
18

Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

huggingface/skills11k1 repo~6.9kAutomated safety check: PassApache-2.06 days ago
19
19.Debug InferenceOfficial

Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

NVIDIA/OpenShell15k—~1.9kAutomated safety check: PassApache-2.0today
20

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~2.5kAutomated safety check: PassNo licence2 days ago
21

Routes a NanoClaw agent group to a local Ollama model instead of the Anthropic API, using environment overrides and an optional block on Anthropic hosts.

nanocoai/nanoclaw31k1 repo~1.5kAutomated safety check: PassMITyesterday
22

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

vllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0today
23

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMITtoday
24

A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

timescale/pg-aiguide1.9k1 repo~3.8kAutomated safety check: PassApache-2.05 days ago
25

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.09 days ago
26

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
27
27.Adapt New LLMOfficial

Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

intel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0yesterday
28

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
29

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNo licence2 days ago
30

Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

danyuchn/pii-guard249—~2.7kAutomated safety check: PassMIT5 days ago
31

Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands.

internet-court/internet-court-skill6.4k1 repo~1.9kAutomated safety check: PassUnknown1 mo ago
32

Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…

amd/skills395—~2.6kAutomated safety check: PassMITtoday
33

A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…

warpfront/hipfire652—~2.6kAutomated safety check: PassUnknowntoday
34

Use iDeer as a daily paper-reading workflow for chatbot-first users such as Codex, Gemini, or ChatGPT.

AI45Lab/iDeer416—~3kAutomated safety check: NotesAGPL-3.02 mo ago
35

Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

ModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassUnknowntoday
36

Run wally's LLM e2e on Apple Neural Engine (NeuRT) and Snapdragon Hexagon NPU (QHexRT) devices.

RunanywhereAI/wally1.6k—~845Automated safety check: PassMITtoday
37

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

intel/auto-round1.6k—~1.9kAutomated safety check: PassApache-2.0yesterday
38

Adapt and port new LLM model architectures to this xinfer project.

guoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT28 days ago
39

Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner).

tryonlabs/opentryon551—~1.1kAutomated safety check: PassUnknownyesterday
40

Check if package.json files are in sync with pnpm-lock.yaml.

Hyk260/PureChat546—~594Automated safety check: PassMIT20 days ago
41

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

huggingface/skills11k3 repos~945Automated safety check: PassApache-2.06 days ago
42

Query images with a local Ollama vision model without loading the image into the main agent context.

gridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0yesterday
43

Build, extend, and debug AI agents in Egroo using the Microsoft Agent Framework (C .NET).

jihadkhawaja/Egroo178—~1.9kAutomated safety check: PassApache-2.06 mo ago
44

Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu.

apache/dubbo-go-pixiu568—~2.5kAutomated safety check: PassApache-2.04 days ago
45
45.Bank

Answer questions about the user's own bank accounts and money using the bank MCP tools (listaccounts, getbalances, gettransactions, watches).

noskillish/bankmcp276—~1.1kAutomated safety check: PassMIT8 days ago
46

Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

EGalahad/sim2real145—~1.1kAutomated safety check: PassNo licence9 days ago
47

A skill your agent uses when a user asks how many tokens (or how much context/cost) ContextPilot has saved, or wants a ContextPilot savings status/summary inside Hermes Agent — e.g.

EfficientContext/ContextPilot140—~1.4kAutomated safety check: PassMIT4 days ago
48

Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

dstackai/dstack2.3k—~403Automated safety check: PassMPL-2.0today

Questions, answered from the data.

What is the best LLM inference and serving skill?

LLM Torch Profiler Analysis from sgl-project/sglang ranks first of the 364 LLM inference and serving skills listed here, with the highest score: its repository has 37k GitHub stars, 2 other GitHub owners carry a copy, its SKILL.md loads about 6.4k tokens and it passes the automated safety check with no findings. Next come SageMaker Serving Image Selection and Aider Delegate.

Which LLM inference and serving skills are official?

48 of the 364 LLM inference and serving skills are official, published by the vendor's own GitHub organization: SageMaker Serving Image Selection, Hugging Face Local Model Evals, Adapt New Diffusion Model, SageMaker Production Defaults, Debug Inference and 43 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.