Topic · AI & LLM Engineering
Best LLM inference and serving skills for Claude Code, Codex and other agents.
- skills
- 364
- official
- 48
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. | huggingface/ | 11k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 3 | Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself. | amElnagdy/ | 2.3k | 2 repos | ~3k | Automated safety check: Pass | MIT | 16 days ago |
| 4 | Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment. | Jeffallan/ | 12k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 4 days ago |
| 5 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 6 | Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements. | vllm-project/ | 2.9k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | 7.Perfup Autonomous performance optimization: research, PoC, benchmark, implement, review, PR | raullenchai/ | 3.9k | — | ~1.6k | Automated safety check: Notes | Unknown | today |
| 8 | Bump or upgrade the pinned versions of Helmor's bundled agent CLIs, SDKs, and supporting binaries — Claude Code + claude-agent-sdk (lockstep), Codex, Cursor SDK, OpenCode, Kimi, Pi, and gh / glab /… | dohooo/ | 1.3k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 9 | Run the pre-release end-to-end sweep of every user-facing surface — the 33 samples under samples/ (booted from their packaged artifacts and driven in a real browser via chrome-devtools MCP), the… | Atmosphere/ | 3.8k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. | R6410418/ | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 11 | Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm. | guqiong96/ | 464 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 12 | 12.Create Skill Create a new skill in the current repository. An agent skill from Hyk260/PureChat. | Hyk260/ | 546 | 1 repo | ~823 | Automated safety check: Pass | MIT | 20 days ago |
| 13 | Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence. | BBuf/ | 900 | — | ~2.3k | Automated safety check: Pass | No licence | 2 days ago |
| 14 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | today |
| 15 | Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using… | perminder-klair/ | 1.4k | — | ~2.4k | Automated safety check: Notes | MIT | today |
| 16 | Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT). | intel/ | 1.6k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 17 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 464 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 18 | Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups. | huggingface/ | 11k | 1 repo | ~6.9k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 19 | Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. | NVIDIA/ | 15k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 900 | — | ~2.5k | Automated safety check: Pass | No licence | 2 days ago |
| 21 | Routes a NanoClaw agent group to a local Ollama model instead of the Anthropic API, using environment overrides and an optional block on Anthropic hosts. | nanocoai/ | 31k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 22 | Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. | vllm-project/ | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 23 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | today |
| 24 | A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search. | timescale/ | 1.9k | 1 repo | ~3.8k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 25 | 25.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 26 | A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing… | Mesh-LLM/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box. | intel/ | 1.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 28 | 28.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 29 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 900 | — | ~2.8k | Automated safety check: Pass | No licence | 2 days ago |
| 30 | Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents. | danyuchn/ | 249 | — | ~2.7k | Automated safety check: Pass | MIT | 5 days ago |
| 31 | Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands. | internet-court/ | 6.4k | 1 repo | ~1.9k | Automated safety check: Pass | Unknown | 1 mo ago |
| 32 | Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the… | amd/ | 395 | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 33 | 33.Astrea A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform… | warpfront/ | 652 | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 34 | Use iDeer as a daily paper-reading workflow for chatbot-first users such as Codex, Gemini, or ChatGPT. | AI45Lab/ | 416 | — | ~3k | Automated safety check: Notes | AGPL-3.0 | 2 mo ago |
| 35 | Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems. | ModelCloud/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Unknown | today |
| 36 | Run wally's LLM e2e on Apple Neural Engine (NeuRT) and Snapdragon Hexagon NPU (QHexRT) devices. | RunanywhereAI/ | 1.6k | — | ~845 | Automated safety check: Pass | MIT | today |
| 37 | Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor). | intel/ | 1.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 38 | 38.Add Model Adapt and port new LLM model architectures to this xinfer project. | guoqingbao/ | 333 | — | ~4.2k | Automated safety check: Notes | MIT | 28 days ago |
| 39 | Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner). | tryonlabs/ | 551 | — | ~1.1k | Automated safety check: Pass | Unknown | yesterday |
| 40 | Check if package.json files are in sync with pnpm-lock.yaml. | Hyk260/ | 546 | — | ~594 | Automated safety check: Pass | MIT | 20 days ago |
| 41 | Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. | huggingface/ | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 42 | 42.Vision Query images with a local Ollama vision model without loading the image into the main agent context. | gridaco/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 43 | Build, extend, and debug AI agents in Egroo using the Microsoft Agent Framework (C .NET). | jihadkhawaja/ | 178 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 44 | Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu. | apache/ | 568 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 45 | 45.Bank Answer questions about the user's own bank accounts and money using the bank MCP tools (listaccounts, getbalances, gettransactions, watches). | noskillish/ | 276 | — | ~1.1k | Automated safety check: Pass | MIT | 8 days ago |
| 46 | Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable. | EGalahad/ | 145 | — | ~1.1k | Automated safety check: Pass | No licence | 9 days ago |
| 47 | A skill your agent uses when a user asks how many tokens (or how much context/cost) ContextPilot has saved, or wants a ContextPilot savings status/summary inside Hermes Agent — e.g. | EfficientContext/ | 140 | — | ~1.4k | Automated safety check: Pass | MIT | 4 days ago |
| 48 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | today |
Questions, answered from the data.
What is the best LLM inference and serving skill?
LLM Torch Profiler Analysis from sgl-project/sglang ranks first of the 364 LLM inference and serving skills listed here, with the highest score: its repository has 37k GitHub stars, 2 other GitHub owners carry a copy, its SKILL.md loads about 6.4k tokens and it passes the automated safety check with no findings. Next come SageMaker Serving Image Selection and Aider Delegate.
Which LLM inference and serving skills are official?
48 of the 364 LLM inference and serving skills are official, published by the vendor's own GitHub organization: SageMaker Serving Image Selection, Hugging Face Local Model Evals, Adapt New Diffusion Model, SageMaker Production Defaults, Debug Inference and 43 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23