Search
LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. | huggingface/ | 11k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself. | amElnagdy/ | 2.3k | 2 repos | ~3k | Automated safety check: Pass | MIT | 2 days ago |
| 4 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. | vllm-project/ | 7.1k | — | ~7.5k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements. | vllm-project/ | 2.9k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | 7.Perfup Autonomous performance optimization: research, PoC, benchmark, implement, review, PR | raullenchai/ | 3.9k | — | ~1.6k | Automated safety check: Notes | Unknown | today |
| 8 | Bump or upgrade the pinned versions of Helmor's bundled agent CLIs, SDKs, and supporting binaries — Claude Code + claude-agent-sdk (lockstep), Codex, Cursor SDK, OpenCode, Kimi, Pi, and gh / glab /… | dohooo/ | 1.3k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 9 | Run the pre-release end-to-end sweep of every user-facing surface — the 33 samples under samples/ (booted from their packaged artifacts and driven in a real browser via chrome-devtools MCP), the… | Atmosphere/ | 3.8k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands. | internet-court/ | 6.5k | 1 repo | ~1.9k | Automated safety check: Pass | Unknown | 1 mo ago |
| 11 | Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. | R6410418/ | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 12 | Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm. | guqiong96/ | 465 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 13 | 13.Create Skill Create a new skill in the current repository. An agent skill from Hyk260/PureChat. | Hyk260/ | 546 | 1 repo | ~823 | Automated safety check: Pass | MIT | 23 days ago |
| 14 | Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence. | BBuf/ | 925 | — | ~2.3k | Automated safety check: Pass | No licence | 4 days ago |
| 15 | Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups. | huggingface/ | 11k | 1 repo | ~6.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 16 | 16.Quantization Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. | vllm-project/ | 7.1k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | today |
| 18 | Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using… | perminder-klair/ | 1.4k | — | ~2.4k | Automated safety check: Notes | MIT | yesterday |
| 19 | Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT). | intel/ | 1.6k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 465 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 21 | Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. | NVIDIA/ | 16k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 22 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 925 | — | ~2.5k | Automated safety check: Pass | No licence | 4 days ago |
| 23 | Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. | vllm-project/ | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | 24.Review PR Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings. | vllm-project/ | 7.1k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 26 | 26.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 11 days ago |
| 27 | A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing… | Mesh-LLM/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 28 | Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box. | intel/ | 1.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 29 | 29.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 30 | Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. | huggingface/ | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 31 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 925 | — | ~2.8k | Automated safety check: Pass | No licence | 4 days ago |
| 32 | Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents. | danyuchn/ | 249 | — | ~2.7k | Automated safety check: Pass | MIT | 7 days ago |
| 33 | Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the… | amd/ | 406 | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 34 | 34.Astrea A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform… | warpfront/ | 655 | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 35 | Use iDeer as a daily paper-reading workflow for chatbot-first users such as Codex, Gemini, or ChatGPT. | AI45Lab/ | 416 | — | ~3k | Automated safety check: Notes | AGPL-3.0 | 2 mo ago |
| 36 | Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems. | ModelCloud/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Unknown | yesterday |
| 37 | Run wally's LLM e2e on Apple Neural Engine (NeuRT) and Snapdragon Hexagon NPU (QHexRT) devices. | RunanywhereAI/ | 1.6k | — | ~845 | Automated safety check: Pass | MIT | today |
| 38 | Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor). | intel/ | 1.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 39 | 39.Add Model Adapt and port new LLM model architectures to this xinfer project. | guoqingbao/ | 334 | — | ~4.2k | Automated safety check: Notes | MIT | 1 mo ago |
| 40 | Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner). | tryonlabs/ | 551 | — | ~1.1k | Automated safety check: Pass | Unknown | yesterday |
| 41 | Check if package.json files are in sync with pnpm-lock.yaml. | Hyk260/ | 546 | — | ~594 | Automated safety check: Pass | MIT | 23 days ago |
| 42 | 42.Vision Query images with a local Ollama vision model without loading the image into the main agent context. | gridaco/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 43 | Build, extend, and debug AI agents in Egroo using the Microsoft Agent Framework (C .NET). | jihadkhawaja/ | 178 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 44 | 44.Bank Answer questions about the user's own bank accounts and money using the bank MCP tools (listaccounts, getbalances, gettransactions, watches). | noskillish/ | 277 | — | ~1.1k | Automated safety check: Pass | MIT | 11 days ago |
| 45 | Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu. | apache/ | 568 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 46 | Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT… | vllm-project/ | 7.1k | — | ~7k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable. | EGalahad/ | 145 | — | ~1.1k | Automated safety check: Pass | No licence | 11 days ago |
| 48 | A skill your agent uses when a user asks how many tokens (or how much context/cost) ContextPilot has saved, or wants a ContextPilot savings status/summary inside Hermes Agent — e.g. | EfficientContext/ | 140 | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |