Search
SGLang · LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. | huggingface/ | 11k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 3 | Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence. | BBuf/ | 938 | — | ~2.3k | Automated safety check: Pass | No licence | 6 days ago |
| 4 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 5 | Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. | NVIDIA/ | 16k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 6 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 938 | — | ~2.5k | Automated safety check: Pass | No licence | 6 days ago |
| 7 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | 8.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 9 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 6 days ago |
| 10 | Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems. | ModelCloud/ | 1.3k | — | ~1.1k | Automated safety check: Pass | Unknown | today |
| 11 | Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor). | intel/ | 1.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 12 | Naming conventions for SGLang speculative decoding identifiers. | sgl-project/ | 37k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | Validate changed quantization math, packed formats, and inference kernels against a Torch oracle; adjudicate near-tie mismatches before rejecting optimized arithmetic. | ModelCloud/ | 1.3k | — | ~1.4k | Automated safety check: Pass | Unknown | today |
| 14 | 14.Recif Eval Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. | sohu-mptc/ | 107 | — | ~915 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | 15.Dev Bump Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron. | Netis/ | 102 | — | ~983 | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 16 | Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding… | azrtydxb/ | 108 | — | ~980 | Automated safety check: Pass | Apache-2.0 | today |
| 18 | Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron. | Netis/ | 102 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 19 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 938 | — | ~3.9k | Automated safety check: Pass | No licence | 6 days ago |
| 20 | Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. | amd/ | 408 | — | ~1.7k | Automated safety check: Notes | MIT | yesterday |
| 21 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 938 | — | ~7.5k | Automated safety check: Pass | No licence | 6 days ago |
| 22 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 23 | A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. | BBuf/ | 938 | — | ~1.5k | Automated safety check: Pass | No licence | 6 days ago |
| 24 | Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and… | ai-dynamo/ | 8.3k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom. | AMD-AGI/ | 219 | — | ~7.1k | Automated safety check: Notes | Unknown | yesterday |
| 26 | Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson. | NVIDIA/ | 3.6k | 1 repo | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 27 | Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin. | NVIDIA/ | 3.6k | 1 repo | ~3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 28 | Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. | agentsope/ | 436 | — | ~6.1k | Automated safety check: Pass | MIT | 2 days ago |
| 29 | Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy. | agentsope/ | 436 | — | ~6.1k | Automated safety check: Pass | MIT | 2 days ago |
| 30 | Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics. | benchflow-ai/ | 1.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 31 | End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. | amd/ | 182 | — | ~6.2k | Automated safety check: Pass | MIT | 13 days ago |
| 32 | LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 105 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 33 | 33.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 34 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |