AI model or service
SGLang agent skills, page 2
SGLang skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Dev Bump Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron. | Netis/ | 102 | — | ~983 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 50 | How SGLang's runtime configuration and process-global state are organized (RuntimeContext tiers, publish + namespace config bags, the pristine ServerArgs seed, override entry points… | sgl-project/ | 37k | 2 repos | ~11k | Automated safety check: Pass | Apache-2.0 | today |
| 51 | Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding… | azrtydxb/ | 108 | — | ~980 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 52 | Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron. | Netis/ | 102 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 53 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 911 | — | ~3.9k | Automated safety check: Pass | No licence | 3 days ago |
| 54 | Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodalgen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the… | sgl-project/ | 37k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | today |
| 55 | Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer. | amd/ | 398 | — | ~1.7k | Automated safety check: Notes | MIT | today |
| 56 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 911 | — | ~7.5k | Automated safety check: Pass | No licence | 3 days ago |
| 57 | Capture torch.profiler traces on a running FlashRec server. An agent skill from sohu-mptc/FlashRec. | sohu-mptc/ | 107 | — | ~312 | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 58 | Run a 4-hour multi-node Hyperloom Qwen3-30B-A3B optimization (Infera PD-disaggregated or RayJob aggregated) with --nodes 2 and sglang MoE tuning on MI325X. | AMD-AGI/ | 217 | — | ~1.8k | Automated safety check: Pass | Unknown | today |
| 59 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 398 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 60 | Audit the existing test tree and CI configuration for improvements, using a catalog of patterns previously applied in this repo. | sgl-project/ | 37k | — | ~520 | Automated safety check: Pass | Apache-2.0 | today |
| 61 | A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. | BBuf/ | 911 | — | ~1.5k | Automated safety check: Pass | No licence | 3 days ago |
| 62 | 62.Upgrade Deps Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL. | areal-project/ | 5.8k | — | ~6k | Automated safety check: Pass | Apache-2.0 | today |
| 63 | Audit SGLang startup logs, save evidence, and propose cleanup for user review. | sgl-project/ | 37k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 64 | 64.Watchdog Hold / release GPUs by running a standard SGLang inference server on them (lab convention — occupy a card with a real serving job that fills ~all of its memory AND is driven at stable near-full… | cua-lite/ | 106 | — | ~1.9k | Automated safety check: Pass | No licence | 12 days ago |
| 65 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 66 | 66.Slime User Guide for using SLIME (LLM post-training framework for RL Scaling). | yzlnew/ | 149 | — | ~3.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 67 | Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations. | sgl-project/ | 37k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 68 | Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and… | ai-dynamo/ | 8.2k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 69 | Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom. | AMD-AGI/ | 217 | — | ~7.2k | Automated safety check: Notes | Unknown | today |
| 70 | 70.Slime RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | today |
| 71 | Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. | agentsope/ | 459 | — | ~6.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 72 | Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy. | agentsope/ | 459 | — | ~6.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 73 | Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it… | autonomous-ai/ | 1.1k | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 74 | Generate videos using a local SGLang-Diffusion server (Wan2.2, Hunyuan, FastWan, etc.). | LeoYeAI/ | 2.2k | — | ~676 | Automated safety check: Pass | MIT | 2 mo ago |
| 75 | Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics. | benchflow-ai/ | 1.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 76 | End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks. | amd/ | 181 | — | ~6.2k | Automated safety check: Pass | MIT | 10 days ago |
| 77 | LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 103 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 78 | 78.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |
| 79 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |