Search

AI & LLM Engineering · SGLang · For developers

54 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

sgl-project/sglang37k3 repos~2.1kAutomated safety check: PassApache-2.0today
2

Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

sgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0today
3

Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

huggingface/skills11k1 repo~4.6kAutomated safety check: PassApache-2.02 days ago
4

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
5

Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.3kAutomated safety check: PassNo licence6 days ago
6

Code style for SGLang large classes Scheduler, TokenizerManager, and ModelRunner: frozen-code conventions and init orchestration style.

sgl-project/sglang37k2 repos~1.9kAutomated safety check: PassApache-2.0today
7

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.02 days ago
8
8.Debug InferenceOfficial

Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

NVIDIA/OpenShell16k—~1.9kAutomated safety check: PassApache-2.0yesterday
9

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
10

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
11

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence6 days ago
12

Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

ModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassUnknowntoday
13

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

intel/auto-round1.6k—~1.9kAutomated safety check: PassApache-2.0yesterday
14

A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
15

A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

sgl-project/sglang37k2 repos~5kAutomated safety check: PassApache-2.0today
16

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

guqiong96/Lsglang1441 repo~10kAutomated safety check: PassApache-2.06 days ago
17

Naming conventions for SGLang speculative decoding identifiers.

sgl-project/sglang37k2 repos~1.6kAutomated safety check: PassApache-2.0today
18

Validate changed quantization math, packed formats, and inference kernels against a Torch oracle; adjudicate near-tie mismatches before rejecting optimized arithmetic.

ModelCloud/GPTQModel1.3k—~1.4kAutomated safety check: PassUnknowntoday
19

Sets up a structured debugging session for a Dynamo bug — pull the report from a Linear ticket, GitHub issue, or pasted text, capture the environment, create a persistent worklog markdown file, and…

ai-dynamo/dynamo8.3k—~1.2kAutomated safety check: PassApache-2.0today
20

Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

sgl-project/sglang37k2 repos~3.4kAutomated safety check: PassApache-2.0today
21

Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

sgl-project/sglang37k2 repos~4.9kAutomated safety check: PassApache-2.0today
22

Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware…

sgl-project/sglang37k2 repos~2.5kAutomated safety check: PassApache-2.0today
23

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

sohu-mptc/FlashRec107—~915Automated safety check: PassApache-2.0yesterday
24

Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron.

Netis/heron102—~983Automated safety check: PassApache-2.06 days ago
25

Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
26

Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

azrtydxb/Fastllm-proxy108—~980Automated safety check: PassApache-2.0today
27

Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron.

Netis/heron102—~1.4kAutomated safety check: PassApache-2.06 days ago
28

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~3.9kAutomated safety check: PassNo licence6 days ago
29

Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodalgen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the…

sgl-project/sglang37k—~7.1kAutomated safety check: PassApache-2.0today
30

Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids).

sgl-project/sglang37k—~3.4kAutomated safety check: PassApache-2.0today
31

Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

amd/skills408—~1.7kAutomated safety check: NotesMITyesterday
32

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~7.5kAutomated safety check: PassNo licence6 days ago
33

Run a 4-hour multi-node Hyperloom Qwen3-30B-A3B optimization (Infera PD-disaggregated or RayJob aggregated) with --nodes 2 and sglang MoE tuning on MI325X.

AMD-AGI/Hyperloom220—~1.8kAutomated safety check: PassUnknowntoday
34

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITyesterday
35

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~1.5kAutomated safety check: PassNo licence6 days ago
36

Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL.

areal-project/AReaL5.8k—~6kAutomated safety check: PassApache-2.0yesterday
37

Audit SGLang startup logs, save evidence, and propose cleanup for user review.

sgl-project/sglang37k—~1.5kAutomated safety check: PassApache-2.0today
38

Hold / release GPUs by running a standard SGLang inference server on them (lab convention — occupy a card with a real serving job that fills ~all of its memory AND is driven at stable near-full…

cua-lite/cua-lite108—~1.9kAutomated safety check: PassNo licenceyesterday
39

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups

sgl-project/sglang37k—~13kAutomated safety check: PassApache-2.0today
40

Guide for using SLIME (LLM post-training framework for RL Scaling).

yzlnew/infra-skills149—~3.2kAutomated safety check: PassNo licence3 mo ago
41

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

ai-dynamo/dynamo8.3k—~4.9kAutomated safety check: PassApache-2.0today
42

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom219—~7.1kAutomated safety check: NotesUnknownyesterday
43

Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

NVIDIA/skills3.6k1 repo~2.3kAutomated safety check: NotesApache-2.0yesterday
44

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.6k1 repo~2.9kAutomated safety check: PassApache-2.0yesterday
45

RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1711 repo~2.7kAutomated safety check: PassMIT3 days ago
46
46.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.6k1 repo~3kAutomated safety check: NotesApache-2.0yesterday
47

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
48

Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago