Search
NVIDIA AI Platform · Performance optimization
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 11 days ago |
| 3 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 925 | — | ~2.8k | Automated safety check: Pass | No licence | 4 days ago |
| 4 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 5 | Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 6 | Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel. | alibaba/ | 166 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 7 | Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use. | davila7/ | 32k | 10 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 8 | Profile GPU kernels using NCU (NVIDIA) or rocprof (AMD) to collect performance metrics. | ZJLi2013/ | 102 | — | ~696 | Automated safety check: Pass | No licence | 6 mo ago |
| 9 | Profile a target (script, process, GPU, memory, interconnect) using external tools and code instrumentation. | AI4Scientist/ | 128 | 4 repos | ~1.1k | Automated safety check: Pass | No licence | 4 mo ago |
| 10 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 11 | Evidence-gated workflow for MoE performance optimization in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |