Search

NVIDIA AI Platform · Performance optimization

12 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

sgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0today
2

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.011 days ago
3

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNo licence4 days ago
4

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT3 mo ago
5

Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

slowlyC/agent-gpu-skills169—~1.8kAutomated safety check: PassMIT2 mo ago
6

Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.

alibaba/atrex-kernel-agent166—~5.2kAutomated safety check: PassApache-2.010 days ago
7

Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use.

davila7/claude-code-templates32k10 repos~2.4kAutomated safety check: PassMITtoday
8

Profile GPU kernels using NCU (NVIDIA) or rocprof (AMD) to collect performance metrics.

ZJLi2013/awesome-kernel-skills102—~696Automated safety check: PassNo licence6 mo ago
9

Profile a target (script, process, GPU, memory, interconnect) using external tools and code instrumentation.

AI4Scientist/nano-scientist1284 repos~1.1kAutomated safety check: PassNo licence4 mo ago
10

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

majiayu000/spellbook287—~1.1kAutomated safety check: PassMITyesterday
11

Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

NVIDIA/skills3.5k—~3.3kAutomated safety check: PassApache-2.0today
12

CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills253—~1.6kAutomated safety check: NotesMIT3 mo ago