Search

CUDA · Performance optimization

14 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

stas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.03 days ago
2

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
3

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.011 days ago
4

CUDA kernel development, debugging, and performance optimization for Claude Code.

technillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNo licence9 mo ago
5
5.Cpp

Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics.

crazyguitar/cppcheatsheet290—~1.8kAutomated safety check: PassMIT2 days ago
6

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNo licence4 days ago
7

Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

slowlyC/agent-gpu-skills169—~1.8kAutomated safety check: PassMIT2 mo ago
8

A skill your agent uses for performance profiling and optimization.

ByteDance-Seed/VeOmni2.2k—~1.7kAutomated safety check: PassApache-2.0today
9

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2kAutomated safety check: PassNo licence4 days ago
10

Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly.

sablin39/tilelang-cuda-skills145—~1.6kAutomated safety check: PassNo licence24 days ago
11

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills406—~2.3kAutomated safety check: PassMITtoday
12

Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

NVIDIA/skills3.5k—~3.3kAutomated safety check: PassApache-2.0today
13

AI for Science 场景下的昇腾 NPU Profiling 采集与性能分析 Skill,用于在华为 Ascend NPU 上使用 torchnpu.profiler 采集 L0、L1、L2 级性能数据,分析训练或推理中的算子耗时、调用栈、内存与瓶颈,并指导后续调优。

ascend-ai-coding/awesome-ascend-skills174—~3kAutomated safety check: PassNo licencetoday
14

CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills253—~1.6kAutomated safety check: NotesMIT3 mo ago