Search
CUDA · Performance optimization
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 3 days ago |
| 2 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 3 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 11 days ago |
| 4 | 4.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 5 | 5.Cpp Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics. | crazyguitar/ | 290 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 6 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 925 | — | ~2.8k | Automated safety check: Pass | No licence | 4 days ago |
| 7 | Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 8 | A skill your agent uses for performance profiling and optimization. | ByteDance-Seed/ | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 925 | — | ~2k | Automated safety check: Pass | No licence | 4 days ago |
| 10 | Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly. | sablin39/ | 145 | — | ~1.6k | Automated safety check: Pass | No licence | 24 days ago |
| 11 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 406 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 12 | Evidence-gated workflow for MoE performance optimization in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | AI for Science 场景下的昇腾 NPU Profiling 采集与性能分析 Skill,用于在华为 Ascend NPU 上使用 torchnpu.profiler 采集 L0、L1、L2 级性能数据,分析训练或推理中的算子耗时、调用栈、内存与瓶颈,并指导后续调优。 | ascend-ai-coding/ | 174 | — | ~3k | Automated safety check: Pass | No licence | today |
| 14 | CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |