Search
Development · GPU and accelerator computing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 3 | Adds Liger Kernel support for a new HuggingFace Transformers model, or modifies existing monkey-patching. | linkedin/ | 6.7k | — | ~1.3k | Automated safety check: Pass | BSD-2-Clause | today |
| 4 | 4.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 5 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 925 | — | ~2.8k | Automated safety check: Pass | No licence | 4 days ago |
| 6 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 7 | Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches. | pytorch/ | 104k | — | ~3.5k | Automated safety check: Pass | Unknown | today |
| 8 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 925 | — | ~2k | Automated safety check: Pass | No licence | 4 days ago |
| 9 | 9.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 23 days ago |
| 10 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 11 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 925 | — | ~3.9k | Automated safety check: Pass | No licence | 4 days ago |
| 13 | Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX). | facebookexperimental/ | 201 | — | ~644 | Automated safety check: Pass | MIT | today |
| 14 | Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming). | yzlnew/ | 149 | — | ~2.4k | Automated safety check: Pass | No licence | 3 mo ago |
| 15 | Author Triton kernels with automatic warp specialization (AutoWS). | facebookexperimental/ | 201 | — | ~7.1k | Automated safety check: Pass | MIT | today |
| 16 | Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store… | facebookexperimental/ | 201 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 17 | A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under… | NVIDIA/ | 3.5k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 18 | CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 19 | CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |
| 20 | 20.Hip Rocm HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |