Search
By ZJLi2013
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Optimize fused cross-entropy loss kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~461 | Automated safety check: Pass | No licence | 6 mo ago |
| 2 | Optimize FlashAttention-style fused attention kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~697 | Automated safety check: Pass | No licence | 6 mo ago |
| 3 | Optimize Fused Mixture-of-Experts (MoE) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~699 | Automated safety check: Pass | No licence | 6 mo ago |
| 4 | Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~1.1k | Automated safety check: Pass | No licence | 6 mo ago |
| 5 | Orchestrates continuous kernel optimization by chaining profiling, bottleneck diagnosis, tier-based optimization, verification, and benchmarking into an iterative loop. | ZJLi2013/ | 102 | — | ~2.5k | Automated safety check: Pass | No licence | 6 mo ago |
| 6 | Unified kernel benchmarking protocol producing JSON results with latency, TFLOPS, GBps, and comparison against PyTorch baselines. | ZJLi2013/ | 102 | — | ~567 | Automated safety check: Pass | No licence | 6 mo ago |
| 7 | Profile GPU kernels using NCU (NVIDIA) or rocprof (AMD) to collect performance metrics. | ZJLi2013/ | 102 | — | ~696 | Automated safety check: Pass | No licence | 6 mo ago |
| 8 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 9 | Optimize RMS Normalization kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~728 | Automated safety check: Pass | No licence | 6 mo ago |
| 10 | Optimize fused softmax kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~742 | Automated safety check: Pass | No licence | 6 mo ago |
| 11 | Optimize Rotary Position Embedding (RoPE) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~421 | Automated safety check: Pass | No licence | 6 mo ago |
| 12 | Classifies GPU kernel bottlenecks (memory, compute, latency) from profiling metrics and applies GEAK-style workload guidance: what to prefer, consider, or deprioritize. | ZJLi2013/ | 102 | — | ~992 | Automated safety check: Pass | No licence | 6 mo ago |