Search
By Krusty84
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably. | Krusty84/ | 106 | — | ~564 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 2 | Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering. | Krusty84/ | 106 | — | ~597 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 3 | Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs. | Krusty84/ | 106 | — | ~662 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 4 | Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation. | Krusty84/ | 106 | — | ~753 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 5 | Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs. | Krusty84/ | 106 | — | ~636 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 6 | Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles. | Krusty84/ | 106 | — | ~659 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 7 | Build Ascend-friendly MoE gather, scatter, and router-weight-gradient kernels with vector-core-sized grids, UB-aware row/column tiling, and CANN slice extensions. | Krusty84/ | 106 | — | ~696 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 8 | Build padded MoE gather, scatter, and router-weight-gradient kernels for Triton-Ascend using real cumulative bins, padded cumulative bins, UB slice assembly, and vector-core-aware tiling. | Krusty84/ | 106 | — | ~652 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 9 | Build a single-program Triton-Ascend matrix multiplication with fused bias using two-dimensional pointer grids and tl.dot. | Krusty84/ | 106 | — | ~576 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 10 | Build and debug masked one-dimensional elementwise kernels for Triton-Ascend, including launch wrappers and PyTorch/NPU correctness checks. | Krusty84/ | 106 | — | ~584 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 11 | Refactor fused causal Conv1d state updates for Triton-Ascend by replacing negative-offset cat emulation and tl.where selection with transposed UB assembly through extension.insertslice. | Krusty84/ | 106 | — | ~646 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 12 | Apply Triton-Ascend maxautotune to search base kernel configurations together with Ascend compiler options for vector, cube, or mixed operators. | Krusty84/ | 106 | — | ~595 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 13 | Optimize bounded discrete reads in Triton-Ascend by staging a contiguous source vector in UB and selecting elements with tl.gather. | Krusty84/ | 106 | — | ~517 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 14 | Optimize irregular KV-cache access in Triton-Ascend grouped decode attention by vector-loading contiguous inner dimensions, assembling discrete outer rows with CANN slice extensions, and transposing… | Krusty84/ | 106 | — | ~658 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 15 | Optimize integer vector kernels for Triton-Ascend by preferring int32 arithmetic when the value and index ranges permit it. | Krusty84/ | 106 | — | ~545 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 16 | Prevent Triton-Ascend Unified Buffer overflow by streaming large logical vectors through bounded compile-time tiles with masked loop iterations. | Krusty84/ | 106 | — | ~558 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 17 | Generate TTIR for Triton-Ascend kernel configurations and rank them with the Ascend costmodel backend before empirical autotuning. | Krusty84/ | 106 | — | ~721 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 18 | Reorder independent Triton-Ascend loads to expose overlap with loop-carried stores and reduce dependency stalls. | Krusty84/ | 106 | — | ~558 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 19 | Use Triton-Ascend tl.load carepadding=False safely to skip initialization of masked tail lanes and remove unnecessary padding work. | Krusty84/ | 106 | — | ~480 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 20 | Tile two-dimensional torch.gather(dim=1) kernels for Triton-Ascend so the logical grid stays near the physical vector-core count while each program handles multiple batch rows and loops over K. | Krusty84/ | 106 | — | ~583 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 21 | Add standard or advanced autotuning to Triton-Ascend kernels, define shape keys and candidate meta-parameters, and structure kernels so automatic split and tiling analysis can recognize their axes. | Krusty84/ | 106 | — | ~813 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 22 | Validate Triton-Ascend kernel outputs against PyTorch references with dtype-aware tolerances, exact integer checks, bfloat16 promotion, and boolean handling. | Krusty84/ | 106 | — | ~649 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 23 | Vectorize explicit integer comparisons in Triton-Ascend by casting bounded index vectors to float32 before tl.where or similar compute expressions. | Krusty84/ | 106 | — | ~498 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |