Gemm Kernel Optimization is an agent skill from ZJLi2013/awesome-kernel-skills. Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs. Covers tiled blocking, L2 cache grouping, tensor core utilization, and dual-platform autotune. Use when writing or optimizing matmul, linear layers, or any GEMM-based kernel.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `test_gemm.py` and `triton_template.py`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with NVIDIA AI Platform.