Agent skill

Gemm Kernel Optimization

by ZJLi2013 in ZJLi2013/awesome-kernel-skills

Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs.

No licenceAuto-check passedAI & LLM Engineering

Install Gemm Kernel Optimization

skills CLI
$ npx skills add ZJLi2013/awesome-kernel-skills --skill gemm-kernel-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZJLi2013/awesome-kernel-skills gemm-kernel-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZJLi2013/awesome-kernel-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/kernels/gemm .claude/skills/gemm-kernel-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemm-kernel-optimization
GitHub stars
102
Token cost
~1.1k tokens
SKILL.md length
357 words
Files
3
Skills in repo
12
Repo updated
First seen
Licence
None found

At a glance

Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs.

  • Works in 4 steps: Partition output C into [BLOCK_M,… → Accumulate in FP32: iterate K dimension… → Each iteration: load A tile [BLOCK_M,… → …
  • Optimizing matmul
  • SKILL.md covers Overview, Core Technique, Autotune Configs and Verification, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Gemm Kernel Optimization is an agent skill from ZJLi2013/awesome-kernel-skills. Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs. Covers tiled blocking, L2 cache grouping, tensor core utilization, and dual-platform autotune. Use when writing or optimizing matmul, linear layers, or any GEMM-based kernel.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `test_gemm.py` and `triton_template.py`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with NVIDIA AI Platform.

When your agent uses it

  • Optimizing matmul
  • Any GEMM-based kernel

Example prompts

  • “/gemm-kernel-optimization”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Partition output C into [BLOCK_M, BLOCK_N] tiles, one per program instance
  2. Accumulate in FP32: iterate K dimension in BLOCK_K chunks
  3. Each iteration: load A tile [BLOCK_M, BLOCK_K], B tile [BLOCK_K, BLOCK_N], call tl.dot
  4. Store result tile back to C (cast to output dtype)

What it can do on your machine

Read from SKILL.md and the folder at commit aba7662. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • triton-lang.org
    • github.com
    • research.colfax-intl.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemm Kernel Optimization loads about 1.1k tokens when it runs. Until then it costs about 72 tokens; SKILL.md has 357 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 357 words (~1,112 tokens).

“General Matrix Multiply (C = A @ B) is the most compute-intensive operation in deep learning. GEMM is compute-bound: the key metric is TFLOPS (target: >60% of hardware peak).”

— opening of SKILL.md by ZJLi2013
name
gemm-kernel-optimization

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files in skills/kernels/gemm of ZJLi2013/awesome-kernel-skills.

  • SKILL.md
  • test_gemm.py
  • triton_template.py

Open the folder on GitHubat commit aba7662

Compare with similar skills

Gemm Kernel Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemm Kernel Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemm Kernel Optimization this skillZJLi2013/awesome-kernel-skills102—~1.1kAutomated safety check: PassNone
Fla Triton To Gluonfla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMIT
DGX Spark Memory and Thermal Opswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
DGX Spark Training Gotchaswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Physical AI Video Augmentation on OSMONVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0

Similar skills

  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • OpenVLA-OFT Fine-Tuning

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

    13k GitHub starsUsed in 1 repo~3.7k tokens
    AI & LLM EngineeringAuto-check passed

More from ZJLi2013/awesome-kernel-skills

All 12 skills in this repo
  • Cross Entropy Kernel

    ZJLi2013/awesome-kernel-skills

    Optimize fused cross-entropy loss kernels in Triton for NVIDIA and AMD GPUs.

    102 GitHub stars~461 tokensUpdated 6 mo ago
    Auto-check passed
  • Flash Attention Kernel

    ZJLi2013/awesome-kernel-skills

    Optimize FlashAttention-style fused attention kernels in Triton for NVIDIA and AMD GPUs.

    102 GitHub stars~697 tokensUpdated 6 mo ago
    Auto-check passed
  • Fused Moe Kernel

    ZJLi2013/awesome-kernel-skills

    Optimize Fused Mixture-of-Experts (MoE) kernels in Triton for NVIDIA and AMD GPUs.

    102 GitHub stars~699 tokensUpdated 6 mo ago
    Auto-check passed
  • Iterative Kernel Optimization Loop

    ZJLi2013/awesome-kernel-skills

    Orchestrates continuous kernel optimization by chaining profiling, bottleneck diagnosis, tier-based optimization, verification, and benchmarking into an iterative loop.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    Auto-check passed
  • Kernel Benchmark

    ZJLi2013/awesome-kernel-skills

    Unified kernel benchmarking protocol producing JSON results with latency, TFLOPS, GBps, and comparison against PyTorch baselines.

    102 GitHub stars~567 tokensUpdated 6 mo ago
    Auto-check passed
  • Kernel Profiling

    ZJLi2013/awesome-kernel-skills

    Profile GPU kernels using NCU (NVIDIA) or rocprof (AMD) to collect performance metrics.

    102 GitHub stars~696 tokensUpdated 6 mo ago
    Auto-check passed

Questions about Gemm Kernel Optimization

What does Gemm Kernel Optimization do?

Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs. Gemm Kernel Optimization is an agent skill from ZJLi2013/awesome-kernel-skills. Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs.

When should I use Gemm Kernel Optimization?

Gemm Kernel Optimization fits situations like: optimizing matmul; any GEMM-based kernel.

How do I install Gemm Kernel Optimization in Claude Code?

Run `npx skills add ZJLi2013/awesome-kernel-skills --skill gemm-kernel-optimization -a claude-code`. Or copy the skill folder (skills/kernels/gemm in ZJLi2013/awesome-kernel-skills) into .claude/skills/gemm-kernel-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Gemm Kernel Optimization in Codex?

Run `npx skills add ZJLi2013/awesome-kernel-skills --skill gemm-kernel-optimization -a codex`. Or copy the skill folder (skills/kernels/gemm in ZJLi2013/awesome-kernel-skills) into .agents/skills/gemm-kernel-optimization in your project. Codex loads it when a task matches its description.

Can I use Gemm Kernel Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZJLi2013/awesome-kernel-skills --skill gemm-kernel-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemm-kernel-optimization, .gemini/skills/gemm-kernel-optimization, .github/skills/gemm-kernel-optimization and .opencode/skills/gemm-kernel-optimization in your project.

What does Gemm Kernel Optimization need to run?

Going by SKILL.md and its folder, Gemm Kernel Optimization needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Gemm Kernel Optimization access the network?

SKILL.md names 3 domains. As links in the text: triton-lang.org, github.com and research.colfax-intl.com. This is read from the text; nothing was executed.

Is Gemm Kernel Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemm Kernel Optimization use?

No licence was found for Gemm Kernel Optimization or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Gemm Kernel Optimization use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemm Kernel Optimization?

Skills that share tags, products or a category with Gemm Kernel Optimization: Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars), DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemm Kernel Optimization?

ZJLi2013 (a GitHub user) maintains it in ZJLi2013/awesome-kernel-skills, which has 102 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on March 31, 2026.

Source: ZJLi2013/awesome-kernel-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.