Agent skill

Kernel Benchmark

by ZJLi2013 in ZJLi2013/awesome-kernel-skills

Unified kernel benchmarking protocol producing JSON results with latency, TFLOPS, GBps, and comparison against PyTorch baselines.

No licenceAuto-check passedAI & LLM Engineering

Install Kernel Benchmark

skills CLI
$ npx skills add ZJLi2013/awesome-kernel-skills --skill kernel-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZJLi2013/awesome-kernel-skills kernel-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZJLi2013/awesome-kernel-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/system/benchmark .claude/skills/kernel-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kernel-benchmark
GitHub stars
102
Token cost
~567 tokens
SKILL.md length
181 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
None found

At a glance

Unified kernel benchmarking protocol producing JSON results with latency, TFLOPS, GBps, and comparison against PyTorch baselines.

  • Works in 5 steps: Warmup → Timed Runs → Metrics Computation → …
  • Benchmarking kernels
  • SKILL.md covers Overview, Protocol and Agent Instructions
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Kernel Benchmark is an agent skill from ZJLi2013/awesome-kernel-skills. Unified kernel benchmarking protocol producing JSON results with latency, TFLOPS, GBps, and comparison against PyTorch baselines. Covers warmup, timing, and cross-platform reporting. Use when benchmarking kernels, measuring performance, or comparing implementations.

Its SKILL.md is about 570 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch.

When your agent uses it

  • Benchmarking kernels
  • Measuring performance
  • Comparing implementations

Example prompts

  • “/kernel-benchmark”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Warmup
  2. Timed Runs
  3. Metrics Computation
  4. Baseline Comparison
  5. Output Format

What it can do on your machine

Read from SKILL.md and the folder at commit aba7662. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kernel Benchmark loads about 567 tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 181 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~567

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 181 words (~567 tokens).

“Consistent benchmarking methodology across all kernels and platforms. Every benchmark produces a JSON result for automated tracking.”

— opening of SKILL.md by ZJLi2013
name
kernel-benchmark

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/system/benchmark of ZJLi2013/awesome-kernel-skills.

Open the folder on GitHubat commit aba7662

Compare with similar skills

Kernel Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kernel Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kernel Benchmark this skillZJLi2013/awesome-kernel-skills102—~567Automated safety check: PassNone
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Add Torch Shapes Examplefacebook/pyrefly7.1k—~1.3kAutomated safety check: PassMIT
Interview Cheatsheetwanshuiyin/ARIS-in-AI-Offer5741 repos~3.4kAutomated safety check: NotesMIT
Ghstack CIpytorch/pytorch104k—~1.4kAutomated safety check: PassCustom licence

Similar skills

  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Torch Shapes Example

    facebook/pyrefly

    Official

    A skill your agent uses when adding a new PyTorch model to Pyrefly's shape-tracking example corpus under tensor-shapes/pyrefly-torch-stubs/examples — i.e.

    7.1k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Interview Cheatsheet

    wanshuiyin/ARIS-in-AI-Offer

    Generate a long-form Chinese interview-prep cheat sheet on a specific ML/LLM topic — formulas with derivations, from-scratch PyTorch code, comparison tables, and 25 高频面试题 (L1 必会 / L2 进阶 / L3 顶级 lab).

    574 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check: notes
  • Ghstack CI

    pytorch/pytorch

    Manage CI for PyTorch ghstack stacks by running CI where its results are useful now and deferring other PRs with [no-ci].

    104k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from ZJLi2013/awesome-kernel-skills

All 12 skills in this repo
  • Cross Entropy Kernel

    ZJLi2013/awesome-kernel-skills

    Optimize fused cross-entropy loss kernels in Triton for NVIDIA and AMD GPUs.

    102 GitHub stars~461 tokensUpdated 6 mo ago
    Auto-check passed
  • Flash Attention Kernel

    ZJLi2013/awesome-kernel-skills

    Optimize FlashAttention-style fused attention kernels in Triton for NVIDIA and AMD GPUs.

    102 GitHub stars~697 tokensUpdated 6 mo ago
    Auto-check passed
  • Fused Moe Kernel

    ZJLi2013/awesome-kernel-skills

    Optimize Fused Mixture-of-Experts (MoE) kernels in Triton for NVIDIA and AMD GPUs.

    102 GitHub stars~699 tokensUpdated 6 mo ago
    Auto-check passed
  • Gemm Kernel Optimization

    ZJLi2013/awesome-kernel-skills

    Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs.

    102 GitHub stars~1.1k tokensUpdated 6 mo ago
    Auto-check passed
  • Iterative Kernel Optimization Loop

    ZJLi2013/awesome-kernel-skills

    Orchestrates continuous kernel optimization by chaining profiling, bottleneck diagnosis, tier-based optimization, verification, and benchmarking into an iterative loop.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    Auto-check passed
  • Kernel Profiling

    ZJLi2013/awesome-kernel-skills

    Profile GPU kernels using NCU (NVIDIA) or rocprof (AMD) to collect performance metrics.

    102 GitHub stars~696 tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Kernel Benchmark

What does Kernel Benchmark do?

Unified kernel benchmarking protocol producing JSON results with latency, TFLOPS, GBps, and comparison against PyTorch baselines. Kernel Benchmark is an agent skill from ZJLi2013/awesome-kernel-skills. Unified kernel benchmarking protocol producing JSON results with latency, TFLOPS, GBps, and comparison against PyTorch baselines.

When should I use Kernel Benchmark?

Kernel Benchmark fits situations like: benchmarking kernels; measuring performance; comparing implementations.

How do I install Kernel Benchmark in Claude Code?

Run `npx skills add ZJLi2013/awesome-kernel-skills --skill kernel-benchmark -a claude-code`. Or copy the skill folder (skills/system/benchmark in ZJLi2013/awesome-kernel-skills) into .claude/skills/kernel-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Kernel Benchmark in Codex?

Run `npx skills add ZJLi2013/awesome-kernel-skills --skill kernel-benchmark -a codex`. Or copy the skill folder (skills/system/benchmark in ZJLi2013/awesome-kernel-skills) into .agents/skills/kernel-benchmark in your project. Codex loads it when a task matches its description.

Can I use Kernel Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZJLi2013/awesome-kernel-skills --skill kernel-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kernel-benchmark, .gemini/skills/kernel-benchmark, .github/skills/kernel-benchmark and .opencode/skills/kernel-benchmark in your project.

What does Kernel Benchmark need to run?

SKILL.md names no scripts, command-line tools or credentials: Kernel Benchmark is instructions for the agent only. Our summary lists: Python 3.

Does Kernel Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Kernel Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kernel Benchmark use?

No licence was found for Kernel Benchmark or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Kernel Benchmark use?

About 567 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kernel Benchmark?

Skills that share tags, products or a category with Kernel Benchmark: Add Uint Support (pytorch/pytorch, 104k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Torch Shapes Example (facebook/pyrefly, 7.1k stars) and Interview Cheatsheet (wanshuiyin/ARIS-in-AI-Offer, 574 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kernel Benchmark?

ZJLi2013 (a GitHub user) maintains it in ZJLi2013/awesome-kernel-skills, which has 102 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on March 31, 2026.

Source: ZJLi2013/awesome-kernel-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.