Agent skill

Kernel Microbenchmark

by guqiong96 in guqiong96/Lvllm

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Kernel Microbenchmark

skills CLI
$ npx skills add guqiong96/Lvllm --skill kernel-microbenchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guqiong96/Lvllm kernel-microbenchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/kernel-microbenchmark .claude/skills/kernel-microbenchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kernel-microbenchmark
GitHub stars
464
Used in
2 other repos
Token cost
~1.5k tokens
SKILL.md length
784 words
Files
4
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

  • Works in 6 steps: Create an isolated repro or benchmark… → Check correctness before timing. Keep… → Time only the operation under study.… → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Workflow, Benchmark Defaults, Sanity-Check Reference Numbers and Multi-GPU Benchmarks, plus 1 more section
  • Runs Python scripts from its folder

What it does

Kernel Microbenchmark is an agent skill from guqiong96/Lvllm. Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL sanity checks.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `agents/openai.yaml`, `benchmarks/cupti_microbenchmark.py` and `benchmarks/multi_gpu_gemm_rs.py`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and CUDA. The repository describes itself as: LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/kernel-microbenchmark”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Create an isolated repro or benchmark when the existing harness is noisy.
  2. Check correctness before timing. Keep tolerances explicit.
  3. Time only the operation under study. Exclude allocation, compilation, random
  4. Compare against a baseline and report enough metadata to reproduce the
  5. Treat explanations as hypotheses until backed by an artifact: ablation,
  6. If the result changes the conclusion, preserve the compact lesson in a note,

What it can do on your machine

Read from SKILL.md and the folder at commit 43ffc42. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kernel Microbenchmark loads about 1.5k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 784 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guqiong96/Lvllm at commit 43ffc42, republished under its Apache-2.0 licence (© guqiong96). 784 words, ~1,531 tokens.

Download SKILL.mdSave it as .claude/skills/kernel-microbenchmark/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
kernel-microbenchmark
description
Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL sanity checks.

Kernel Microbenchmark

Workflow

  1. Create an isolated repro or benchmark when the existing harness is noisy.
  2. Check correctness before timing. Keep tolerances explicit.
  3. Time only the operation under study. Exclude allocation, compilation, random input generation, logging, and host-device transfers unless those are the target.
  4. Compare against a baseline and report enough metadata to reproduce the result: GPU, dtype, shape, command, branch/commit, and relevant env vars.
  5. Treat explanations as hypotheses until backed by an artifact: ablation, generated PTX/SASS, profiler output, or controlled benchmark.
  6. If the result changes the conclusion, preserve the compact lesson in a note, comment, benchmark table, or final summary. For experiments, a short Question / Change / Correctness / Result / Observation / Next note is usually enough.

Benchmark Defaults

  • Use FlashInfer CUPTI timing by default, with CUDA graph and cold L2 cache: from flashinfer.testing import bench_gpu_time_with_cupti.
  • For compute-heavy kernels, report TFLOPS with the FLOP formula in the benchmark. For memory-heavy kernels, report estimated bytes moved and GB/s. For mixed kernels, report the most honest metric available and call out the caveats. TFLOPS and memory bandwidth should be computed from the theoretical best for the operation, not from a particular kernel implementation. For example, memory bandwidth should assume all data is read exactly once from global memory.
  • When comparing across shapes, prefer throughput metrics such as TFLOPS or GB/s as the primary table columns; keep latency for absolute cost.
  • If a result exceeds expected peak/SOL, first inspect units, FLOP/byte formulas, skipped work, sparsity, caching, and whether the baseline is doing the same operation.
  • Force compilation/autotuning before measuring compiled kernels.
  • Seed inputs when correctness comparisons matter.
  • Keep metadata setup, plan construction, allocation, random input generation, and logging outside the timed region unless that overhead is the experiment.

Sanity-Check Reference Numbers

Use these as rough reference points for large, well-shaped workloads, not as gold standards, guaranteed peaks, or hard limits. Hardware SKU, clocks, shape, precision conventions, and the FLOP/byte accounting can move the result. A large gap is a prompt to investigate, not proof that a kernel is poor.

Kernel regimeHardwareRough reference
Memory-bound, large batchBlackwell6 TB/s
BF16 GEMMBlackwell2 PFLOP/s
BF16 attentionB2001.6 PFLOP/s
FP8 GEMMBlackwell4 PFLOP/s
FP8 attentionB3002.8 PFLOP/s

The BF16 attention reference is approximately the 1613 TFLOP/s result reported by the FlashAttention-4 paper. Compare kernels only with matching workload and throughput conventions.

Show full SKILL.md (390 more words)Show less

Multi-GPU Benchmarks

  • State whether the run is local or multi-node and report the GPU topology, world size, GPUs per node, collective backend, and relevant library versions.
  • Compare like-for-like TP configurations. Report both per-rank and global dimensions, and do not compare results from different TP sizes without an explicit normalization or scaling question.
  • Define the timed operation boundary before benchmarking. If the production wrapper performs input staging, flag resets, generation barriers, padding, or output copies, keep them in the timed region for an end-to-end comparison. Use a separate, clearly labeled ablation for kernel-only timing.
  • Check distributed correctness before timing. Seed each rank deliberately, form the reference with the same collective semantics, and synchronize before reading or comparing outputs.
  • Use a device-side barrier immediately before each measured replay. Keep this common synchronization outside the timed interval, but keep barriers required by the candidate implementation inside it.
  • With CUDA graphs, warm up before capture, coordinate capture across ranks, rotate pointer-distinct graphs when cache reuse matters, and ensure every collective is issued in the same order on every rank.
  • Measure every rank and reduce each sample with MAX; report the median of those per-sample maxima. A rank-local event time is not a distributed latency.
  • Reset reusable symmetric-memory flags before each invocation and establish a device-side generation barrier before peers may signal them. Otherwise a fast rank can signal before a slow rank resets its flags, causing a lost arrival and intermittent deadlock.
  • Preserve symmetric-memory handles, CUDA graphs, streams, and graph outputs for the full measurement lifetime. Allocate and rendezvous symmetric buffers in identical order and with identical shapes on all ranks.
  • Treat hangs and isolated millisecond outliers as synchronization bugs or rank skew until disproven. Add stage markers, bounded synchronization checks, and per-rank diagnostics before blaming compilation or kernel performance.
  • For multi-node runs, record the scheduler allocation and verify the fabric supports the required multicast or NVLink-domain assumptions. Do not describe a two-node TP run as equivalent to a local NVLink-domain run without checking.
  • Stabilize GPU clocks or run enough untimed work to reach a steady state. Alternate candidate order so clock, thermal, and rank-skew effects are shared.

Included Examples

Adapt the operation, cases, work formula, and correctness tolerances rather than copying either example unchanged.

© guqiong96, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in .agents/skills/kernel-microbenchmark of guqiong96/Lvllm.

  • SKILL.md
  • agents/openai.yaml
  • benchmarks/cupti_microbenchmark.py
  • benchmarks/multi_gpu_gemm_rs.py

Open the folder on GitHubat commit 43ffc42

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in guqiong96/Lvllm, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Kernel Microbenchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kernel Microbenchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kernel Microbenchmark this skillguqiong96/Lvllm4642 repos~1.5kAutomated safety check: PassApache-2.0
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone
Production Add Diffusion Modelvllm-project/vllm-omni7.1k—~5.5kAutomated safety check: PassApache-2.0
Vllm Deploy Simplevllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Production Add Diffusion Model

    vllm-project/vllm-omni

    Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.

    7.1k GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    398 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from guqiong96/Lvllm

  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    464 GitHub starsUsed in 2 repos~349 tokens
    Auto-check passed
  • Triton Kernel Writing

    guqiong96/Lvllm

    Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

    464 GitHub starsUsed in 1 repo~831 tokens
    Auto-check passed

Works with

Questions about Kernel Microbenchmark

What does Kernel Microbenchmark do?

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…. Kernel Microbenchmark is an agent skill from guqiong96/Lvllm. Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL sanity checks.

When should I use Kernel Microbenchmark?

Kernel Microbenchmark fits situations like: tasks that involve LLM inference and serving.

How do I install Kernel Microbenchmark in Claude Code?

Run `npx skills add guqiong96/Lvllm --skill kernel-microbenchmark -a claude-code`. Or copy the skill folder (.agents/skills/kernel-microbenchmark in guqiong96/Lvllm) into .claude/skills/kernel-microbenchmark in your project. Claude Code loads it when a task matches its description.

How do I install Kernel Microbenchmark in Codex?

Run `npx skills add guqiong96/Lvllm --skill kernel-microbenchmark -a codex`. Or copy the skill folder (.agents/skills/kernel-microbenchmark in guqiong96/Lvllm) into .agents/skills/kernel-microbenchmark in your project. Codex loads it when a task matches its description.

Can I use Kernel Microbenchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guqiong96/Lvllm --skill kernel-microbenchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kernel-microbenchmark, .gemini/skills/kernel-microbenchmark, .github/skills/kernel-microbenchmark and .opencode/skills/kernel-microbenchmark in your project.

What does Kernel Microbenchmark need to run?

Going by SKILL.md and its folder, Kernel Microbenchmark needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Kernel Microbenchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Kernel Microbenchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kernel Microbenchmark use?

Kernel Microbenchmark is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kernel Microbenchmark use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kernel Microbenchmark?

Skills that share tags, products or a category with Kernel Microbenchmark: Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars) and Production Add Diffusion Model (vllm-project/vllm-omni, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kernel Microbenchmark?

guqiong96 (a GitHub user) maintains it in guqiong96/Lvllm, which has 464 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 22, 2026.

Source: guqiong96/Lvllm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.