Kernel Microbenchmark is an agent skill from guqiong96/Lvllm. Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL sanity checks.
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `agents/openai.yaml`, `benchmarks/cupti_microbenchmark.py` and `benchmarks/multi_gpu_gemm_rs.py`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and CUDA. The repository describes itself as: LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA… The licence is Apache-2.0.