Hugging Face Local Model Evals
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.
$ npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install guqiong96/Lvllm triton-kernel-writing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/triton-kernel-writing .claude/skills/triton-kernel-writing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "triton-kernel-writing" agent skill from https://github.com/guqiong96/Lvllm/tree/main/.agents/skills/triton-kernel-writing into .claude/skills/triton-kernel-writing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-kernel-writing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/guqiong96/Lvllm/tree/main/.agents/skills/triton-kernel-writingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install guqiong96/Lvllm triton-kernel-writing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/triton-kernel-writing .agents/skills/triton-kernel-writing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "triton-kernel-writing" agent skill from https://github.com/guqiong96/Lvllm/tree/main/.agents/skills/triton-kernel-writing into .agents/skills/triton-kernel-writing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-kernel-writing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install guqiong96/Lvllm triton-kernel-writing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/triton-kernel-writing .cursor/skills/triton-kernel-writing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "triton-kernel-writing" agent skill from https://github.com/guqiong96/Lvllm/tree/main/.agents/skills/triton-kernel-writing into .cursor/skills/triton-kernel-writing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-kernel-writing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/guqiong96/Lvllm.git --path .agents/skills/triton-kernel-writing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install guqiong96/Lvllm triton-kernel-writing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/triton-kernel-writing .gemini/skills/triton-kernel-writing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "triton-kernel-writing" agent skill from https://github.com/guqiong96/Lvllm/tree/main/.agents/skills/triton-kernel-writing into .gemini/skills/triton-kernel-writing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-kernel-writing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install guqiong96/Lvllm triton-kernel-writingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/triton-kernel-writing .github/skills/triton-kernel-writing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "triton-kernel-writing" agent skill from https://github.com/guqiong96/Lvllm/tree/main/.agents/skills/triton-kernel-writing into .github/skills/triton-kernel-writing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-kernel-writing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install guqiong96/Lvllm triton-kernel-writing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guqiong96/Lvllm.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/triton-kernel-writing .opencode/skills/triton-kernel-writing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "triton-kernel-writing" agent skill from https://github.com/guqiong96/Lvllm/tree/main/.agents/skills/triton-kernel-writing into .opencode/skills/triton-kernel-writing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "triton-kernel-writing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
triton-kernel-writingWrite or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.
Triton Kernel Writing is an agent skill from guqiong96/Lvllm. Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.
Its SKILL.md is about 830 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing and LLM inference and serving. It works with vLLM. The repository describes itself as: LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 43ffc42. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
triton-lang.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Triton Kernel Writing loads about 831 tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 412 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from guqiong96/Lvllm at commit 43ffc42, republished under its Apache-2.0 licence (© guqiong96). 412 words, ~831 tokens.
.claude/skills/triton-kernel-writing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.torch.compile as a possible
implementation to inspect. Print Inductor's generated code with
TORCH_LOGS="output_code" .venv/bin/python <script> or enable
torch._logging.set_logs(output_code=True) before the compiled function
runs. Treat generated code as a reference, not as proof of correctness or
optimality.BLOCK_SIZE, or use
a small, legible heuristic when workloads need different choices. Use
triton.autotune only when tuning is critical to performance, such as for a
matrix multiplication. Otherwise prioritize simple code and fast startup.do_not_specialize, especially those
that may alternate between values such as 0 and 1, which can produce
different specialization keys.tl.debug_barrier() between the write and read. The barrier
synchronizes threads in the block; it does not synchronize separate program
instances.grid[1] and grid[2] must be at most 65,535. Choose or flatten the grid
order so those dimensions cannot exceed the limit for supported shapes.
For example, num_tokens is commonly 8K or 16K, but users may configure 32K
or more. If num_tokens is a grid dimension, it is safe to put it in
grid[0] (or tile it).int64 for offset arithmetic when an index can exceed 32-bit range,
especially for KV-cache addressing. Cast operands before multiplication or
addition so an intermediate does not overflow in 32-bit arithmetic.[num_tokens, num_heads] grid can be a good low-latency mapping for decode,
but it can be very slow for prefill. If the kernel serves prefill, consider
tiling tokens or otherwise increasing the work and locality per program.$kernel-microbenchmark for benchmark construction, measurement, and
interpretation.num_tokens covering decode and representative prefill
workloads. Include relevant head counts and dimensions when they affect the
launch shape, and do not select an implementation or tuning heuristic from a
single setup.© guqiong96, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .agents/skills/triton-kernel-writing of guqiong96/Lvllm.
Open the folder on GitHubat commit 43ffc42
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in guqiong96/Lvllm, which our catalogue first saw on October 7, 2026.
Triton Kernel Writing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Triton Kernel Writing this skillguqiong96/Lvllm | 465 | 1 repos | ~831 | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| LLM Serving Capacity PlannerBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.5k | Automated safety check: Pass | None | |
| Ascend Model Adapter for vLLMvllm-project/vllm-ascend | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.8k | Automated safety check: Pass | None |
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
BBuf/AI-Infra-Auto-Driven-SKILLS
Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.
vllm-project/vllm-ascend
Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
BBuf/AI-Infra-Auto-Driven-SKILLS
Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.
guqiong96/Lvllm
Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.
guqiong96/Lvllm
Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…
Works with
Categories
Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. Triton Kernel Writing is an agent skill from guqiong96/Lvllm. Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.
Triton Kernel Writing fits situations like: tasks that involve GPU and accelerator computing; tasks that involve LLM inference and serving.
Run `npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a claude-code`. Or copy the skill folder (.agents/skills/triton-kernel-writing in guqiong96/Lvllm) into .claude/skills/triton-kernel-writing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a codex`. Or copy the skill folder (.agents/skills/triton-kernel-writing in guqiong96/Lvllm) into .agents/skills/triton-kernel-writing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guqiong96/Lvllm --skill triton-kernel-writing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/triton-kernel-writing, .gemini/skills/triton-kernel-writing, .github/skills/triton-kernel-writing and .opencode/skills/triton-kernel-writing in your project.
Going by SKILL.md and its folder, Triton Kernel Writing needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: triton-lang.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Triton Kernel Writing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 831 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Triton Kernel Writing: Hugging Face Local Model Evals (huggingface/skills, 11k stars), LLM Serving Capacity Planner (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Ascend Model Adapter for vLLM (vllm-project/vllm-ascend, 2.9k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
guqiong96 (a GitHub user) maintains it in guqiong96/Lvllm, which has 465 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 22, 2026.
Source: guqiong96/Lvllm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.