Agent skill

Liger Kernel Perf

by linkedin in linkedin/Liger-Kernel

Optimizes the performance of existing Liger Kernel Triton kernels.

BSD-2-ClauseAuto-check passedAI & LLM Engineering

Install Liger Kernel Perf

skills CLI
$ npx skills add linkedin/Liger-Kernel --skill liger-kernel-perf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install linkedin/Liger-Kernel liger-kernel-perf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/linkedin/Liger-Kernel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/liger-kernel-perf .claude/skills/liger-kernel-perf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
liger-kernel-perf
GitHub stars
6.6k
Token cost
~1.5k tokens
SKILL.md length
639 words
Files
7
Skills in repo
3
Repo updated
First seen
Licence
BSD-2-Clause

At a glance

Optimizes the performance of existing Liger Kernel Triton kernels.

  • Works in 3 steps: Profile → Optimize → Finalize
  • A user asks to optimize
  • SKILL.md covers Mode Detection, Input Parsing, Pre-Flight Validation and Pipeline, plus 2 more sections
  • Calls ruff, pip and python

What it does

Liger Kernel Perf is an agent skill from linkedin/Liger-Kernel. Optimizes the performance of existing Liger Kernel Triton kernels. Profiles kernels, diagnoses bottlenecks (memory-bound vs compute-bound), generates multiple optimization variants with benchmarking, and applies the best variant while maintaining correctness. Supports GPU architecture-specific optimization (Ampere, Hopper, Blackwell). Use when a user asks to optimize, speed up, tune, profile, or reduce memory of an existing Liger kernel.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `finalizer.md`, `optimization-strategies.md` and `optimizer.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with Mistral AI. The repository describes itself as: Efficient Triton Kernels for LLM Training. The licence is BSD-2-Clause.

When your agent uses it

  • A user asks to optimize
  • Reduce memory of an existing Liger kernel

Example prompts

  • “Use the liger-kernel-perf skill to optimiz the performance of existing Liger Kernel Triton kernels”
  • “/liger-kernel-perf”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Profile
  2. Optimize
  3. Finalize

What it can do on your machine

Read from SKILL.md and the folder at commit d5f2817. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ruff
    • pip
    • python
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Liger Kernel Perf loads about 1.5k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 639 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from linkedin/Liger-Kernel at commit d5f2817, republished under its BSD-2-Clause licence (© linkedin). 639 words, ~1,536 tokens.

Download SKILL.mdSave it as .claude/skills/liger-kernel-perf/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
liger-kernel-perf
description
Optimizes the performance of existing Liger Kernel Triton kernels. Profiles kernels, diagnoses bottlenecks (memory-bound vs compute-bound), generates multiple optimization variants with benchmarking, and applies the best variant while maintaining correctness. Supports GPU architecture-specific optimization (Ampere, Hopper, Blackwell). Use when a user asks to optimize, speed up, tune, profile, or reduce memory of an existing Liger kernel.

Liger Kernel Perf

Optimizes existing Liger Kernel Triton kernels through a 3-stage pipeline: Profile, Optimize, Finalize. Supports interactive mode (human checkpoints between stages) and autonomous mode (runs end-to-end). NVIDIA GPUs only.

Mode Detection

  • Interactive mode (default): Human checkpoints between each stage
  • Autonomous mode: User says "just optimize it", "run without asking me", "optimize autonomously" → all stages run end-to-end, user sees only the final report

Input Parsing

Extract from the user's request:

FieldDescriptionDefault
target_kernelWhich kernel to optimize (e.g., "rms_norm", "cross_entropy")Required
optimization_goalspeed / memory / balancedbalanced
scopeSpecific pass (forward/backward), input regime, or generalgeneral
target_gpuAmpere / Hopper / Blackwell / auto-detectauto-detect
autonomyinteractive / autonomousinteractive
max_variantsMax optimization variants to try8
target_metricOptional concrete target (e.g., "forward under 0.3ms at hidden_size=4096")none

Pre-Flight Validation

Before starting the pipeline, validate:

  1. Kernel file exists: src/liger_kernel/ops/{kernel}.py
  2. Benchmark script exists: benchmark/scripts/benchmark_{kernel}.py
  3. Test file exists: test/transformers/test_{kernel}.py
  4. GPU is available and CUDA works
  5. Project is installed in dev mode (pip install -e ".[dev]")

If any validation fails, report clearly and stop.

Pipeline

Stage 1: Profile

Follow the Profiler workflow in profiler.md. If the host runtime supports parallel subagents, this stage may be delegated to one; otherwise execute the workflow directly.

This stage:

  1. Creates the workspace directory optimization/{kernel}/
  2. Copies the original kernel as a snapshot
  3. Runs baseline benchmarks using the existing benchmark script
  4. Detects GPU architecture (or uses user-specified target)
  5. Optionally runs NCU profiling (if ncu is available)
  6. Analyzes the kernel code (tier classification, patterns, optimization opportunities)
  7. Classifies the bottleneck: memory-bound vs compute-bound
  8. Produces an optimization profile with a recommended strategy order
  9. Saves profile to optimization/{kernel}/profile.md

Human checkpoint (interactive mode): Present the optimization profile with bottleneck diagnosis and proposed strategy order. Confirm before proceeding.

Stage 2: Optimize

Follow the Optimizer workflow in optimizer.md.

This stage runs an autonomous optimization loop:

  1. Read the optimization profile and original kernel
  2. Always try parameter tuning first (BLOCK_SIZE, num_warps, num_stages manual sweep -- NOT @triton.autotune)
  3. Then apply diagnosis-driven techniques from optimization-strategies.md
  4. For each variant: a. Generate the variant code → optimization/{kernel}/{kernel}_vN.py b. Write the variant lab notebook → optimization/{kernel}/{kernel}_vN_notes.md c. Run quick smoke test (single shape, float32, forward+backward) → discard on failure d. Run the full existing benchmark script → optimization/{kernel}/benchmarks/vN_results.csv e. Check guardrails (no catastrophic regressions) f. Update the variant notes with actual results
  5. Read all prior variant notes before generating the next variant
  6. Stop when: budget exhausted, 2 consecutive variants with <1% improvement, or target metric met
  7. Produce a comparison table of ALL variants

Human checkpoint (interactive mode): Present the comparison table across all variants. User approves the winner (or skill picks best if autonomous).

Show full SKILL.md (195 more words)Show less
Stage 3: Finalize

Follow the Finalizer workflow in finalizer.md.

This stage:

  1. Applies the winning variant in-place to src/liger_kernel/ops/{kernel}.py
  2. Runs the full test suite: python -m pytest test/transformers/test_{kernel}.py -xvs (hard gate)
  3. Runs checkstyle: make checkstyle (auto-fix with ruff check . --fix && ruff format .)
  4. Generates 3-way comparison plots (original liger vs optimized liger vs huggingface baseline) using benchmarks_visualizer.py
  5. Generates the final optimization report → optimization/{kernel}/report.md
  6. Creates a PR with only the kernel code changes (no plots or optimization workspace files)
  7. Presents the before/after summary with plots

Human checkpoint (interactive mode): Present the final report with before/after numbers, comparison plots, and test results.

Guardrails

These apply to EVERY variant, regardless of mode:

GuardrailThresholdAction
Non-target metric regression>5% worseReject variant
Cross-pass regression>10% on one pass to marginally improve otherReject variant
Smoke test failureAny correctness failureDiscard variant immediately
Full test suite failureAnyDo NOT apply winner, report failure, stop
Checkstyle failureAnyAuto-fix with ruff, retry once

Reference Files

© linkedin, BSD-2-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files in .agents/skills/liger-kernel-perf of linkedin/Liger-Kernel.

  • SKILL.md
  • finalizer.md
  • optimization-strategies.md
  • optimizer.md
  • profiler.md
  • templates/optimization-profile.md
  • templates/variant-notes.md

Open the folder on GitHubat commit d5f2817

Compare with similar skills

Liger Kernel Perf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Liger Kernel Perf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Liger Kernel Perf this skilllinkedin/Liger-Kernel6.6k—~1.5kAutomated safety check: PassBSD-2-Clause
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Fla Triton To Gluonfla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMIT
DGX Spark Memory and Thermal Opswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    212 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from linkedin/Liger-Kernel

  • Liger Autopatch

    linkedin/Liger-Kernel

    Adds Liger Kernel support for a new HuggingFace Transformers model, or modifies existing monkey-patching.

    6.6k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Liger Kernel Dev

    linkedin/Liger-Kernel

    Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.

    6.6k GitHub stars~799 tokensUpdated today
    Auto-check passed

Works with

Questions about Liger Kernel Perf

What does Liger Kernel Perf do?

Optimizes the performance of existing Liger Kernel Triton kernels. Liger Kernel Perf is an agent skill from linkedin/Liger-Kernel. Optimizes the performance of existing Liger Kernel Triton kernels.

When should I use Liger Kernel Perf?

Liger Kernel Perf fits situations like: A user asks to optimize; reduce memory of an existing Liger kernel.

How do I install Liger Kernel Perf in Claude Code?

Run `npx skills add linkedin/Liger-Kernel --skill liger-kernel-perf -a claude-code`. Or copy the skill folder (.agents/skills/liger-kernel-perf in linkedin/Liger-Kernel) into .claude/skills/liger-kernel-perf in your project. Claude Code loads it when a task matches its description.

How do I install Liger Kernel Perf in Codex?

Run `npx skills add linkedin/Liger-Kernel --skill liger-kernel-perf -a codex`. Or copy the skill folder (.agents/skills/liger-kernel-perf in linkedin/Liger-Kernel) into .agents/skills/liger-kernel-perf in your project. Codex loads it when a task matches its description.

Can I use Liger Kernel Perf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add linkedin/Liger-Kernel --skill liger-kernel-perf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/liger-kernel-perf, .gemini/skills/liger-kernel-perf, .github/skills/liger-kernel-perf and .opencode/skills/liger-kernel-perf in your project.

What does Liger Kernel Perf need to run?

Going by SKILL.md and its folder, Liger Kernel Perf needs the command-line tools its instructions call (ruff, pip, python and make). Our summary lists: Python 3.

Does Liger Kernel Perf access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Liger Kernel Perf safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Liger Kernel Perf use?

Liger Kernel Perf is published under the BSD-2-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Liger Kernel Perf use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Liger Kernel Perf?

Skills that share tags, products or a category with Liger Kernel Perf: Hugging Face LLM Trainer (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars) and DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Liger Kernel Perf?

linkedin (a GitHub organization) maintains it in linkedin/Liger-Kernel, which has 6,649 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: linkedin/Liger-Kernel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.