Agent skill

Tune Triton Ascend Kernels

by Krusty84 in Krusty84/triton-ascend-agent-dev-kit

Add standard or advanced autotuning to Triton-Ascend kernels, define shape keys and candidate meta-parameters, and structure kernels so automatic split and tiling analysis can recognize their axes.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Tune Triton Ascend Kernels

skills CLI
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill tune-triton-ascend-kernels -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Krusty84/triton-ascend-agent-dev-kit tune-triton-ascend-kernels --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tune-triton-ascend-kernels .claude/skills/tune-triton-ascend-kernels && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tune-triton-ascend-kernels
GitHub stars
106
Token cost
~813 tokens
SKILL.md length
257 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add standard or advanced autotuning to Triton-Ascend kernels, define shape keys and candidate meta-parameters, and structure kernels so automatic split and tiling analysis can recognize their axes.

  • Works in 6 steps: Establish a correct untuned kernel and… → Choose standard autotune when candidate… → Choose advanced autotune when… → …
  • An agent needs to benchmark block-size
  • SKILL.md covers Goal, Workflow, Implementation Pattern and Ascend Guardrails, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tune Triton Ascend Kernels is an agent skill from Krusty84/triton-ascend-agent-dev-kit. Add standard or advanced autotuning to Triton-Ascend kernels, define shape keys and candidate meta-parameters, and structure kernels so automatic split and tiling analysis can recognize their axes. Use when an agent needs to benchmark block-size or multibuffer choices, use configs=[] for generated vector-kernel candidates, or diagnose why Triton-Ascend autotune cannot infer a tunable axis.

Its SKILL.md is about 810 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Accessibility. The repository describes itself as: A toolkit for AI agents used for development on Triton-Ascend for Ascend NPU. The licence is Apache-2.0.

When your agent uses it

  • An agent needs to benchmark block-size
  • Multibuffer choices
  • Use configs=[] for generated vector-kernel candidates
  • Diagnose why Triton-Ascend autotune cannot infer a tunable axis

Example prompts

  • “/tune-triton-ascend-kernels”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Establish a correct untuned kernel and reference test.
  2. Choose standard autotune when candidate configurations are known.
  3. Choose advanced autotune when Triton-Ascend should infer vector split or tiling candidates.
  4. Put only values that can change the winner in key.
  5. Derive the launch grid from meta-parameters and omit tunable meta-parameters at launch.
  6. Validate the tuned result, then benchmark representative shapes.

What it can do on your machine

Read from SKILL.md and the folder at commit 4ab5ee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tune Triton Ascend Kernels loads about 813 tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 257 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~813

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Krusty84/triton-ascend-agent-dev-kit at commit 4ab5ee7, republished under its Apache-2.0 licence (© Krusty84). 257 words, ~813 tokens.

Download SKILL.mdSave it as .claude/skills/tune-triton-ascend-kernels/SKILL.md (or your agent's skills folder).
name
tune-triton-ascend-kernels
description
Add standard or advanced autotuning to Triton-Ascend kernels, define shape keys and candidate meta-parameters, and structure kernels so automatic split and tiling analysis can recognize their axes. Use when an agent needs to benchmark block-size or multibuffer choices, use configs=[] for generated vector-kernel candidates, or diagnose why Triton-Ascend autotune cannot infer a tunable axis.

Tune Triton-Ascend Kernels

Goal

Select a correct kernel configuration for each relevant input shape without hardcoding one block size for all workloads.

Workflow

  1. Establish a correct untuned kernel and reference test.
  2. Choose standard autotune when candidate configurations are known.
  3. Choose advanced autotune when Triton-Ascend should infer vector split or tiling candidates.
  4. Put only values that can change the winner in key.
  5. Derive the launch grid from meta-parameters and omit tunable meta-parameters at launch.
  6. Validate the tuned result, then benchmark representative shapes.

Implementation Pattern

Use explicit candidates for standard autotune:

python
@triton.autotune(
    configs=[
        triton.Config({"XS": 128, "multibuffer": True}),
        triton.Config({"XS": 8192, "multibuffer": False}),
    ],
    key=["numel"],
)
@triton.jit
def kernel(out_ptr, x_ptr, numel, XS: tl.constexpr):
    offsets = tl.program_id(0) * XS + tl.arange(0, XS)
    mask = offsets < numel
    values = tl.load(x_ptr + offsets, mask=mask)
    tl.store(out_ptr + offsets, values, mask=mask)

Use generated candidates for advanced vector autotune:

python
@triton.autotune(
    configs=[],
    key={"x": "n_elements"},
    split_params={"x": "BLOCK_SIZE"},
    tiling_params={},
    low_dims=["x"],
    persistent_reduction=False,
    dual_reduction=False,
)
@triton.jit
def kernel(out_ptr, x_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
    offsets = tl.program_id(0) * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
    mask = offsets < n_elements
    values = tl.load(x_ptr + offsets, mask=mask)
    tl.store(out_ptr + offsets, values, mask=mask)

Ascend Guardrails

  • Standard Triton-Ascend autotune supports block-size and multibuffer choices; do not assume community num_warps or num_stages tuning maps to Ascend hardware.
  • Treat configs=[] as generated-candidate mode only when split_params or tiling_params identifies at least one axis.
  • Make a split parameter multiply tl.program_id and compare its resulting offsets with the matching key in a mask.
  • Make a tiling parameter participate in tl.arange and a loop range, then compare its offsets with the matching key.
  • Candidate meta-parameters must remain omitted from the kernel launch; explicitly passed values are excluded from automatic inference.
  • Keep advanced automatic tiling to supported vector kernels; do not apply it to cube operators.
  • Set TRITON_BENCH_METHOD="npu" only when profiler timing is needed for very fast kernels; expect substantially longer tuning.

Verification

Compare the selected kernel with the PyTorch reference for every key shape and inspect that retuning happens when the key changes. Benchmark the winner separately; autotune success does not replace correctness tests.

© Krusty84, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tune-triton-ascend-kernels of Krusty84/triton-ascend-agent-dev-kit.

Open the folder on GitHubat commit 4ab5ee7

Compare with similar skills

Tune Triton Ascend Kernels next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tune Triton Ascend Kernels compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tune Triton Ascend Kernels this skillKrusty84/triton-ascend-agent-dev-kit106—~813Automated safety check: PassApache-2.0
Align Urdf Mjcfi2rt-robotics/i2rt168—~4.4kAutomated safety check: PassMIT
Streaming Responsethedaviddias/ux-patterns-for-developers259—~1kAutomated safety check: PassCustom licence
Color Pickerthedaviddias/ux-patterns-for-developers259—~1.3kAutomated safety check: PassCustom licence
Date Pickerthedaviddias/ux-patterns-for-developers259—~1.5kAutomated safety check: PassCustom licence
Hambuger Menuthedaviddias/ux-patterns-for-developers259—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Align Urdf Mjcf

    i2rt-robotics/i2rt

    Align a MuJoCo MJCF robot model with a URDF source while preserving kinematics, geometry, masses, centers of mass, and inertia tensors.

    168 GitHub stars~4.4k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Streaming Response

    thedaviddias/ux-patterns-for-developers

    A skill your agent uses when implementing real-time AI response streaming.

    259 GitHub stars~1k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Color Picker

    thedaviddias/ux-patterns-for-developers

    A skill your agent uses when implementing select colors with visual feedback.

    259 GitHub stars~1.3k tokensUpdated 5 days ago
    Frontend & DesignAuto-check passed
  • Date Picker

    thedaviddias/ux-patterns-for-developers

    A skill your agent uses when implementing select dates from a calendar interface.

    259 GitHub stars~1.5k tokensUpdated 5 days ago
    Frontend & DesignAuto-check passed
  • Hambuger Menu

    thedaviddias/ux-patterns-for-developers

    A skill your agent uses when you need to display a menu icon for mobile devices.

    259 GitHub stars~1.2k tokensUpdated 5 days ago
    Frontend & DesignAuto-check passed
  • Signup

    thedaviddias/ux-patterns-for-developers

    A skill your agent uses when implementing user registration and account creation.

    259 GitHub stars~1.2k tokensUpdated 5 days ago
    Frontend & DesignAuto-check passed

More from Krusty84/triton-ascend-agent-dev-kit

All 23 skills in this repo
  • Triton-Ascend Token Pool Assignment

    Krusty84/triton-ascend-agent-dev-kit

    Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

    106 GitHub stars~564 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Batch Token Reorder

    Krusty84/triton-ascend-agent-dev-kit

    Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

    106 GitHub stars~597 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Binned MoE Routing

    Krusty84/triton-ascend-agent-dev-kit

    Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

    106 GitHub stars~662 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Attention

    Krusty84/triton-ascend-agent-dev-kit

    Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

    106 GitHub stars~753 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Softmax

    Krusty84/triton-ascend-agent-dev-kit

    Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

    106 GitHub stars~636 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Layer Norm

    Krusty84/triton-ascend-agent-dev-kit

    Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles.

    106 GitHub stars~659 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Tune Triton Ascend Kernels

What does Tune Triton Ascend Kernels do?

Add standard or advanced autotuning to Triton-Ascend kernels, define shape keys and candidate meta-parameters, and structure kernels so automatic split and tiling analysis can recognize their axes. Tune Triton Ascend Kernels is an agent skill from Krusty84/triton-ascend-agent-dev-kit. Add standard or advanced autotuning to Triton-Ascend kernels, define shape keys and candidate meta-parameters, and structure kernels so automatic split and tiling analysis can recognize their axes.

When should I use Tune Triton Ascend Kernels?

Tune Triton Ascend Kernels fits situations like: an agent needs to benchmark block-size; multibuffer choices; use configs=[] for generated vector-kernel candidates; diagnose why Triton-Ascend autotune cannot infer a tunable axis.

How do I install Tune Triton Ascend Kernels in Claude Code?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill tune-triton-ascend-kernels -a claude-code`. Or copy the skill folder (skills/tune-triton-ascend-kernels in Krusty84/triton-ascend-agent-dev-kit) into .claude/skills/tune-triton-ascend-kernels in your project. Claude Code loads it when a task matches its description.

How do I install Tune Triton Ascend Kernels in Codex?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill tune-triton-ascend-kernels -a codex`. Or copy the skill folder (skills/tune-triton-ascend-kernels in Krusty84/triton-ascend-agent-dev-kit) into .agents/skills/tune-triton-ascend-kernels in your project. Codex loads it when a task matches its description.

Can I use Tune Triton Ascend Kernels in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill tune-triton-ascend-kernels -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tune-triton-ascend-kernels, .gemini/skills/tune-triton-ascend-kernels, .github/skills/tune-triton-ascend-kernels and .opencode/skills/tune-triton-ascend-kernels in your project.

What does Tune Triton Ascend Kernels need to run?

SKILL.md names no scripts, command-line tools or credentials: Tune Triton Ascend Kernels is instructions for the agent only. Our summary lists: Python 3.

Does Tune Triton Ascend Kernels access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tune Triton Ascend Kernels safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tune Triton Ascend Kernels use?

Tune Triton Ascend Kernels is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tune Triton Ascend Kernels use?

About 813 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tune Triton Ascend Kernels?

Skills that share tags, products or a category with Tune Triton Ascend Kernels: Align Urdf Mjcf (i2rt-robotics/i2rt, 168 stars), Streaming Response (thedaviddias/ux-patterns-for-developers, 259 stars), Color Picker (thedaviddias/ux-patterns-for-developers, 259 stars) and Date Picker (thedaviddias/ux-patterns-for-developers, 259 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tune Triton Ascend Kernels?

Krusty84 (a GitHub user) maintains it in Krusty84/triton-ascend-agent-dev-kit, which has 106 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 15, 2026.

Source: Krusty84/triton-ascend-agent-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.