Agent skill

Triton-Ascend Fused Attention

by Krusty84 in Krusty84/triton-ascend-agent-dev-kit

Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Triton-Ascend Fused Attention

skills CLI
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-attention --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/build-triton-ascend-fused-attention .claude/skills/build-triton-ascend-fused-attention && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
build-triton-ascend-fused-attention
GitHub stars
106
Token cost
~753 tokens
SKILL.md length
328 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

  • Works in 6 steps: Validate matching Q, K, and V shapes,… → Choose BLOCK_M for query rows and… → Map each task to a batch, head, and… → …
  • Writing a forward scaled dot-product attention kernel for Ascend NPU
  • SKILL.md covers Goal, Workflow, Implementation Pattern and Ascend Guardrails, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill describes computing softmax(QK^T * scale)V for BNSD tensors without materializing the full attention matrix. The workflow validates shapes, strides, dtype and supported head dimensions, picks BLOCK_M and BLOCK_N, maps each task to a batch, head and query-block tuple, then streams K and V blocks through an online-softmax loop. Causal attention handles preceding blocks and the diagonal block as separate stages, applying the triangular mask only on the diagonal, and the output is normalized by the running denominator.

Ascend-specific guardrails include taking base offsets from each tensor's own strides, building tl.make_block_ptr pointers that match the BNSD layout, keeping maxima, denominators and accumulators in float32, and using a global float32 scratch accumulator with CANN extract_slice and insert_slice when HEAD_DIM is 256 and the local accumulator would exceed UB capacity. Verification compares against torch_npu.npu_fusion_attention across causal modes, head dimensions, block sizes, float16 and bfloat16, and long contexts.

When your agent uses it

  • Writing a forward scaled dot-product attention kernel for Ascend NPU
  • Supporting causal and non-causal masking in a Triton-Ascend kernel
  • Handling large head dimensions without overflowing on-chip memory
  • Checking a custom attention kernel against torch_npu.npu_fusion_attention

Example prompts

  • “Write a causal fused attention forward kernel in Triton-Ascend for BNSD tensors with head dimension 128.”
  • “My attention kernel runs out of on-chip memory at head dimension 256, so apply the scratch accumulator approach.”
  • “Add a test that compares our kernel with torch_npu.npu_fusion_attention in float16 and bfloat16.”
  • “Choose BLOCK_M and BLOCK_N values that divide the sequence length evenly.”

Requirements

  • Triton-Ascend on an Ascend NPU
  • `torch_npu` for the reference comparison

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Validate matching Q, K, and V shapes, compatible strides, supported dtype, and HEAD_DIM in {16, 32, 64, 128, 256}.
  2. Choose BLOCK_M for query rows and BLOCK_N for key/value rows. Require N_CTX to be divisible by both in the compact implementation.
  3. Map each task to a batch, head, and query-block tuple and create block pointers for Q, K, V, and output.
  4. Load one Q block and stream K/V blocks through an online-softmax inner loop.
  5. For causal attention, process preceding blocks and the diagonal block as separate stages; apply the triangular mask only to the diagonal…
  6. Normalize the accumulated output by the running denominator and store it in the output dtype.

What it can do on your machine

Read from SKILL.md and the folder at commit 4ab5ee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton-Ascend Fused Attention loads about 753 tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 328 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~753

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Krusty84/triton-ascend-agent-dev-kit at commit 4ab5ee7, republished under its Apache-2.0 licence (© Krusty84). 328 words, ~753 tokens.

Download SKILL.mdSave it as .claude/skills/build-triton-ascend-fused-attention/SKILL.md (or your agent's skills folder).
name
build-triton-ascend-fused-attention
description
Build and adapt a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend using tiled Q/K/V block pointers, online softmax, causal staging, and float32 accumulation. Use when an agent needs forward scaled dot-product attention in BNSD layout on Ascend NPU, must support causal or non-causal masking, or must handle large head dimensions without overflowing on-chip memory.

Build Triton-Ascend Fused Attention

Goal

Compute softmax(QKᵀ * scale)V without materializing the full attention matrix. Support BNSD tensors shaped Z by H by N_CTX by HEAD_DIM.

Workflow

  1. Validate matching Q, K, and V shapes, compatible strides, supported dtype, and HEAD_DIM in {16, 32, 64, 128, 256}.
  2. Choose BLOCK_M for query rows and BLOCK_N for key/value rows. Require N_CTX to be divisible by both in the compact implementation.
  3. Map each task to a batch, head, and query-block tuple and create block pointers for Q, K, V, and output.
  4. Load one Q block and stream K/V blocks through an online-softmax inner loop.
  5. For causal attention, process preceding blocks and the diagonal block as separate stages; apply the triangular mask only to the diagonal stage.
  6. Normalize the accumulated output by the running denominator and store it in the output dtype.

Implementation Pattern

For every query row, initialize m_i=-inf, l_i=1, and a float32 accumulator. For each K/V tile:

text
scores = dot(q, transpose(k)) * scale
scores = scores + causal_mask_if_diagonal
m_new = max(m_i, row_max(scores))
p = exp(scores - m_new)
alpha = exp(m_i - m_new)
l_i = l_i * alpha + row_sum(p)
acc = acc * alpha + dot(cast(p, k.dtype), v)
m_i = m_new
output = acc / l_i

Use a stage bitmask compatible with STAGE=3 for causal attention and STAGE=1 for non-causal attention. Advance K and V block pointers by BLOCK_N after each iteration.

Ascend Guardrails

  • Compute base offsets from each tensor's own strides; do not assume Q strides also describe K, V, or output.
  • Use tl.make_block_ptr with shape, strides, offsets, block_shape, and order consistent with BNSD layout.
  • Keep score maxima, denominators, and output accumulators in float32.
  • For HEAD_DIM=256, use a global float32 scratch accumulator and CANN extract_slice/insert_slice updates when a full local accumulator exceeds UB capacity.
  • Treat the example core count of 20 as target-specific; confirm the available NPU execution resources before fixing a persistent grid size.
  • Import triton.language.extra.cann.extension only when the large-head scratch path needs it.
  • Expose this as forward-only unless a real backward kernel is implemented.

Verification

Compare with torch_npu.npu_fusion_attention using input_layout="BNSD", identical scale, and a matching causal mask. Test both causal modes, supported head dimensions, multiple BLOCK_M/BLOCK_N pairs, float16 and bfloat16, and long contexts. Start with atol=rtol=1e-2 and tighten when evidence supports it.

© Krusty84, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/build-triton-ascend-fused-attention of Krusty84/triton-ascend-agent-dev-kit.

Open the folder on GitHubat commit 4ab5ee7

Compare with similar skills

Triton-Ascend Fused Attention next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton-Ascend Fused Attention compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton-Ascend Fused Attention this skillKrusty84/triton-ascend-agent-dev-kit106—~753Automated safety check: PassApache-2.0
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Liger Kernel Devlinkedin/Liger-Kernel6.7k—~799Automated safety check: PassBSD-2-Clause
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence

Similar skills

  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Liger Kernel Dev

    linkedin/Liger-Kernel

    Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.

    6.7k GitHub stars~799 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from Krusty84/triton-ascend-agent-dev-kit

All 23 skills in this repo
  • Triton-Ascend Token Pool Assignment

    Krusty84/triton-ascend-agent-dev-kit

    Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

    106 GitHub stars~564 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Batch Token Reorder

    Krusty84/triton-ascend-agent-dev-kit

    Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

    106 GitHub stars~597 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Binned MoE Routing

    Krusty84/triton-ascend-agent-dev-kit

    Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

    106 GitHub stars~662 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Softmax

    Krusty84/triton-ascend-agent-dev-kit

    Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

    106 GitHub stars~636 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Layer Norm

    Krusty84/triton-ascend-agent-dev-kit

    Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles.

    106 GitHub stars~659 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Moe Gather Scatter

    Krusty84/triton-ascend-agent-dev-kit

    Build Ascend-friendly MoE gather, scatter, and router-weight-gradient kernels with vector-core-sized grids, UB-aware row/column tiling, and CANN slice extensions.

    106 GitHub stars~696 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Triton-Ascend Fused Attention

What does Triton-Ascend Fused Attention do?

Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation. The skill describes computing softmax(QK^T * scale)V for BNSD tensors without materializing the full attention matrix. The workflow validates shapes, strides, dtype and supported head dimensions, picks BLOCK_M and BLOCK_N, maps each task to a batch, head and query-block tuple, then streams K and V blocks through an online-softmax loop.

When should I use Triton-Ascend Fused Attention?

Triton-Ascend Fused Attention fits situations like: writing a forward scaled dot-product attention kernel for Ascend NPU; supporting causal and non-causal masking in a Triton-Ascend kernel; handling large head dimensions without overflowing on-chip memory; checking a custom attention kernel against torch_npu.npu_fusion_attention.

How do I install Triton-Ascend Fused Attention in Claude Code?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a claude-code`. Or copy the skill folder (skills/build-triton-ascend-fused-attention in Krusty84/triton-ascend-agent-dev-kit) into .claude/skills/build-triton-ascend-fused-attention in your project. Claude Code loads it when a task matches its description.

How do I install Triton-Ascend Fused Attention in Codex?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a codex`. Or copy the skill folder (skills/build-triton-ascend-fused-attention in Krusty84/triton-ascend-agent-dev-kit) into .agents/skills/build-triton-ascend-fused-attention in your project. Codex loads it when a task matches its description.

Can I use Triton-Ascend Fused Attention in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-triton-ascend-fused-attention, .gemini/skills/build-triton-ascend-fused-attention, .github/skills/build-triton-ascend-fused-attention and .opencode/skills/build-triton-ascend-fused-attention in your project.

What does Triton-Ascend Fused Attention need to run?

SKILL.md names no scripts, command-line tools or credentials: Triton-Ascend Fused Attention is instructions for the agent only. Our summary lists: Triton-Ascend on an Ascend NPU; `torch_npu` for the reference comparison.

Does Triton-Ascend Fused Attention access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton-Ascend Fused Attention safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton-Ascend Fused Attention use?

Triton-Ascend Fused Attention is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton-Ascend Fused Attention use?

About 753 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton-Ascend Fused Attention?

Skills that share tags, products or a category with Triton-Ascend Fused Attention: Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Cuda Index Width (pytorch/pytorch, 104k stars), Liger Kernel Dev (linkedin/Liger-Kernel, 6.7k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton-Ascend Fused Attention?

Krusty84 (a GitHub user) maintains it in Krusty84/triton-ascend-agent-dev-kit, which has 106 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 15, 2026.

Source: Krusty84/triton-ascend-agent-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.