Paddle Design Compiler
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-attention --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/build-triton-ascend-fused-attention .claude/skills/build-triton-ascend-fused-attention && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "build-triton-ascend-fused-attention" agent skill from https://github.com/Krusty84/triton-ascend-agent-dev-kit/tree/main/skills/build-triton-ascend-fused-attention into .claude/skills/build-triton-ascend-fused-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-triton-ascend-fused-attention", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Krusty84/triton-ascend-agent-dev-kit/tree/main/skills/build-triton-ascend-fused-attentionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-attention --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/build-triton-ascend-fused-attention .agents/skills/build-triton-ascend-fused-attention && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "build-triton-ascend-fused-attention" agent skill from https://github.com/Krusty84/triton-ascend-agent-dev-kit/tree/main/skills/build-triton-ascend-fused-attention into .agents/skills/build-triton-ascend-fused-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-triton-ascend-fused-attention", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-attention --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/build-triton-ascend-fused-attention .cursor/skills/build-triton-ascend-fused-attention && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "build-triton-ascend-fused-attention" agent skill from https://github.com/Krusty84/triton-ascend-agent-dev-kit/tree/main/skills/build-triton-ascend-fused-attention into .cursor/skills/build-triton-ascend-fused-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-triton-ascend-fused-attention", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Krusty84/triton-ascend-agent-dev-kit.git --path skills/build-triton-ascend-fused-attention--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-attention --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/build-triton-ascend-fused-attention .gemini/skills/build-triton-ascend-fused-attention && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "build-triton-ascend-fused-attention" agent skill from https://github.com/Krusty84/triton-ascend-agent-dev-kit/tree/main/skills/build-triton-ascend-fused-attention into .gemini/skills/build-triton-ascend-fused-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-triton-ascend-fused-attention", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-attentionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/build-triton-ascend-fused-attention .github/skills/build-triton-ascend-fused-attention && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "build-triton-ascend-fused-attention" agent skill from https://github.com/Krusty84/triton-ascend-agent-dev-kit/tree/main/skills/build-triton-ascend-fused-attention into .github/skills/build-triton-ascend-fused-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-triton-ascend-fused-attention", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-attention --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/build-triton-ascend-fused-attention .opencode/skills/build-triton-ascend-fused-attention && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "build-triton-ascend-fused-attention" agent skill from https://github.com/Krusty84/triton-ascend-agent-dev-kit/tree/main/skills/build-triton-ascend-fused-attention into .opencode/skills/build-triton-ascend-fused-attention/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "build-triton-ascend-fused-attention", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
build-triton-ascend-fused-attentionGuides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.
The skill describes computing softmax(QK^T * scale)V for BNSD tensors without materializing the full attention matrix. The workflow validates shapes, strides, dtype and supported head dimensions, picks BLOCK_M and BLOCK_N, maps each task to a batch, head and query-block tuple, then streams K and V blocks through an online-softmax loop. Causal attention handles preceding blocks and the diagonal block as separate stages, applying the triangular mask only on the diagonal, and the output is normalized by the running denominator.
Ascend-specific guardrails include taking base offsets from each tensor's own strides, building tl.make_block_ptr pointers that match the BNSD layout, keeping maxima, denominators and accumulators in float32, and using a global float32 scratch accumulator with CANN extract_slice and insert_slice when HEAD_DIM is 256 and the local accumulator would exceed UB capacity. Verification compares against torch_npu.npu_fusion_attention across causal modes, head dimensions, block sizes, float16 and bfloat16, and long contexts.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4ab5ee7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Triton-Ascend Fused Attention loads about 753 tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 328 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Krusty84/triton-ascend-agent-dev-kit at commit 4ab5ee7, republished under its Apache-2.0 licence (© Krusty84). 328 words, ~753 tokens.
.claude/skills/build-triton-ascend-fused-attention/SKILL.md (or your agent's skills folder).Compute softmax(QKᵀ * scale)V without materializing the full attention matrix. Support BNSD tensors shaped Z by H by N_CTX by HEAD_DIM.
For every query row, initialize m_i=-inf, l_i=1, and a float32 accumulator. For each K/V tile:
scores = dot(q, transpose(k)) * scale
scores = scores + causal_mask_if_diagonal
m_new = max(m_i, row_max(scores))
p = exp(scores - m_new)
alpha = exp(m_i - m_new)
l_i = l_i * alpha + row_sum(p)
acc = acc * alpha + dot(cast(p, k.dtype), v)
m_i = m_new
output = acc / l_iUse a stage bitmask compatible with STAGE=3 for causal attention and STAGE=1 for non-causal attention. Advance K and V block pointers by BLOCK_N after each iteration.
Compare with torch_npu.npu_fusion_attention using input_layout="BNSD", identical scale, and a matching causal mask. Test both causal modes, supported head dimensions, multiple BLOCK_M/BLOCK_N pairs, float16 and bfloat16, and long contexts. Start with atol=rtol=1e-2 and tighten when evidence supports it.
© Krusty84, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/build-triton-ascend-fused-attention of Krusty84/triton-ascend-agent-dev-kit.
Open the folder on GitHubat commit 4ab5ee7
Triton-Ascend Fused Attention next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Triton-Ascend Fused Attention this skillKrusty84/triton-ascend-agent-dev-kit | 106 | — | ~753 | Automated safety check: Pass | Apache-2.0 | |
| Paddle Design CompilerPaddlePaddle/Paddle | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Liger Kernel Devlinkedin/Liger-Kernel | 6.7k | — | ~799 | Automated safety check: Pass | BSD-2-Clause | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| Metal Kernelpytorch/pytorch | 104k | — | ~4.9k | Automated safety check: Pass | Custom licence |
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
linkedin/Liger-Kernel
Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
pytorch/pytorch
Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
Krusty84/triton-ascend-agent-dev-kit
Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.
Krusty84/triton-ascend-agent-dev-kit
Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.
Krusty84/triton-ascend-agent-dev-kit
Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.
Krusty84/triton-ascend-agent-dev-kit
Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.
Krusty84/triton-ascend-agent-dev-kit
Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles.
Krusty84/triton-ascend-agent-dev-kit
Build Ascend-friendly MoE gather, scatter, and router-weight-gradient kernels with vector-core-sized grids, UB-aware row/column tiling, and CANN slice extensions.
Categories
Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation. The skill describes computing softmax(QK^T * scale)V for BNSD tensors without materializing the full attention matrix. The workflow validates shapes, strides, dtype and supported head dimensions, picks BLOCK_M and BLOCK_N, maps each task to a batch, head and query-block tuple, then streams K and V blocks through an online-softmax loop.
Triton-Ascend Fused Attention fits situations like: writing a forward scaled dot-product attention kernel for Ascend NPU; supporting causal and non-causal masking in a Triton-Ascend kernel; handling large head dimensions without overflowing on-chip memory; checking a custom attention kernel against torch_npu.npu_fusion_attention.
Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a claude-code`. Or copy the skill folder (skills/build-triton-ascend-fused-attention in Krusty84/triton-ascend-agent-dev-kit) into .claude/skills/build-triton-ascend-fused-attention in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a codex`. Or copy the skill folder (skills/build-triton-ascend-fused-attention in Krusty84/triton-ascend-agent-dev-kit) into .agents/skills/build-triton-ascend-fused-attention in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-attention -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-triton-ascend-fused-attention, .gemini/skills/build-triton-ascend-fused-attention, .github/skills/build-triton-ascend-fused-attention and .opencode/skills/build-triton-ascend-fused-attention in your project.
SKILL.md names no scripts, command-line tools or credentials: Triton-Ascend Fused Attention is instructions for the agent only. Our summary lists: Triton-Ascend on an Ascend NPU; `torch_npu` for the reference comparison.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Triton-Ascend Fused Attention is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 753 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Triton-Ascend Fused Attention: Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Cuda Index Width (pytorch/pytorch, 104k stars), Liger Kernel Dev (linkedin/Liger-Kernel, 6.7k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Krusty84 (a GitHub user) maintains it in Krusty84/triton-ascend-agent-dev-kit, which has 106 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 15, 2026.
Source: Krusty84/triton-ascend-agent-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.