Agent skill

Triton-Ascend Fused Softmax

by Krusty84 in Krusty84/triton-ascend-agent-dev-kit

Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Triton-Ascend Fused Softmax

skills CLI
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-softmax -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-fused-softmax --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/build-triton-ascend-fused-softmax .claude/skills/build-triton-ascend-fused-softmax && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
build-triton-ascend-fused-softmax
GitHub stars
106
Token cost
~636 tokens
SKILL.md length
178 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

  • Works in 5 steps: Interpret the input as n_rows by n_cols… → Choose BLOCK_SIZE as the next power of… → Let each program process rows separated… → …
  • Fusing max, exponentiation, sum and normalization for a 2-D tensor on an Ascend NPU
  • SKILL.md covers Goal, Workflow, Implementation Pattern and Ascend Guardrails, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

This skill shows how to write a softmax kernel over the last dimension of a two-dimensional tensor for Triton-Ascend, fusing the max, exponentiation, sum and normalization steps so each input row is read once and each output row is written once. The steps are to treat the input as n_rows by n_cols and keep its row stride, set BLOCK_SIZE to the next power of two at or above n_cols, and let each program handle rows spaced by the program count.

Rows are loaded padded with -inf so the padding cannot affect the maximum or the denominator, the row maximum is subtracted before exponentiation, and the result is stored with the original mask. Guardrails say to reject zero-column inputs and shapes whose padded row exceeds on-chip resources. To verify, compare against torch.softmax over dim 1 with tolerances suited to the dtype, and test odd shapes such as 1823 by 781, very small rows and rows near the largest BLOCK_SIZE.

When your agent uses it

  • Fusing max, exponentiation, sum and normalization for a 2-D tensor on an Ascend NPU
  • Adapting the row-wise reduction pattern to a similar operator
  • Checking a softmax kernel's padding, stride and mask handling

Example prompts

  • “Write a fused softmax kernel in Triton-Ascend for a 2-D tensor with strided rows.”
  • “Adapt the softmax reduction pattern to a row-wise log-softmax kernel for Ascend.”
  • “Test my Triton softmax against torch.softmax on irregular shapes.”

Requirements

  • Triton-Ascend and an Ascend NPU
  • PyTorch, for the torch.softmax comparison

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Interpret the input as n_rows by n_cols and preserve its row stride.
  2. Choose BLOCK_SIZE as the next power of two greater than or equal to n_cols.
  3. Let each program process rows separated by tl.num_programs(0).
  4. Load a padded row with -inf outside n_cols.
  5. Subtract the row maximum, exponentiate, reduce the sum, normalize, and store with the original mask.

What it can do on your machine

Read from SKILL.md and the folder at commit 4ab5ee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton-Ascend Fused Softmax loads about 636 tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 178 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~636

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Krusty84/triton-ascend-agent-dev-kit at commit 4ab5ee7, republished under its Apache-2.0 licence (© Krusty84). 178 words, ~636 tokens.

Download SKILL.mdSave it as .claude/skills/build-triton-ascend-fused-softmax/SKILL.md (or your agent's skills folder).
name
build-triton-ascend-fused-softmax
description
Build a fused row-wise softmax kernel for Triton-Ascend with stable reductions, power-of-two padding, strided rows, and masked memory operations. Use when an agent needs to fuse max, exponentiation, sum, and normalization for a two-dimensional NPU tensor or adapt this reduction pattern to a similar row-wise operator.

Build Triton-Ascend Fused Softmax

Goal

Compute softmax across the last dimension while reading each input row once and writing each output row once.

Workflow

  1. Interpret the input as n_rows by n_cols and preserve its row stride.
  2. Choose BLOCK_SIZE as the next power of two greater than or equal to n_cols.
  3. Let each program process rows separated by tl.num_programs(0).
  4. Load a padded row with -inf outside n_cols.
  5. Subtract the row maximum, exponentiate, reduce the sum, normalize, and store with the original mask.

Implementation Pattern

python
@triton.jit
def softmax_kernel(out, x, x_stride, out_stride, n_rows, n_cols,
                   BLOCK_SIZE: tl.constexpr):
    first_row = tl.program_id(0)
    row_step = tl.num_programs(0)
    cols = tl.arange(0, BLOCK_SIZE)
    mask = cols < n_cols
    for row in tl.range(first_row, n_rows, row_step):
        values = tl.load(x + row * x_stride + cols, mask=mask,
                         other=-float("inf"))
        shifted = values - tl.max(values, axis=0)
        numerator = tl.exp(shifted)
        result = numerator / tl.sum(numerator, axis=0)
        tl.store(out + row * out_stride + cols, result, mask=mask)


def softmax(x):
    assert x.ndim == 2
    n_rows, n_cols = x.shape
    block = triton.next_power_of_2(n_cols)
    out = torch.empty_like(x)
    programs = min(32, n_rows)
    softmax_kernel[(programs,)](
        out, x, x.stride(0), out.stride(0), n_rows, n_cols,
        BLOCK_SIZE=block,
    )
    return out

Ascend Guardrails

  • Use -inf for padded values so they cannot affect the maximum or denominator.
  • Perform the max shift before tl.exp; Triton exponentiation is fast and approximate.
  • Reject n_cols == 0 and shapes whose padded row exceeds available on-chip resources.
  • Preserve explicit strides or require a contiguous last dimension.
  • Keep the program count independent of n_rows when using the row-striding loop; cap it by n_rows.

Verification

Compare with torch.softmax(x, dim=1) using dtype-appropriate tolerances. Test irregular dimensions such as 1823 by 781, very small rows, and rows near the largest supported BLOCK_SIZE.

© Krusty84, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/build-triton-ascend-fused-softmax of Krusty84/triton-ascend-agent-dev-kit.

Open the folder on GitHubat commit 4ab5ee7

Compare with similar skills

Triton-Ascend Fused Softmax next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton-Ascend Fused Softmax compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton-Ascend Fused Softmax this skillKrusty84/triton-ascend-agent-dev-kit106—~636Automated safety check: PassApache-2.0
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT
LLM Serving Framework BenchmarkBBuf/AI-Infra-Auto-Driven-SKILLS938—~7.5kAutomated safety check: PassNone
Magpie Kernel Evaluatoramd/skills408—~2.3kAutomated safety check: PassMIT
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0

Similar skills

  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 4 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Serving Framework Benchmark

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

    938 GitHub stars~7.5k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    408 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed

More from Krusty84/triton-ascend-agent-dev-kit

All 23 skills in this repo
  • Triton-Ascend Token Pool Assignment

    Krusty84/triton-ascend-agent-dev-kit

    Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

    106 GitHub stars~564 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Batch Token Reorder

    Krusty84/triton-ascend-agent-dev-kit

    Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

    106 GitHub stars~597 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Binned MoE Routing

    Krusty84/triton-ascend-agent-dev-kit

    Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

    106 GitHub stars~662 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Attention

    Krusty84/triton-ascend-agent-dev-kit

    Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

    106 GitHub stars~753 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Layer Norm

    Krusty84/triton-ascend-agent-dev-kit

    Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles.

    106 GitHub stars~659 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Moe Gather Scatter

    Krusty84/triton-ascend-agent-dev-kit

    Build Ascend-friendly MoE gather, scatter, and router-weight-gradient kernels with vector-core-sized grids, UB-aware row/column tiling, and CANN slice extensions.

    106 GitHub stars~696 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Triton-Ascend Fused Softmax

What does Triton-Ascend Fused Softmax do?

Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs. This skill shows how to write a softmax kernel over the last dimension of a two-dimensional tensor for Triton-Ascend, fusing the max, exponentiation, sum and normalization steps so each input row is read once and each output row is written once. The steps are to treat the input as n_rows by n_cols and keep its row stride, set BLOCK_SIZE to the next power of two at or above n_cols, and let each program handle rows spaced by the program count.

When should I use Triton-Ascend Fused Softmax?

Triton-Ascend Fused Softmax fits situations like: fusing max, exponentiation, sum and normalization for a 2-D tensor on an Ascend NPU; adapting the row-wise reduction pattern to a similar operator; checking a softmax kernel's padding, stride and mask handling.

How do I install Triton-Ascend Fused Softmax in Claude Code?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-softmax -a claude-code`. Or copy the skill folder (skills/build-triton-ascend-fused-softmax in Krusty84/triton-ascend-agent-dev-kit) into .claude/skills/build-triton-ascend-fused-softmax in your project. Claude Code loads it when a task matches its description.

How do I install Triton-Ascend Fused Softmax in Codex?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-softmax -a codex`. Or copy the skill folder (skills/build-triton-ascend-fused-softmax in Krusty84/triton-ascend-agent-dev-kit) into .agents/skills/build-triton-ascend-fused-softmax in your project. Codex loads it when a task matches its description.

Can I use Triton-Ascend Fused Softmax in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-fused-softmax -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-triton-ascend-fused-softmax, .gemini/skills/build-triton-ascend-fused-softmax, .github/skills/build-triton-ascend-fused-softmax and .opencode/skills/build-triton-ascend-fused-softmax in your project.

What does Triton-Ascend Fused Softmax need to run?

SKILL.md names no scripts, command-line tools or credentials: Triton-Ascend Fused Softmax is instructions for the agent only. Our summary lists: Triton-Ascend and an Ascend NPU; PyTorch, for the torch.softmax comparison.

Does Triton-Ascend Fused Softmax access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton-Ascend Fused Softmax safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton-Ascend Fused Softmax use?

Triton-Ascend Fused Softmax is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton-Ascend Fused Softmax use?

About 636 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton-Ascend Fused Softmax?

Skills that share tags, products or a category with Triton-Ascend Fused Softmax: TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars), LLM Serving Framework Benchmark (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Magpie Kernel Evaluator (amd/skills, 408 stars) and Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton-Ascend Fused Softmax?

Krusty84 (a GitHub user) maintains it in Krusty84/triton-ascend-agent-dev-kit, which has 106 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 15, 2026.

Source: Krusty84/triton-ascend-agent-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.