Agent skill

Triton-Ascend Batch Token Reorder

by Krusty84 in Krusty84/triton-ascend-agent-dev-kit

Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Triton-Ascend Batch Token Reorder

skills CLI
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill batch-triton-ascend-token-reorder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Krusty84/triton-ascend-agent-dev-kit batch-triton-ascend-token-reorder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/batch-triton-ascend-token-reorder .claude/skills/batch-triton-ascend-token-reorder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
batch-triton-ascend-token-reorder
GitHub stars
106
Token cost
~597 tokens
SKILL.md length
194 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

  • Works in 5 steps: Treat input and output as S by D and… → Assign BLOCK_SIZE consecutive output… → Load their indices as one vector. → …
  • Implementing a gather-by-index kernel for MoE token reordering on Ascend NPU
  • SKILL.md covers Goal, Workflow, Implementation Pattern and Ascend Guardrails, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The goal is implementing output[j] = input[indices[j]] on Ascend NPU without falling back to one program per token, which creates too many logical cores and scattered writes on GPU-style tiling. Input and output are treated as an S by D matrix and indices as a length-S integer vector; each program is assigned BLOCK_SIZE consecutive output rows, loads their indices as one vector, loops over the valid rows to load each indexed D-element source row, inserts it into a BLOCK_SIZE by D buffer in UB memory using extension.insert_slice, and stores the assembled tile with row and column masks.

Guardrails call for validating every selected index as within bounds before loading, computing the output row offset as row times D since omitting that multiplication overlaps rows, and sizing BLOCK_SIZE against the UB memory cost of the full tile plus a temporary row. D must be a legal compile-time tile size or padded with column masks. The pattern only removes scattered output writes; source reads stay random. Verification compares results against a plain input[indices] gather across permutations, repeated indices, a partial final block and invalid-index cases that must fail before the kernel launches.

When your agent uses it

  • Implementing a gather-by-index kernel for MoE token reordering on Ascend NPU
  • Replacing a one-program-per-token GPU tiling pattern that creates too many logical cores on Ascend
  • Writing a Triton-Ascend kernel that needs coalesced output stores despite random source reads

Example prompts

  • “Write a Triton-Ascend kernel that reorders tokens by an index array with contiguous output writes.”
  • “My one-program-per-token MoE gather kernel is too slow on Ascend; batch it per program instead.”
  • “Add bounds and overlap checks to this Ascend token-reorder kernel before I test it.”

Requirements

  • Triton-Ascend with the extension.insert_slice and extension.get_element APIs

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Treat input and output as S by D and indices as a length-S integer vector.
  2. Assign BLOCK_SIZE consecutive output rows to each program.
  3. Load their indices as one vector.
  4. Loop over valid rows, load each indexed D-element source row, and insert it into a BLOCK_SIZE by D UB tensor.
  5. Store the assembled tile with row and column masks.

What it can do on your machine

Read from SKILL.md and the folder at commit 4ab5ee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton-Ascend Batch Token Reorder loads about 597 tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 194 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~597

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Krusty84/triton-ascend-agent-dev-kit at commit 4ab5ee7, republished under its Apache-2.0 licence (© Krusty84). 194 words, ~597 tokens.

Download SKILL.mdSave it as .claude/skills/batch-triton-ascend-token-reorder/SKILL.md (or your agent's skills folder).
name
batch-triton-ascend-token-reorder
description
Batch MoE token reordering in Triton-Ascend with contiguous index loads, random source-row reads, extension.insert_slice assembly in UB, and coalesced output stores. Use when an agent must implement output[j] = input[indices[j]] on Ascend NPU and one-program-per-token GPU tiling creates too many logical cores or scattered writes.

Batch Triton-Ascend Token Reorder

Goal

Gather several indexed token rows per program, assemble them in an on-chip buffer, and write one contiguous output tile.

Workflow

  1. Treat input and output as S by D and indices as a length-S integer vector.
  2. Assign BLOCK_SIZE consecutive output rows to each program.
  3. Load their indices as one vector.
  4. Loop over valid rows, load each indexed D-element source row, and insert it into a BLOCK_SIZE by D UB tensor.
  5. Store the assembled tile with row and column masks.

Implementation Pattern

python
import triton.language.extra.cann.extension as extension

rows = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE)
row_mask = rows < S
selected = tl.load(indices + rows, mask=row_mask, other=0)
tile = tl.zeros((BLOCK_SIZE, D), dtype=x_ptr.dtype.element_ty)

for i in range(0, BLOCK_SIZE):
    if pid * BLOCK_SIZE + i < S:
        source_row = extension.get_element(selected, (i,))
        cols = tl.arange(0, D)
        values = tl.load(x_ptr + source_row * D + cols)
        tile = extension.insert_slice(
            tile, values[None, :], (i, 0), (1, D), (1, 1)
        )

cols = tl.arange(0, D)
out_offsets = rows[:, None] * D + cols[None, :]
tl.store(out_ptr + out_offsets, tile, mask=row_mask[:, None])

Ascend Guardrails

  • Use the extension namespace; current APIs are extension.insert_slice and extension.get_element.
  • Validate every selected index as 0 <= index < input_rows before loading.
  • Compute the output row offset as row * D; omitting the multiplication overlaps rows.
  • Size BLOCK_SIZE by the UB cost of the full BLOCK_SIZE by D tile and temporary row.
  • Require D to be a legal compile-time tile or add padded columns and column masks.
  • Use this pattern when output writes are contiguous; it does not remove random source reads.

Verification

Compare with input[indices] for permutations and repeated indices. Include a partial final block and invalid-index tests that must fail before kernel launch.

© Krusty84, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/batch-triton-ascend-token-reorder of Krusty84/triton-ascend-agent-dev-kit.

Open the folder on GitHubat commit 4ab5ee7

Compare with similar skills

Triton-Ascend Batch Token Reorder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton-Ascend Batch Token Reorder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton-Ascend Batch Token Reorder this skillKrusty84/triton-ascend-agent-dev-kit106—~597Automated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
0G Compute Network Guideinternet-court/internet-court-skill6.6k1 repos~1.9kAutomated safety check: PassCustom licence
LLM Serving Capacity PlannerBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.5kAutomated safety check: PassNone
Ascend Model Adapter for vLLMvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • 0G Compute Network Guide

    internet-court/internet-court-skill

    Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands.

    6.6k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    938 GitHub stars~2.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from Krusty84/triton-ascend-agent-dev-kit

All 23 skills in this repo
  • Triton-Ascend Token Pool Assignment

    Krusty84/triton-ascend-agent-dev-kit

    Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

    106 GitHub stars~564 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Binned MoE Routing

    Krusty84/triton-ascend-agent-dev-kit

    Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

    106 GitHub stars~662 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Attention

    Krusty84/triton-ascend-agent-dev-kit

    Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

    106 GitHub stars~753 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Softmax

    Krusty84/triton-ascend-agent-dev-kit

    Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

    106 GitHub stars~636 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Layer Norm

    Krusty84/triton-ascend-agent-dev-kit

    Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles.

    106 GitHub stars~659 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Moe Gather Scatter

    Krusty84/triton-ascend-agent-dev-kit

    Build Ascend-friendly MoE gather, scatter, and router-weight-gradient kernels with vector-core-sized grids, UB-aware row/column tiling, and CANN slice extensions.

    106 GitHub stars~696 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Triton-Ascend Batch Token Reorder

What does Triton-Ascend Batch Token Reorder do?

Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering. The goal is implementing output[j] = input[indices[j]] on Ascend NPU without falling back to one program per token, which creates too many logical cores and scattered writes on GPU-style tiling.insert_slice, and stores the assembled tile with row and column masks.

When should I use Triton-Ascend Batch Token Reorder?

Triton-Ascend Batch Token Reorder fits situations like: implementing a gather-by-index kernel for MoE token reordering on Ascend NPU; replacing a one-program-per-token GPU tiling pattern that creates too many logical cores on Ascend; writing a Triton-Ascend kernel that needs coalesced output stores despite random source reads.

How do I install Triton-Ascend Batch Token Reorder in Claude Code?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill batch-triton-ascend-token-reorder -a claude-code`. Or copy the skill folder (skills/batch-triton-ascend-token-reorder in Krusty84/triton-ascend-agent-dev-kit) into .claude/skills/batch-triton-ascend-token-reorder in your project. Claude Code loads it when a task matches its description.

How do I install Triton-Ascend Batch Token Reorder in Codex?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill batch-triton-ascend-token-reorder -a codex`. Or copy the skill folder (skills/batch-triton-ascend-token-reorder in Krusty84/triton-ascend-agent-dev-kit) into .agents/skills/batch-triton-ascend-token-reorder in your project. Codex loads it when a task matches its description.

Can I use Triton-Ascend Batch Token Reorder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill batch-triton-ascend-token-reorder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/batch-triton-ascend-token-reorder, .gemini/skills/batch-triton-ascend-token-reorder, .github/skills/batch-triton-ascend-token-reorder and .opencode/skills/batch-triton-ascend-token-reorder in your project.

What does Triton-Ascend Batch Token Reorder need to run?

SKILL.md names no scripts, command-line tools or credentials: Triton-Ascend Batch Token Reorder is instructions for the agent only. Our summary lists: Triton-Ascend with the extension.insert_slice and extension.get_element APIs.

Does Triton-Ascend Batch Token Reorder access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton-Ascend Batch Token Reorder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton-Ascend Batch Token Reorder use?

Triton-Ascend Batch Token Reorder is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton-Ascend Batch Token Reorder use?

About 597 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton-Ascend Batch Token Reorder?

Skills that share tags, products or a category with Triton-Ascend Batch Token Reorder: Hugging Face Local Model Evals (huggingface/skills, 11k stars), 0G Compute Network Guide (internet-court/internet-court-skill, 6.6k stars), LLM Serving Capacity Planner (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars) and Ascend Model Adapter for vLLM (vllm-project/vllm-ascend, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton-Ascend Batch Token Reorder?

Krusty84 (a GitHub user) maintains it in Krusty84/triton-ascend-agent-dev-kit, which has 106 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 15, 2026.

Source: Krusty84/triton-ascend-agent-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.