Agent skill

Triton-Ascend Binned MoE Routing

by Krusty84 in Krusty84/triton-ascend-agent-dev-kit

Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Triton-Ascend Binned MoE Routing

skills CLI
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-binned-moe-routing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Krusty84/triton-ascend-agent-dev-kit build-triton-ascend-binned-moe-routing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/build-triton-ascend-binned-moe-routing .claude/skills/build-triton-ascend-binned-moe-routing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
build-triton-ascend-binned-moe-routing
GitHub stars
106
Token cost
~662 tokens
SKILL.md length
265 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

  • Works in 6 steps: Require bin_ids to be nondecreasing… → Interpret bins[e] as the inclusive… → For gather, compute each assignment's… → …
  • Porting MegaBlocks binned routing to Triton-Ascend
  • SKILL.md covers Goal, Workflow, Implementation Pattern and Ascend Guardrails, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Sorted token-to-expert assignments are mapped into a tensor shaped by experts, expert capacity and hidden size, any overflow past each expert's capacity is dropped, and the inverse scatter and router-weight gradient are implemented. The first step is to require bin_ids to be nondecreasing expert IDs aligned with the indices and to read bins[e] as an inclusive cumulative count.

Gather computes each assignment's offset inside its expert and keeps only offsets below capacity, buffering rows in UB and flushing the buffer when the expert changes within a sub-block. Scatter writes unique assignment slots, optionally scales, and reduces the TOP_K slots per token. The gradient is a dot product of each kept row with its token's gradient. Work is split across vector-core programs, and checks cover monotonic bins, zero-initialized output, aligned tiles and float32 accumulation. Verification compares all three kernels with references on empty experts, overflow, exact capacity and expert transitions inside sub-blocks.

When your agent uses it

  • Porting MegaBlocks binned routing to Triton-Ascend
  • Packing tokens into fixed expert-capacity tensors on an Ascend NPU
  • Restoring top-k token outputs from expert bins
  • Handling sub-blocks that cross expert boundaries in a kernel

Example prompts

  • “Write the binned gather kernel for Triton-Ascend with truncation at expert capacity.”
  • “Implement the router-weight gradient for the binned scatter and add reference comparisons.”
  • “My gather kernel loses rows when a sub-block spans two experts. Fix the buffer flush.”

Requirements

  • Triton-Ascend for an Ascend NPU

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Require bin_ids to be nondecreasing expert IDs aligned with indices.
  2. Interpret bins[e] as the inclusive cumulative assignment count through expert e.
  3. For gather, compute each assignment's offset within its expert and keep only offsets below expert_capacity.
  4. Buffer rows for the current expert in UB; flush them when the expert ID changes inside a sub-block.
  5. For scatter, load valid expert-capacity rows, store to unique assignment slots, optionally scale, and reduce TOP_K slots per token.
  6. For wgrad, dot each kept expert row with the gradient of indices[i] // TOP_K.

What it can do on your machine

Read from SKILL.md and the folder at commit 4ab5ee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton-Ascend Binned MoE Routing loads about 662 tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~662

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Krusty84/triton-ascend-agent-dev-kit at commit 4ab5ee7, republished under its Apache-2.0 licence (© Krusty84). 265 words, ~662 tokens.

Download SKILL.mdSave it as .claude/skills/build-triton-ascend-binned-moe-routing/SKILL.md (or your agent's skills folder).
name
build-triton-ascend-binned-moe-routing
description
Build capacity-limited binned MoE gather, scatter, and router-weight-gradient kernels for Triton-Ascend using sorted expert bins, UB sub-block buffering, and vector-core-aware tiling. Use when an agent ports MegaBlocks binned routing, packs tokens into fixed expert-capacity tensors, restores top-k token outputs, or must handle sub-blocks that cross expert boundaries.

Build Triton-Ascend Binned MoE Routing

Goal

Map sorted token assignments into a tensor shaped experts by expert_capacity by hidden_size, truncate overflow per expert, and implement the inverse scatter and router-weight gradient.

Workflow

  1. Require bin_ids to be nondecreasing expert IDs aligned with indices.
  2. Interpret bins[e] as the inclusive cumulative assignment count through expert e.
  3. For gather, compute each assignment's offset within its expert and keep only offsets below expert_capacity.
  4. Buffer rows for the current expert in UB; flush them when the expert ID changes inside a sub-block.
  5. For scatter, load valid expert-capacity rows, store to unique assignment slots, optionally scale, and reduce TOP_K slots per token.
  6. For wgrad, dot each kept expert row with the gradient of indices[i] // TOP_K.

Implementation Pattern

text
expert_start = 0 if expert == 0 else bins[expert - 1]
offset_in_expert = sorted_position - expert_start
if offset_in_expert < expert_capacity:
    expert_row = expert * expert_capacity + offset_in_expert
    source_token = indices[sorted_position] // TOP_K

Split the sorted-position range across num_vectorcore programs. Within each program, iterate over SUB_BLOCK_SIZE assignments and BLOCK_X feature tiles; use extension.insert_slice for gather buffers and extension.extract_slice for scatter rows.

Ascend Guardrails

  • Validate bins as monotonic, bins[-1] == len(indices), and bin_ids consistent with bin boundaries.
  • Zero-initialize expert output so unused capacity has defined values.
  • Flush a partial UB buffer before switching experts and after the final expert in a sub-block.
  • Never read or scatter assignments beyond expert_capacity.
  • Align BLOCK_X for the input dtype and include bin vectors, UB row buffers, and multibuffering in the memory budget.
  • Require indices to identify unique top-k assignment slots for non-atomic scatter.
  • Accumulate weighted values and wgrad reductions in float32.

Verification

Compare gather, scatter, and wgrad with references across empty experts, expert overflow, exact capacity, sub-block expert transitions, TOP_K values, unaligned hidden sizes, and large token counts.

© Krusty84, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/build-triton-ascend-binned-moe-routing of Krusty84/triton-ascend-agent-dev-kit.

Open the folder on GitHubat commit 4ab5ee7

Compare with similar skills

Triton-Ascend Binned MoE Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton-Ascend Binned MoE Routing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton-Ascend Binned MoE Routing this skillKrusty84/triton-ascend-agent-dev-kit106—~662Automated safety check: PassApache-2.0
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Liger Kernel Devlinkedin/Liger-Kernel6.7k—~799Automated safety check: PassBSD-2-Clause
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0

Similar skills

  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Liger Kernel Dev

    linkedin/Liger-Kernel

    Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.

    6.7k GitHub stars~799 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from Krusty84/triton-ascend-agent-dev-kit

All 23 skills in this repo
  • Triton-Ascend Token Pool Assignment

    Krusty84/triton-ascend-agent-dev-kit

    Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

    106 GitHub stars~564 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Batch Token Reorder

    Krusty84/triton-ascend-agent-dev-kit

    Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

    106 GitHub stars~597 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Attention

    Krusty84/triton-ascend-agent-dev-kit

    Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

    106 GitHub stars~753 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Softmax

    Krusty84/triton-ascend-agent-dev-kit

    Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

    106 GitHub stars~636 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Layer Norm

    Krusty84/triton-ascend-agent-dev-kit

    Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles.

    106 GitHub stars~659 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Moe Gather Scatter

    Krusty84/triton-ascend-agent-dev-kit

    Build Ascend-friendly MoE gather, scatter, and router-weight-gradient kernels with vector-core-sized grids, UB-aware row/column tiling, and CANN slice extensions.

    106 GitHub stars~696 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Triton-Ascend Binned MoE Routing

What does Triton-Ascend Binned MoE Routing do?

Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs. Sorted token-to-expert assignments are mapped into a tensor shaped by experts, expert capacity and hidden size, any overflow past each expert's capacity is dropped, and the inverse scatter and router-weight gradient are implemented. The first step is to require bin_ids to be nondecreasing expert IDs aligned with the indices and to read bins[e] as an inclusive cumulative count.

When should I use Triton-Ascend Binned MoE Routing?

Triton-Ascend Binned MoE Routing fits situations like: porting MegaBlocks binned routing to Triton-Ascend; packing tokens into fixed expert-capacity tensors on an Ascend NPU; restoring top-k token outputs from expert bins; handling sub-blocks that cross expert boundaries in a kernel.

How do I install Triton-Ascend Binned MoE Routing in Claude Code?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-binned-moe-routing -a claude-code`. Or copy the skill folder (skills/build-triton-ascend-binned-moe-routing in Krusty84/triton-ascend-agent-dev-kit) into .claude/skills/build-triton-ascend-binned-moe-routing in your project. Claude Code loads it when a task matches its description.

How do I install Triton-Ascend Binned MoE Routing in Codex?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-binned-moe-routing -a codex`. Or copy the skill folder (skills/build-triton-ascend-binned-moe-routing in Krusty84/triton-ascend-agent-dev-kit) into .agents/skills/build-triton-ascend-binned-moe-routing in your project. Codex loads it when a task matches its description.

Can I use Triton-Ascend Binned MoE Routing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill build-triton-ascend-binned-moe-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-triton-ascend-binned-moe-routing, .gemini/skills/build-triton-ascend-binned-moe-routing, .github/skills/build-triton-ascend-binned-moe-routing and .opencode/skills/build-triton-ascend-binned-moe-routing in your project.

What does Triton-Ascend Binned MoE Routing need to run?

SKILL.md names no scripts, command-line tools or credentials: Triton-Ascend Binned MoE Routing is instructions for the agent only. Our summary lists: Triton-Ascend for an Ascend NPU.

Does Triton-Ascend Binned MoE Routing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton-Ascend Binned MoE Routing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton-Ascend Binned MoE Routing use?

Triton-Ascend Binned MoE Routing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton-Ascend Binned MoE Routing use?

About 662 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton-Ascend Binned MoE Routing?

Skills that share tags, products or a category with Triton-Ascend Binned MoE Routing: Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Cuda Index Width (pytorch/pytorch, 104k stars), Liger Kernel Dev (linkedin/Liger-Kernel, 6.7k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton-Ascend Binned MoE Routing?

Krusty84 (a GitHub user) maintains it in Krusty84/triton-ascend-agent-dev-kit, which has 106 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 15, 2026.

Source: Krusty84/triton-ascend-agent-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.