Agent skill

Triton-Ascend Token Pool Assignment

by Krusty84 in Krusty84/triton-ascend-agent-dev-kit

Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Triton-Ascend Token Pool Assignment

skills CLI
$ npx skills add Krusty84/triton-ascend-agent-dev-kit --skill assign-triton-ascend-token-pools -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Krusty84/triton-ascend-agent-dev-kit assign-triton-ascend-token-pools --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Krusty84/triton-ascend-agent-dev-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/assign-triton-ascend-token-pools .claude/skills/assign-triton-ascend-token-pools && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
assign-triton-ascend-token-pools
GitHub stars
106
Token cost
~564 tokens
SKILL.md length
169 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

  • Works in 5 steps: Launch one program per request. → Load kv_start, kv_end, and the… → Sum end-start for all earlier requests… → …
  • Porting SGLang-style KV cache assignment to Ascend NPU kernels
  • SKILL.md covers Goal, Workflow, Implementation Pattern and Ascend Guardrails, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The goal is to copy each request's contiguous cache segment into the correct row and token range of a request-to-token table, porting SGLang-style KV cache assignment to Ascend hardware. The workflow launches one program per request, loads its kv_start, kv_end and destination pool row, sums the lengths of all earlier requests to find this request's source offset in the concatenated cache, and copies the resulting span in fixed BLOCK_SIZE chunks using int32 metadata when every length, offset and pool index fits.

A set of Ascend-specific guardrails follows: use saved offsets rather than the first block's offsets when masking every loop iteration, compute each iteration's offset as i times BLOCK_SIZE plus a base rather than mutating a vector offset across iterations, keep BS_UPPER a safe compile-time bound with entries beyond the program ID masked out, validate that start is less than or equal to end within the pool length, and switch to int64 whenever a prefix sum or flattened address could exceed int32 range.

Verification compares the kernel's output against a plain CPU loop over requests, specifically exercising spans longer than BLOCK_SIZE, a partial final chunk, empty spans, non-uniform lengths and a boundary case near int32 overflow.

When your agent uses it

  • Porting SGLang-style KV cache assignment to Ascend NPU kernels
  • Replacing a loop-carried vector offset that compiles poorly on Ascend
  • Copying variable-length request spans out of a concatenated cache correctly

Example prompts

  • “Port this SGLang token pool assignment kernel to Triton-Ascend.”
  • “Fix the loop-carried offset bug in our Ascend token pool kernel.”
  • “Verify the token pool kernel against a CPU loop, including the int32 boundary.”

Requirements

  • A Triton-Ascend development setup targeting Ascend NPU hardware

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Launch one program per request.
  2. Load kv_start, kv_end, and the destination request-pool row.
  3. Sum end-start for all earlier requests to find this request's source offset in the concatenated cache.
  4. Use int32 metadata when all lengths, offsets, and pool indices fit.
  5. Copy kv_end-kv_start elements in fixed BLOCK_SIZE chunks with derived per-iteration offsets.

What it can do on your machine

Read from SKILL.md and the folder at commit 4ab5ee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Triton-Ascend Token Pool Assignment loads about 564 tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 169 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~564

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Krusty84/triton-ascend-agent-dev-kit at commit 4ab5ee7, republished under its Apache-2.0 licence (© Krusty84). 169 words, ~564 tokens.

Download SKILL.mdSave it as .claude/skills/assign-triton-ascend-token-pools/SKILL.md (or your agent's skills folder).
name
assign-triton-ascend-token-pools
description
Implement and optimize Triton-Ascend request-to-token-pool assignment using int32 metadata, masked prefix-length reduction, and fixed-offset block loops. Use when an agent ports SGLang-style KV cache assignment to Ascend NPU, copies variable request spans from a concatenated cache, or needs to replace loop-carried vector offsets that compile poorly.

Assign Triton-Ascend Token Pools

Goal

Copy each request's contiguous cache segment into the correct row and token range of a request-to-token table.

Workflow

  1. Launch one program per request.
  2. Load kv_start, kv_end, and the destination request-pool row.
  3. Sum end-start for all earlier requests to find this request's source offset in the concatenated cache.
  4. Use int32 metadata when all lengths, offsets, and pool indices fit.
  5. Copy kv_end-kv_start elements in fixed BLOCK_SIZE chunks with derived per-iteration offsets.

Implementation Pattern

python
prefix_offsets = tl.arange(0, BS_UPPER)
prior = prefix_offsets < pid
starts = tl.load(start_ptr + prefix_offsets, mask=prior, other=0)
ends = tl.load(end_ptr + prefix_offsets, mask=prior, other=0)
source_start = tl.sum(ends - starts, axis=0)

base = tl.arange(0, BLOCK_SIZE)
num_loops = tl.cdiv(kv_end - kv_start, BLOCK_SIZE)
for i in range(num_loops):
    load_offsets = source_start + i * BLOCK_SIZE + base
    save_offsets = kv_start + i * BLOCK_SIZE + base
    mask = save_offsets < kv_end
    values = tl.load(out_cache + load_offsets, mask=mask)
    tl.store(token_pool + save_offsets, values, mask=mask)

Ascend Guardrails

  • Use save_offsets, not the initial block's offsets, when masking every loop iteration.
  • Prefer i * BLOCK_SIZE + base over mutating vector offsets across iterations.
  • Keep BS_UPPER a safe compile-time bound and mask entries beyond pid.
  • Validate 0 <= kv_start <= kv_end <= pool_len and ensure destination request rows are unique or intentionally synchronized.
  • Keep int64 when any prefix sum or flattened pool address can exceed int32.

Verification

Compare with a CPU loop over requests. Include spans longer than BLOCK_SIZE, a partial final chunk, empty spans, nonuniform lengths, and an int32-overflow boundary.

© Krusty84, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/assign-triton-ascend-token-pools of Krusty84/triton-ascend-agent-dev-kit.

Open the folder on GitHubat commit 4ab5ee7

Compare with similar skills

Triton-Ascend Token Pool Assignment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Triton-Ascend Token Pool Assignment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Triton-Ascend Token Pool Assignment this skillKrusty84/triton-ascend-agent-dev-kit106—~564Automated safety check: PassApache-2.0
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Leetcuda Cpp Kernelxlite-dev/LeetCUDA12k—~3.5kAutomated safety check: PassGPL-3.0
Liger Autopatchlinkedin/Liger-Kernel6.7k—~1.3kAutomated safety check: PassBSD-2-Clause

Similar skills

  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Leetcuda Cpp Kernel

    xlite-dev/LeetCUDA

    LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、…

    12k GitHub stars~3.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Liger Autopatch

    linkedin/Liger-Kernel

    Adds Liger Kernel support for a new HuggingFace Transformers model, or modifies existing monkey-patching.

    6.7k GitHub stars~1.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • 0G Compute Network Guide

    internet-court/internet-court-skill

    Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands.

    6.6k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed

More from Krusty84/triton-ascend-agent-dev-kit

All 23 skills in this repo
  • Triton-Ascend Batch Token Reorder

    Krusty84/triton-ascend-agent-dev-kit

    Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

    106 GitHub stars~597 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Binned MoE Routing

    Krusty84/triton-ascend-agent-dev-kit

    Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

    106 GitHub stars~662 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Attention

    Krusty84/triton-ascend-agent-dev-kit

    Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

    106 GitHub stars~753 tokensUpdated 1 mo ago
    Auto-check passed
  • Triton-Ascend Fused Softmax

    Krusty84/triton-ascend-agent-dev-kit

    Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

    106 GitHub stars~636 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Layer Norm

    Krusty84/triton-ascend-agent-dev-kit

    Build a fused forward LayerNorm kernel for Triton-Ascend with row-wise mean and variance reductions, float32 accumulation, affine weight and bias, and masked feature tiles.

    106 GitHub stars~659 tokensUpdated 1 mo ago
    Auto-check passed
  • Build Triton Ascend Moe Gather Scatter

    Krusty84/triton-ascend-agent-dev-kit

    Build Ascend-friendly MoE gather, scatter, and router-weight-gradient kernels with vector-core-sized grids, UB-aware row/column tiling, and CANN slice extensions.

    106 GitHub stars~696 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Triton-Ascend Token Pool Assignment

What does Triton-Ascend Token Pool Assignment do?

Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably. The goal is to copy each request's contiguous cache segment into the correct row and token range of a request-to-token table, porting SGLang-style KV cache assignment to Ascend hardware. The workflow launches one program per request, loads its kv_start, kv_end and destination pool row, sums the lengths of all earlier requests to find this request's source offset in the concatenated cache, and copies the resulting span in fixed BLOCK_SIZE chunks using int32 metadata when every length, offset and pool index fits.

When should I use Triton-Ascend Token Pool Assignment?

Triton-Ascend Token Pool Assignment fits situations like: porting SGLang-style KV cache assignment to Ascend NPU kernels; replacing a loop-carried vector offset that compiles poorly on Ascend; copying variable-length request spans out of a concatenated cache correctly.

How do I install Triton-Ascend Token Pool Assignment in Claude Code?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill assign-triton-ascend-token-pools -a claude-code`. Or copy the skill folder (skills/assign-triton-ascend-token-pools in Krusty84/triton-ascend-agent-dev-kit) into .claude/skills/assign-triton-ascend-token-pools in your project. Claude Code loads it when a task matches its description.

How do I install Triton-Ascend Token Pool Assignment in Codex?

Run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill assign-triton-ascend-token-pools -a codex`. Or copy the skill folder (skills/assign-triton-ascend-token-pools in Krusty84/triton-ascend-agent-dev-kit) into .agents/skills/assign-triton-ascend-token-pools in your project. Codex loads it when a task matches its description.

Can I use Triton-Ascend Token Pool Assignment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Krusty84/triton-ascend-agent-dev-kit --skill assign-triton-ascend-token-pools -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/assign-triton-ascend-token-pools, .gemini/skills/assign-triton-ascend-token-pools, .github/skills/assign-triton-ascend-token-pools and .opencode/skills/assign-triton-ascend-token-pools in your project.

What does Triton-Ascend Token Pool Assignment need to run?

SKILL.md names no scripts, command-line tools or credentials: Triton-Ascend Token Pool Assignment is instructions for the agent only. Our summary lists: A Triton-Ascend development setup targeting Ascend NPU hardware.

Does Triton-Ascend Token Pool Assignment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Triton-Ascend Token Pool Assignment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Triton-Ascend Token Pool Assignment use?

Triton-Ascend Token Pool Assignment is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Triton-Ascend Token Pool Assignment use?

About 564 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Triton-Ascend Token Pool Assignment?

Skills that share tags, products or a category with Triton-Ascend Token Pool Assignment: Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Cuda Index Width (pytorch/pytorch, 104k stars) and Leetcuda Cpp Kernel (xlite-dev/LeetCUDA, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Triton-Ascend Token Pool Assignment?

Krusty84 (a GitHub user) maintains it in Krusty84/triton-ascend-agent-dev-kit, which has 106 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on August 15, 2026.

Source: Krusty84/triton-ascend-agent-dev-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.