Official agent skill

Tma Illegal Instruction

by facebookexperimental in facebookexperimental/triton

Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…

OfficialMITAuto-check passedAI & LLM Engineering

Install Tma Illegal Instruction

skills CLI
$ npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton tma-illegal-instruction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tma-illegal-instruction .claude/skills/tma-illegal-instruction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tma-illegal-instruction
GitHub stars
201
Token cost
~1.1k tokens
SKILL.md length
550 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…

  • Works in 4 steps: Find the faoiling TMA p. From the stack… → Reconstruct the failing tile's starting… → Confirm by debug messaging. Determine… → …
  • The user reports CUDA error 716
  • SKILL.md covers Symptom, Diagnosis ladder, Anti-pattern: "just add a mask" and Verify the fix
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tma Illegal Instruction is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store, tl.asyncdescriptorload, async TMA copies) as the source code line. Use when the user reports CUDA error 716, "an illegal instruction was encountered", segfault inside a TMA op, kernel hang followed by an illegal instruction trap, or a crash that only fires on the first or last tile of a launch. Covers the pattern where a TMA…

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering GPU and accelerator computing and Root cause analysis. It works with CUDA. The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • The user reports CUDA error 716
  • An illegal instruction was encountered
  • Segfault inside a TMA op
  • Kernel hang followed by an illegal instruction trap

Example prompts

  • “illegal instruction”
  • “an illegal instruction was encountered”
  • “missing in-kernel mask”
  • “/tma-illegal-instruction”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Find the faoiling TMA p. From the stack trace / sanitizer output / IR
  2. Reconstruct the failing tile's starting offset. For the failing
  3. Confirm by debug messaging. Determine either the grid or value
  4. Only after the structural bug is identified, determine whether the right

What it can do on your machine

Read from SKILL.md and the folder at commit 6f3dd70. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tma Illegal Instruction loads about 1.1k tokens when it runs. Until then it costs about 199 tokens; SKILL.md has 550 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~199
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 6f3dd70, republished under its MIT licence (© facebookexperimental). 550 words, ~1,104 tokens.

Download SKILL.mdSave it as .claude/skills/tma-illegal-instruction/SKILL.md (or your agent's skills folder).
name
tma-illegal-instruction
description
Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (`make_tensor_descriptor`, `TensorDescriptor`, `descriptor.load`, `descriptor.store`, `tl.async_descriptor_load`, async TMA copies) as the source code line. Use when the user reports CUDA error 716, "an illegal instruction was encountered", segfault inside a TMA op, kernel hang followed by an illegal instruction trap, or a crash that only fires on the first or last tile of a launch. Covers the pattern where a TMA store/load is issued at an offset entirely past a tensor's shape — TMA does NOT silently mask out-of-bounds tile accesses; it traps. The root cause is almost never "missing in-kernel mask" — it is commonly a structural launcher / tile-mapping bug.

TMA Illegal Instruction

Symptom

CUDA reports "an illegal instruction was encountered" (error 716), or the kernel crashes inside a TMA op, on a Triton kernel that uses TMA descriptors (TensorDescriptor, tl.make_tensor_descriptor, desc.load(...), desc.store(...), async TMA copies, etc.).

The crash is likely tile-dependent — appears only at certain grid values. This is likely because the tile out of bounds is entirely past the shape of the TME store.

Diagnosis ladder

Walk these in order. Don't skip ahead — the first check is the cheapest and the most often correct.

  1. Find the faoiling TMA p. From the stack trace / sanitizer output / IR dump, identify which descriptor.load(...) or descriptor.store(...) crashed. Note the offsets it was called with (e.g. [pid_m * BM, pid_n * BN]) and the descriptor's declared shape.

  2. Reconstruct the failing tile's starting offset. For the failing program/iteration, compute the literal integer offsets passed to the TMA op. For each axis i of the descriptor, ask: is off_i >= shape_i? If yes, that is the bug. The launcher / tile-mapping logic put a program in a region that does not exist.

  3. Confirm by debug messaging. Determine either the grid or value (could be a jagged tensor) information that is causing the failure. Add a tl.device_print call to the kernel with an if that skips the operation. NOTE: This is the not a proper solution!

  4. Only after the structural bug is identified, determine whether the right fix is launcher/grid dependent or runtime data dependent. If the latter, identify how this shape can be reached.

Anti-pattern: "just add a mask"

The common temptation is to wrap the failing TMA op in if off_m < M and off_n < N: (or to fall back to tl.load with a mask). Resist this. It silences the symptom but:

  • Hides the structural bug — the kernel is still launching programs that own no work, wasting a CTA per stray program.
  • Often masks correctness issues elsewhere — if the kernel reached an out-of-bounds tile, the tile_id it computed for the previous tiles is also suspect.
  • For epilogue stores, the masked-out tile's accumulator was still computed from junk loads further up the kernel — meaning some other tile may have written wrong data that the mask doesn't catch.

In-kernel masks are fine for genuinely ragged shapes (real K not a multiple of BLOCK_K, etc.), but a TMA illegal instruction is a different signal — it says "the launch contract is wrong", not "this iteration is ragged".

Show full SKILL.md (152 more words)Show less

Verify the fix

For the failing tile/iteration, the kernel should be able to assert off_i < shape_i for every TMA op. The verification protocol:

  1. Add temporary tl.device_assert(off_i < shape_i, "...") calls (or print the offsets) before the suspected TMA op and re-run with the same shape that crashed.
  2. Confirm the assert fires at the same iteration the illegal instruction was hitting — that proves you found the actual offending access.
  3. Apply the structural fix (launcher / grid / descriptor).
  4. Re-run the same shape: the asserts no longer fire and the illegal instruction is gone. If the asserts pass but the crash remains, it is a different TMA op or a different bug class — go back to step 1 of the diagnosis ladder.

Removing tl.device_assert after verification is required; the structural fix is what you ship. The code should NOT introduce a new if statement directly over just the TMA operation (that is typically wrong).

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/tma-illegal-instruction of facebookexperimental/triton.

Open the folder on GitHubat commit 6f3dd70

Compare with similar skills

Tma Illegal Instruction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tma Illegal Instruction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tma Illegal Instruction this skillfacebookexperimental/triton201—~1.1kAutomated safety check: PassMIT
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0
Cudasablin39/tilelang-cuda-skills145—~990Automated safety check: PassNone
ONNX Runtime CUDA Attention Patternsmicrosoft/onnxruntime22k—~6.5kAutomated safety check: PassMIT
Doca Gpunetio Ib Write LatNVIDIA/skills3.6k—~3.8kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuda

    sablin39/tilelang-cuda-skills

    Draft, debug, and measure CUDA kernels and host launch workflows.

    145 GitHub stars~990 tokensUpdated 26 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

    22k GitHub stars~6.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under…

    3.6k GitHub stars~3.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuda Profiling

    mohitmishra786/low-level-dev-skills

    CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check: notes

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed

Works with

Questions about Tma Illegal Instruction

What does Tma Illegal Instruction do?

Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…. Tma Illegal Instruction is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization.asyncdescriptorload, async TMA copies) as the source code line.

When should I use Tma Illegal Instruction?

Tma Illegal Instruction fits situations like: the user reports CUDA error 716; an illegal instruction was encountered; segfault inside a TMA op; kernel hang followed by an illegal instruction trap.

How do I install Tma Illegal Instruction in Claude Code?

Run `npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a claude-code`. Or copy the skill folder (.claude/skills/tma-illegal-instruction in facebookexperimental/triton) into .claude/skills/tma-illegal-instruction in your project. Claude Code loads it when a task matches its description.

How do I install Tma Illegal Instruction in Codex?

Run `npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a codex`. Or copy the skill folder (.claude/skills/tma-illegal-instruction in facebookexperimental/triton) into .agents/skills/tma-illegal-instruction in your project. Codex loads it when a task matches its description.

Can I use Tma Illegal Instruction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tma-illegal-instruction, .gemini/skills/tma-illegal-instruction, .github/skills/tma-illegal-instruction and .opencode/skills/tma-illegal-instruction in your project.

What does Tma Illegal Instruction need to run?

SKILL.md names no scripts, command-line tools or credentials: Tma Illegal Instruction is instructions for the agent only.

Does Tma Illegal Instruction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tma Illegal Instruction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tma Illegal Instruction use?

Tma Illegal Instruction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tma Illegal Instruction use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tma Illegal Instruction?

Skills that share tags, products or a category with Tma Illegal Instruction: LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars), Cuda (sablin39/tilelang-cuda-skills, 145 stars) and ONNX Runtime CUDA Attention Patterns (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tma Illegal Instruction?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 10, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.